August 1, 2026

Benchmark? More like bench-brawl

Running Kimi K3 on MI355X at Better Performance per Dollar Than B300

AMD says it found the bargain winner, but the comments section screamed “slop”

TLDR: Wafer says AMD’s MI355X runs a massive new AI model for less money than Nvidia’s B300, even if Nvidia is still faster overall. Commenters immediately turned it into a roast, calling the article “AI slop” and arguing the headline oversold AMD’s win.

A fresh benchmark claimed AMD’s MI355X chips can run the giant new open model Kimi K3 for better bang for your buck than Nvidia’s flashy B300, and on paper that’s a big deal. The model is enormous, expensive to run, and basically the kind of artificial intelligence system that makes companies sweat over power bills. Wafer’s pitch was simple: AMD may not be the absolute speed king, but if you care about what you get for the money, it’s suddenly looking very hard to ignore.

But the real fireworks were in the comments, where readers barely argued about the numbers before going straight for the article’s throat. Multiple posters called it “AI slop,” accusing the write-up of feeling machine-generated, under-edited, and weirdly polished in all the wrong places. One commenter mocked the telltale em-dashes as proof of chatbot fingerprints, while another said if the work is real, the writing made it harder to trust. Ouch.

Then came the counterpunch: some readers argued the bigger issue wasn’t the prose, but the framing. Critics pointed out that Nvidia’s pricier B300 still wins on raw speed in several lines of the table, so calling AMD the winner felt like a marketing spin built around cost math. In other words, this became less a story about chips and more a classic internet cage match: Is AMD actually winning, or did the benchmark headline just win the PR war?

Key Points

  • The article says Kimi K3 has 2.8 trillion parameters and requires more than 1.5TB of VRAM before allocating KV cache for a 1 million-token context window.
  • Wafer reports that an 8x AMD MI355X setup achieved 952 tokens per second per node and 118 tokens per second single-stream on a 1,024-token input and 400-token output benchmark.
  • The article compares this with a two-node, 16-GPU Nvidia B200 deployment at 498 aggregate tokens per second total and a B300 node at 1,568 aggregate tokens per second.
  • According to the article, MI355X matches B300 at 288GB VRAM per GPU and is cheaper in the cited pricing, leading to higher performance per dollar despite lower raw throughput than B300.
  • The article says throughput improvements relied on speculative decoding with RadixArk's Kimi-K3-DSpark and a PyTorch-based fix for a ROCm top-k renormalization issue in sglang.

Hottest takes

"AI slop" — veber-alex
"They even left the em-dashes. Absolute slop" — muragekibicho
"On every row the B300 beat the MI355X" — villgax
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.