July 21, 2026

Fast, Furious, and 56 Minutes Late

Kimi K3: second only to Fable 5 on AA-Briefcase

It nearly tops the charts, but commenters are side-eyeing the price, speed, and the whole test

TLDR: Kimi K3 scored near the very top in a major artificial intelligence office-work test, but it’s pricey and slow. Commenters weren’t just debating whether it’s good — they were fighting over whether the test is fair, private, and worth the money at all.

Kimi K3 just pulled off a flashy entrance: it landed second place on AA-Briefcase, a test that tries to measure how well artificial intelligence handles real office-style work like making spreadsheets, slides, and mock-ups. On paper, that sounds huge — it sits just behind Fable 5 and leaps far ahead of the previous Kimi model. But the comments? Oh, the comments were not ready to simply clap and move on.

The biggest mood in the room was a mix of impressed, suspicious, and extremely budget-conscious. One camp was basically saying, “Nice trophy, but why does it take almost an hour and cost over ten bucks per task?” Another immediately started arguing that the benchmark itself might be shaping the results too much, with one commenter reviving the old debate over the testing setup and claiming different setups used to swing results by more than 10%. In other words: is Kimi amazing, or just amazing under this exact spotlight?

Then came the geopolitical popcorn. One commenter wondered why markets aren’t freaking out the way they did when DeepSeek dropped, now that Chinese models seem to be closing in on top U.S. systems. Others played the optimist, predicting Kimi could become dirt cheap once more providers offer it — the dream of a true “bicycle for the mind.” And naturally, the spreadsheet goblins arrived: one user complained the post was “totally missing cost efficiency,” basically demanding a full value-for-money scorecard before handing out any crowns. So yes, Kimi K3 is a star — but in the comments section, it’s also the star of a very messy reunion special.

Key Points

  • Moonshot AI’s Kimi K3 scored 1543 Elo on Artificial Analysis’ AA-Briefcase benchmark, the second-highest result reported and behind only Claude Fable 5 at 1574.
  • The article says Kimi K3 improved by 727 Elo over the previous-generation Kimi K2.6, rising from 816 to 1543 on AA-Briefcase.
  • AA-Briefcase evaluates agentic knowledge-work tasks on a private dataset involving complex files and deliverables such as spreadsheets, presentations, and UI mock-ups.
  • Kimi K3 achieved a 51% rubric pass rate and an analytical quality Elo of 1754, but its presentation quality Elo of 1471 lagged some competing models.
  • Kimi K3 averaged $10.57 per task and 56.4 minutes per task, driven by 120k output tokens and 83 turns per task, making it relatively expensive and slow in this benchmark.

Hottest takes

"a harness would often affect results ... for more than 10%" — eugene3306
"Can we assume that the test is still private when it was run on many cloud providers?" — threatripper
"totally missing cost efficiency" — felufe
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.