July 24, 2026
Hack flop or benchmark bop?
UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
Experts say Kimi K3 lags badly at cyber break-ins, but commenters smell mystery and spin
TLDR: UK and U.S. evaluators say Kimi K3 is much weaker than the best American AI models at cyber attack tasks, though it still beat one Chinese rival. Commenters are split between “that’s a big gap” and “these vague charts and odd testing choices make the whole report hard to trust.”
A new joint review from the UK’s AI safety team and a U.S. standards group basically says Moonshot AI’s shiny new model, Kimi K3, is not exactly the cyber supervillain some feared. In tests where models were asked to build hacks or move through a fake company network, Kimi K3 trailed the strongest American systems by a wide margin. It still beat another Chinese model, GLM-5.2, but the headline result was blunt: Kimi K3 is behind the frontier pack. One eyebrow-raising detail did get people talking: its safety filters still allowed it to help with offensive cyber tasks during testing.
But the real fireworks were in the comments, where readers immediately put the report itself on trial. One camp was annoyed that the charts lumped everything into vague labels like “Top U.S. Models”, with one commenter fuming that the graphs felt like “two mystery bars, U.S. and China.” Another argued this is just more proof that top closed models still have a huge edge when the tests get serious. Then the skeptics barged in: maybe the evaluation underestimated Kimi K3, maybe the model is unusually “token-hungry,” maybe it hit limits before showing its full strength. And because this is the internet, there was also a deliciously dramatic insinuation that an earlier Moonshot model might have been trained for cyber attacks. Add in confusion over whether this report was even supposed to exist yet, and suddenly the benchmark wasn’t the only thing getting stress-tested.
Key Points
- •UK AISI and the U.S. CAISI jointly conducted a preliminary cyber-capability evaluation of Moonshot AI’s Kimi K3.
- •The assessment found Kimi K3 performed significantly below the most recent frontier cyber-capable models on exploit development and simulated network attack tasks.
- •On “The Last Ones” cyber range task, Kimi K3 reached step 17 of 32 on average, compared with 28.5 steps for the most cyber-capable U.S. models.
- •Kimi K3 performed above GLM-5.2 on the same preliminary cyber evaluations run by UK AISI and CAISI.
- •The report states Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during evaluation.