July 25, 2026
Brains, budgets, and benchmark beef
ARC-AGI Leaderboard
One model is crushing the chart — and commenters are already yelling "rigged"
TLDR: ARC-AGI’s new leaderboard tracks which AI systems solve fresh problems best without costing a fortune, and one model appears way ahead. Commenters are split between amazement, skepticism, and conspiracy-level side-eye, with many questioning whether leaderboard wins mean real-world usefulness.
The new ARC-AGI leaderboard is supposed to be a serious scorecard for how well artificial intelligence handles brand-new problems while staying cheap and efficient. In plain English: it’s not just about getting the right answer, it’s about doing it without burning a mountain of money and computer power. But the community took one look at the chart and immediately turned it into a full-on comment-section soap opera.
The biggest gasp came from people staring at the gap between Opus 5 and the rest of the pack. One commenter said it’s "actually crazy" once you look at the tasks themselves, basically treating the chart like a shock upset at a sports final. But the applause was quickly drowned out by suspicion. Some users wondered whether these systems are just becoming professional puzzle grinders instead of genuinely smarter tools, with one bluntly saying they suspect the models are "just trained on puzzles by now." Ouch.
Then came the conspiracy-flavored drama. One commenter demanded to know why Fable wasn’t on the board at all, then launched into a wild theory that it was so strong it has become an unofficial ceiling no future system is allowed to cross. Meanwhile, another user delivered the painfully relatable gripe of the AI era: why do some models dominate benchmarks, yet somehow people still drift back to older favorites for actual day-to-day work? And just to pour more fuel on the fire, someone chimed in that the same model is also crushing Frontier-Bench, because apparently one leaderboard wasn’t enough to start the discourse.
Key Points
- •ARC-AGI has evolved from ARC-AGI-1 and ARC-AGI-2, which measured passive fluid intelligence, to ARC-AGI-3, which tests adaptive behavior in novel interactive environments.
- •The leaderboard uses a scatter plot to compare cost per task against performance as a measure of efficiency.
- •Reasoning Systems are shown as connected points for the same model at different reasoning levels, illustrating how added reasoning time affects results.
- •Base LLM solutions are defined as single-shot inference from standard models such as GPT-4.5 and Claude 3.7 without extended reasoning.
- •Kaggle Systems are competition-grade entries operating under a stated $50 compute budget for 120 evaluation tasks and are designed specifically for the ARC Prize.