August 4, 2026
Benchmarked and dragged
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
AI scorecards are maxing out, and the internet is already yelling "slop"
TLDR: Researchers found that many popular AI tests are wearing out, with nearly half no longer doing a good job separating the best systems from the rest. In the comments, people fought over whether this means AI progress is stalling, while critics mocked the paper as "slop" and others argued better tests already exist.
A new study says the report cards used to judge artificial intelligence are getting too easy to impress. Looking across 60 tests, researchers found that nearly half are basically hitting the point where newer models all look similarly great, which makes it harder to tell who is actually improving and who is just gaming the quiz. The big twist: the paper argues that carefully designed, expert-made tests hold up better over time than rushed or recycled ones.
But the real fireworks were in the comments, where the community instantly split into camps. One side went full doom mode, with one poster declaring this could be "the end of the road" for large language models, arguing we may be squeezing the last drops out of current methods. Another side was far less philosophical and far more brutal, dismissing the whole thing as "slop" and even roasting the charts for looking like untouched default plotting software. Ouch.
Then came the builders, who basically said: calm down, there are better ways to test these systems. One team pointed to multiplayer game-style evaluations, where AI agents have to cooperate or compete in open-ended settings, saying those feel harder to fake and closer to real-world ability. Another commenter dropped a plug for Agents Last Exam as proof there is still room to grow. So yes, the paper says AI tests are getting stale — but the crowd turned it into a full-on brawl over whether AI is plateauing, whether researchers are phoning it in, and what kind of challenge should come next.
Key Points
- •The paper studies benchmark saturation in AI evaluation and defines the concept explicitly.
- •It analyzes 60 language model benchmarks using 14 properties related to saturation.
- •The authors report that nearly half of the examined benchmarks exhibit saturation.
- •The study finds that saturation becomes more common as benchmarks age.
- •The authors conclude that expert curation improves resilience to saturation, while public test data does not explain that resilience.