July 30, 2026
Leaderboard glow-down
Show HN: I audited my AI leaderboard scale – every score dropped 6-15 points
AI scorekeeper slashed every chatbot’s grade — and commenters instantly smelled chaos
TLDR: AGI Ranker, a site tracking how close AI is to human-level smarts, cut every model’s score by 6–15 points after admitting most of its scale wasn’t truly based on human comparisons. Commenters loved the honesty but mocked the idea that one neat number could still claim “infinite clarity.”
The biggest twist in this Show HN post isn’t just that every major AI model suddenly lost 6 to 15 points on the AGI Ranker scoreboard. It’s that the site’s own creator did the damage himself, publicly, after admitting the old scale made things look more “human-level” than the data really supported. In plain English: the leaderboard that tries to say how close AI is to true human-like intelligence got a lot harsher overnight, and the internet immediately pounced.
The creator says the rankings still keep the same overall order, but the shiny numbers are lower because only one of the ten tests actually has a real measured human comparison. The rest were being judged against the test’s max score, not actual people. That correction won praise from transparency fans — but it also opened the floodgates for sarcasm. The loudest reaction came from a commenter roasting the site’s slogan, “ONE SCORE. INFINITE CLARITY.” before deadpanning that they’d “never had less clarity.” Ouch.
That joke basically became the mood of the thread: part respect, part side-eye. Some readers seemed impressed that someone would publicly downgrade their own project instead of quietly moving on. Others seized on the deeper drama: if the headline number can drop this much after a methodology cleanup, how much trust should anyone put in AGI countdowns at all? In other words, the benchmarks may be scientific, but the comment section turned it into a reality show about credibility, hype, and whether one magic number can ever explain a mess this big.
Key Points
- •AGI Ranker aggregates 10 public AI benchmarks into a single 0-100 AGI Score and says each score is traceable to cited sources.
- •The system weights benchmark results by source quality, giving full weight to independent leaderboards, reduced weight to partially involved third parties and self-reports, and excluding non-verified sources.
- •In v2.0.0, raw benchmark scores were rescaled by ceilings based on either measured human parity or benchmark maximums, correcting prior descriptions that treated most ceilings as best-human.
- •Among currently scored benchmarks, only GPQA Diamond uses a measured human ceiling; the other nine are normalized to benchmark maximums, making the score a mixed scale.
- •The AA Intelligence Index was removed from the AGI Score because its newer version bundled benchmarks already scored directly, which the article says would have double-counted signals.