July 28, 2026

Fake it till you make a score

Don't ask an LLM for a confidence score

Experts say AI confidence scores are fake comfort — and commenters brought knives out

TLDR: The article says chatbot “confidence scores” are mostly meaningless numbers that make answers feel safer than they really are. Commenters split between agreeing it’s fake reassurance, citing research that might salvage the idea, and roasting everything from the evidence to the site’s unreadable design.

A tech essay arguing that you should stop asking chatbots how sure they are lit up the comment section, and honestly, the replies are half the show. The author’s big claim is simple: when an artificial intelligence tool gives you a neat little confidence number — say 87 out of 100 — it may look reassuring, but it doesn’t actually prove the answer is right. In plain English, the number can be more like a security blanket than a safety check.

That take got a lot of nods from readers who said the scores only mean anything in very controlled situations. One commenter basically said, yes, if you compare different chatbots or even different prompts, those numbers can become nonsense fast. But of course, this is the internet, so agreement lasted about five seconds before the pushback arrived. Another reader came in with a counterexample from an earlier Hacker News thread, pointing to research claiming a separate tool could predict when a model was likely to be wrong and even switch to a smarter model.

And then the mood turned spicy. One commenter mocked the evidence as so shaky that the supposed ability might be imaginary. Another snapped, “Stop with the slop,” which is about as close to a digital eye-roll as you can get. In the funniest side quest of the thread, someone ignored the AI debate entirely to complain that the article’s text color was unreadable and they had to turn on reader mode. So yes: the confidence score debate is raging, but the real community consensus may be that bad web design inspires the most confident opinions of all.

Key Points

  • The article argues that LLM-generated self-confidence scores are not scientifically valid indicators of output correctness.
  • It says the pattern appears in chat systems, structured outputs, and agentic workflows that return both a response and a confidence field.
  • The author specifically criticizes continuous confidence scales such as 0 to 100 as lacking a sound basis.
  • The article cites Anthropic research on model planning and introspection, while noting the researchers’ caveat that these abilities are unreliable and context-dependent.
  • It concludes that reflective reasoning may improve performance or error detection, but does not establish a reliable method for quantifying model confidence.

Hottest takes

"Stop with the slop." — empthought
"the 'capability' just another imagining" — chrisjj
"I had to turn on reader mode in Firefox." — SubiculumCode
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.