July 28, 2026
Lie detector or logic meltdown?
Truth is not a direction: a Tarski attack on LLM probes
Researchers chased an AI truth detector, and the comments said “not so fast”
TLDR: The article says a perfect internal AI truth detector probably can’t exist because self-referential logic breaks it. Commenters mostly fired back that this is attacking a straw man, since the real goal is reading what the AI believes, not building a magic truth oracle.
A fresh AI theory post tried to throw cold water on one of the buzziest dreams in machine learning: the idea that a chatbot might have a neat internal “truth meter” hiding somewhere in its brain. The article argues that if a language model can talk about statements, meanings, and even itself, then trying to build a perfect built-in truth detector runs into the same kind of classic logic traps that wrecked earlier dreams of a universal truth machine. In plain English: the author says truth may not be a simple arrow you can point to inside the model.
But in the comments, the crowd basically yelled: “That’s not what the probe was for!” Several readers pushed back hard, saying the whole article goes a bit too dramatic by attacking a claim few people actually make. The hottest rebuttal was that these so-called truth probes are really trying to measure what the AI believes is true, not some godlike final answer about reality. As one commenter put it, even a probe with perfect accuracy could still be measuring a belief that is wrong or inconsistent.
The thread also served up some classic internet seasoning. One person joked that someone at MIRI probably knows the answer but won’t spill it, which is exactly the kind of suspicious-lab humor AI threads love. Another commenter brought up a Zach Weinersmith comic from 11 years ago, essentially saying, “Congrats, satire predicted this mess.” And then came the philosophical side quest: could we just invent a third bucket called paradox and stop the whole truth-vs-lie cage match? The result was less “AI breakthrough” and more comment-section courtroom drama over what anyone was claiming in the first place.
Key Points
- •The article examines claims that LLM embedding spaces may contain a linear direction corresponding to truth.
- •It says researchers have trained classifiers on embeddings of true and false statements, and that these probes appear to work well and generalize to some extent.
- •The article compares this effort to historical attempts to formalize mathematical truth, especially the work associated with Russell, Whitehead, and Hilbert.
- •It uses Gödel’s incompleteness result, Tarski’s theorem on truth predicates, and Turing’s halting problem as examples of limits created by self-reference.
- •The article argues that because transformers encode language in a system that can also express concepts in language, no embedding-space probe can fully capture truth.