The fragile foundations of CoT monitoring

AI’s “show your work” safety plan has people arguing it was shaky all along

TLDR: Researchers are warning that a popular way to inspect AI — making it write out its reasoning — was never designed as a safety feature and may not last. Commenters are split between “any visibility helps” and “this is just the bot telling us a neat story,” which matters because a lot of trust may be riding on it.

The big mood around this piece? Deeply uneasy, slightly exhausted, and very ready to fight in the comments. The article says modern AI safety has been leaning on a lucky idea: if a chatbot writes out its “thinking” step by step, humans might be able to spot danger before it gives a bad answer. But researchers at a recent workshop reportedly came away warning that this setup is fragile. Why? Because that step-by-step output was built to make AI better at hard tasks, not to make it honest. And commenters absolutely pounced on that gap.

The strongest reactions split into two camps. One side basically said, “So the safety plan is to trust the machine’s diary?” They mocked the idea that a system could be forced to reveal all its real reasoning in words, with jokes comparing it to asking a student to “show your work” after they already copied the answer. The other side pushed back that even an imperfect peek is better than total darkness, arguing that losing these visible thought traces would make safety work much harder. That sparked the main drama: capability versus transparency — in plain English, making AI smarter versus making it easier to inspect.

And yes, the humor was brutal. People joked that AI companies will choose “winning” over “explaining” every time, and memed the whole debate as a magician yelling, “Ignore the hidden computation behind the curtain!” The vibe was part serious warning, part communal eye-roll, with many readers landing on the same bleak takeaway: if this safety tool only exists by accident, betting the future on it feels wildly risky.

Key Points

  • The article reports on a workshop focused on chain-of-thought monitorability and its relationship to AI safety.
  • Chain-of-thought reasoning is described as originally developed to improve model performance on hard tasks, not for monitoring.
  • The article cites Wei et al. 2022 as the foundational CoT prompting paper and says OpenAI’s o1 model extended CoT into post-training practice in September 2024.
  • A 2025 paper by Korbak, Balesni, et al. is presented as framing CoT monitorability as a 'fragile opportunity' for AI safety.
  • The article says some recent papers suggest transparency can reduce CoT effectiveness, motivating a call to reduce safety dependence on CoT monitoring.

Hottest takes

"So the plan is: ask the liar to narrate the lie" — @throwawayalignment
"‘Show your work’ for a machine that already got the answer is security theater" — @latent_larry
"We accidentally found one tiny window into the model, and now people want to build the whole safety house around it" — @safety_skeptic
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.