July 29, 2026

Peer review? More like peer roast

The Scientific Literature Is Poisonous to LLMs

Scientists say AI is eating bad research, and the comments instantly turned into a trust-no-one brawl

TLDR: A new essay argues that feeding modern scientific papers into AI can actually make it worse, and points to research suggesting some academic text hurts results. Commenters were split between "obviously, humans are messy" and "hold on, this article sounds shaky too," with others joking that AI may be poisoning science right back.

A spicy essay from Reinvent Science basically dropped a grenade into the "AI will save science" fantasy: maybe the real problem is that modern research papers are such a messy mix of truth, hype, mistakes, and fraud that feeding them to chatbots makes the bots worse. The article points to a 2024 study saying that removing some big piles of academic writing from training data actually improved performance in some cases. Translation for normal people: giving an artificial intelligence more research papers did not always make it smarter.

But the real fireworks were in the comments, where readers split into camps almost immediately. One crowd said, well, duh: humans produce flawed information, so of course machines trained on it inherit the chaos. Another camp was much harsher, rolling their eyes at the article itself and basically asking why anyone should trust a brand-new Substack making such a huge claim with thin receipts. And then came the existential dread faction: if scientific literature is "poisonous" to AI, as one commenter put it, "Then is poisonous to us too." Oof.

The funniest twist? Several readers argued this is not some science-only scandal at all. News, game guides, internet posts, everything online is contradictory and half-labeled. One zinger flipped the whole story on its head: "Not as much as LLMs are poisonous to the scientific literature." So the thread became less "AI in science" and more who is corrupting whom first?

Key Points

  • The article argues that modern scientific literature contains a mix of accurate, inaccurate, honest, and dishonest material that makes it difficult to use as reliable LLM training data.
  • It cites a 2024 study that held model architecture constant while removing training corpora to measure their effects on LLM performance.
  • According to the article, removing ArXiv, PhilPapers, and NIH ExPORTER improved academic-question performance, overall benchmark performance, and reduced toxic output.
  • The article notes that removing PubMed reduced model performance somewhat, which it presents as a complicating result rather than a definitive conclusion.
  • It argues that AI-for-science systems lack access to the informal reputational networks human scientists use to judge reliability, leading to proposals for capturing more social context from researchers.

Hottest takes

"Trust the science!" — pandinus
"Then is poisonous to us too." — sebastianconcpt
"Not as much as LLMs are poisonous to the scientific literature" — iLoveOncall
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.