July 30, 2026

Ug save money? Community say… maybe

Does Speaking to Agents Like Cavemen Save 65% of Tokens? We Test

Turns out “talk like caveman” saves a little money, but the comments went absolutely feral

TLDR: JetBrains tested “caveman mode” on real AI coding jobs and found the flashy 65% claim shrank to about 8.5% savings, with no clear drop in quality. Commenters were split between “hey, 10% still matters” and roasting the whole idea as Flintstones cosplay for software.

The big promise was juicy: make an AI coding helper speak like a cartoon caveman and supposedly slash its word-count by 65%. But when JetBrains actually put the gimmick through a real test, the result was way less prehistoric and way more mundane. In normal coding tasks, the saving was only about 8.5%, with roughly 10% lower cost at best—and even that could get wiped out by one weird, expensive run. The good news? The AI didn’t seem to get worse at its job. So yes, caveman mode is harmless fun. No, it’s not the stone-age coupon code some people hoped for.

And honestly, the real spectacle was in the comments. One camp basically shrugged: nice, but not life-changing. Another argued that 8% to 10% is still real money, especially as newer AI tools get more expensive and wordier. Then came the mockery. One commenter dragged the whole vibe by asking why “caveman dialect” is just broken English “like it’s freaking Flintstone,” while another delivered the quote of the thread: “If you just want a smaller vocabulary, use French?” Brutal. There was also a more serious hot take hiding under the jokes: maybe this whole experiment exposed that AI systems are simply about 10% inefficient in how they turn words into meaning. So the crowd’s verdict was gloriously mixed: funny gimmick, decent tiny savings, absolutely oversold—and now everyone wants to argue about whether the real problem is tokens, language, or just reading lines that sound like “gronkHitThing(true)”

Key Points

  • JetBrains tested the Caveman skill on coding-agent tasks and forced it on in every reply to measure the maximum possible token savings.
  • The benchmark found about 8.5% output-token savings, far below the 65% savings advertised for chat-style responses.
  • Across 82 paired tasks, JetBrains reported no measurable quality degradation: 8 better, 10 worse, 64 tied, with sign test p = 0.82.
  • Caveman changed the agent’s narration style as intended while leaving code artifacts, tool calls, commands, and error strings unchanged.
  • Expected per-task cost savings were around 10%, but total run cost was skewed by a single long-context outlier that made the skill arm more expensive overall.

Hottest takes

"Like wtf is the caveman dilect" — altmanaltman
"If you just want a smaller vocabulary, use French?" — abofh
"there's just a ~10% inefficiency in token to information mapping" — mmastrac
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.