July 19, 2026

Budgeting by setting money on fire

I burned all my tokens researching how to save tokens

He tried to save AI credits and accidentally proved everyone else’s token pain is real

TLDR: A Quesma researcher tried to learn how to cut AI costs and burned through his whole plan before getting an answer, then rebuilt the setup using tools he already paid for. Commenters turned it into a lively debate over whether smart routing really saves money or just creates new ways to waste it.

A developer at Quesma basically lived the most 2026 tech nightmare imaginable: he asked an artificial intelligence research tool how to save on AI usage, and the thing blew through his entire monthly allowance in 30 minutes before it could even finish the job. That instantly turned a dry cost experiment into a full-on community therapy session, with readers piling in to say, yes, this is exactly how modern AI tools eat your wallet while promising to help you budget.

The crowd reaction was half “same here, brother” and half impromptu consulting war. One commenter cheered that the article confirmed their own trick of making models talk to local files in shorter bursts. Another came in with the colder, spicier take: maybe these systems just can’t be trusted to know much about saving money in the first place, because they only know what they’ve already swallowed from the internet. Ouch.

Then the optimization nerds arrived. One argued the real secret is obvious: start with the cheap models first and only bring in the expensive brainpower at the end. Another dropped the kind of painful wisdom only experience can buy: lots of so-called token-saving hacks are actually cache killers, meaning your “money-saving” trick may secretly cost more. And of course, no comment section is complete without a little self-promo drive-by, with one builder sliding in to pitch their own tool as the answer.

The unintentional joke hanging over the whole thread? To learn how to stop burning AI credits, he first had to burn all of them. Peak tech comedy, and everyone knew it.

Key Points

  • Bartosz Kotrys of Quesma says an initial Claude deep-research run exhausted his Claude Max 5x plan in about 30 minutes before producing a final synthesis.
  • The reported run launched 111 agents and queued 123 claims for verification, but only 25 claims were verified before limits were reached.
  • The article’s stated research goal was to understand AI tokenomics, including monitoring systems, AI-spend governance, and optimization practices.
  • Kotrys says he combined Claude, Codex, and Antigravity using an extended local claude-mem shared-memory setup so tools could reuse each other’s findings.
  • The article describes a model-orchestration approach that assigns cheaper or role-specific models such as Claude Opus 4.8, Claude Sonnet 5, GPT-5.5, and Gemini 3.1 Pro to subagent tasks, informed by benchmark sources like Terminal-Bench, SWE-bench Pro, and Artificial Analysis.

Hottest takes

"more or less the same results!" — realaccfromPL
"The best way to save tokens is to start out the deep research pass with cheap models" — bob1029
"most 'token saving' ideas are cache killers" — Arkhetia
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.