July 22, 2026
Dungeon crawl, but make it drama
Can a MUD evaluate LLMs? A $99 proof of concept
AI sent into an old-school dungeon, and the comments instantly turned into a nostalgia riot
TLDR: Researchers tested 13 AI models in an old-school text world for just under $100 and found that changing one judge in the scoring system could seriously reshuffle the winners. Commenters were split between loving the retro MUD idea, mourning lost internet culture, and jokingly demanding even weirder AI dungeon experiments.
A tiny $99 experiment just dropped a big question into the AI world: can you test chatbots better by throwing them into a vintage text adventure instead of a flashy modern simulator? In CrucibleBench, the bots wander a shared fantasy world, talk to guards and townsfolk, build trust, and live with their mistakes for 50 turns. One standout example had a model find a signet ring, politely return it, and earn a recommendation to join the town watch. Cute? Yes. But the real tea is that the researchers say a single judging component reshuffled the rankings so much that the leaderboard changed by up to six spots. In other words: even the scorekeeper may be stirring the pot.
And the community absolutely ran with it. One camp was instantly hit by full nostalgia damage, with commenters basically saying, “Forget the benchmark, I miss MUDs,” referring to old text-based online worlds. Varelion turned the thread into a mini elegy for pre-Discord roleplay culture, mourning how cozy text communities got swallowed by today’s clunky chat apps. Others were delighted by the whole vibe, calling it a fun throwback and a clever learning tool for building AI agents. Then came the chaos gremlins: someone immediately asked whether anybody had plugged an AI into the famously absurd Hitchhiker’s Guide text game, which is exactly the kind of chaotic science the internet wants. The hottest undercurrent? A very online suspicion that benchmarking AI is messy, the judges may be biased, and everyone is one weird text dungeon away from a new argument about what “smart” even means
Key Points
- •CrucibleBench evaluates language models in a persistent MUD with 50-turn runs, hidden social objectives, and replayable transcripts.
- •The Phase 1 proof-of-concept release covers 13 models, 650 runs, and $99.59 in billed cost.
- •The benchmark environment includes 7 command types, 12 rooms, 14 items, and 4 NPCs with trust and suspicion state.
- •A sample run shows GPT-5.4 completing a trust-based objective by returning a signet ring and earning a recommendation in 14 turns.
- •The article's main finding is that an LLM-judge component reordered leaderboard rankings by up to six positions, while aggregate reliability statistics did not expose the instability.