GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

AI physics showdown sparks cheers, side-eyes, and "where are the other bots?"

TLDR: JuliaHub tested top AI models on real-world-style physics problems and says hidden-answer grading matters more than clean-looking code. Commenters immediately turned it into a broader fight over missing rivals, possible bias, and whether the whole contest was outdated on arrival.

A fresh AI face-off tried to answer a deceptively simple question: which chatbot is less likely to fake physics and more likely to match reality? The test came from JuliaHub, which ran OpenAI’s GPT-5.6 family against Anthropic’s Claude Fable 5 on five simulation tasks based on real engineering-style problems. The company says it locked almost everything down and judged results against hidden real-world answers, not just whether the code ran without errors. In plain English: this wasn’t just “does it work,” but “does it make sense in the real world?”

But in the comments, the lab-coat seriousness instantly turned into a mini popcorn fight. One camp gave a polite golf clap — “Nice!” — then immediately asked why Codex wasn’t invited. Another basically said, hold on, where’s Google? because if anyone should crush physics, surely it’s the company people associate with giant brains and multimodal wizardry. Then came the sharpest side-eye of all: “Yet another benchmark to promote their own harness” — aka the classic internet suspicion that every scoreboard is secretly an ad wearing glasses.

And because no tech thread is complete without someone moving the goalposts into the future, one commenter jumped to the bigger sci-fi question: will “world models” eventually leave today’s chatbots in the dust for real-world tasks? Meanwhile, others roasted the whole thing as already aging badly because newer rivals like Kimi 3 and Opus 5 weren’t included. The vibe was less “case closed” and more “drop the full cage match, coward.”

Key Points

  • The article argues that physical-AI evaluation must measure physical correctness against real-world ground truth, not just whether code compiles or passes tests.
  • JuliaHub says it held the Dyad agent harness, problem set, reasoning effort, context window, and token budget constant so that only the model varied.
  • The compared models were OpenAI’s GPT 5.6 variants (terra, sol, luna) and Anthropic’s claude-fable-5.
  • The study used five sealed modeling-and-simulation problems, with three trials per model on four problems and one long-horizon trial on a fifth, totaling 52 graded runs.
  • Scores were based on simulated trajectories against sealed ground truth, then aggregated with difficulty weighting; the article also analyzed transcript tool calls to derive work-style fingerprints.

Hottest takes

"missing Codex" — jespinel
"Yet another benchmark to promote their own harness" — gizmodo59
"already out of data" — hartator
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.