July 18, 2026

Goal mode? More like drama mode

Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?

AI ‘try harder’ button flops as fans argue, nitpick charts, and crown a new beast

TLDR: Claude Fable 5 came out looking strongest in this difficult head-to-head, while the special `/goal` mode proved unreliable rather than magical. Commenters split between hype, skepticism, and comedy, with debates over whether the test means anything at all—and jokes that broken CSS may be the real benchmark.

A fresh AI showdown tried to answer a very online question: if you flip on a model’s special “goal mode”—basically its built-in “focus harder” setting—does it suddenly become a genius? According to this benchmark, not really. The big winner was Claude Fable 5, which the author basically hailed as a monster performer after it beat GPT-5.6 Sol on a brutally difficult network-design puzzle with a mind-bending number of possible answers. The real tea, though, is that the flashy /goal feature wasn’t a magic button. Sometimes it helped, sometimes it absolutely did not, and in one run it seems to have just given a bad idea extra time to ruin everyone’s day.

The comments instantly turned this into a mini tech soap opera. One camp was impressed that someone even did a deep dive on /goal, with one user calling it a perfect test for the feature. Another camp came in swinging with skepticism, saying the results look like noise because the problem is so huge and messy that a handful of runs proves very little. Then there were the lovable chaos goblins: one commenter got distracted by the chart design and complained that “lower is better” while the graph visually looked upside down, while another delivered the line of the thread by saying they’re reading this serious benchmark on one screen and cleaning up catastrophic CSS made by Sol on the other. In other words: one community is debating AI intelligence, another is debating graph readability, and everyone agrees the comments are more fun than the math.

Key Points

  • The article benchmarks six AI models on KIRO, an unpublished NP-hard fiber-network optimization problem, in plain mode and native /goal mode.
  • KIRO requires building valid redundant loop-and-branch networks over directed distance matrices for Grenoble, Nice, and Paris while minimizing total cable length.
  • Azam estimates the search space is enormous, including a Paris lower bound of `11^532` for simplified assignments and about `10^1223` for one restricted valid-solution family.
  • The experiment used a 30-minute optimization budget, 1,900-second outer timeout, maximum reasoning settings, and execution via Harbor 0.1.43 and Docker.
  • Fable 5 achieved the best reported result overall, and the article concludes that /goal changes search behavior rather than serving as a universally helpful 'try harder' mode.

Hottest takes

"A deepdive on the /goal effect on a problem literally made for this." — couAUIA
"I love that we have this on one hand and me cleaning up catastrophic CSS made by Sol on the other." — techpression
"Results seem mostly noise to me." — andai
Made with <3 by @siedrix and @shesho from CDMX. Powered by Forge&Hive.