July 22, 2026
Birds, bikes, and benchmark beef
Are AI Labs Pelicanmaxxing?
AI’s pelican test sparked a comment war over cheating, style, and bird-bike bragging rights
TLDR: A test of more than 1,000 AI drawings found no clear proof that major labs are secretly optimizing for the famous pelican-on-a-bicycle challenge. Commenters still turned it into drama, arguing over stealth cheating, better niche tools, and whether even the scoring method can be trusted.
A goofy internet favorite just turned into serious comment-section theater. Writer Dylan Castillo set out to answer a deliciously petty question: are big artificial intelligence companies secretly training their systems to ace the famous “draw a pelican riding a bicycle” test made popular by Simon Willison? After churning through 1,008 drawings across seven major models, the verdict was basically: no obvious pelican conspiracy here. The famous bird-on-bike pictures didn’t seem clearly better than the other animals and vehicles.
But the real fireworks were in the reactions. One camp was fascinated that each model seemed to have its own signature vibe, with one commenter marveling that every system kept a recognizable style across wildly different prompts. Another camp immediately went into suspicion mode: if a lab were gaming the test, they wouldn’t do it in the most obvious way—they’d quietly train on lots of similar animal-and-vehicle combos so the cheating would look like “general skill.” In other words, the comments turned into a mini detective show.
Then came the snark. Someone basically wanted to downvote the word “pelicanmaxxing” itself, which honestly may be the most internet reaction possible. And of course, a flex arrived: one commenter claimed a specialized image model makes way better-looking pelicans anyway, a classic “your benchmark king is mid” drive-by. Even the scoring system got dragged, with calls for head-to-head matchups like a sports ranking. The community wasn’t just reacting to the bird—they were judging the judges.
Key Points
- •The article investigates whether AI labs may be optimizing specifically for Simon Willison’s “pelican riding a bicycle” SVG benchmark prompt.
- •Dylan Castillo created a 48-prompt grid from 8 animals and 6 vehicles, generating 1,008 SVGs by sampling seven frontier models three times per prompt.
- •The tested models were accessed through OpenRouter and included GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro.
- •The evaluation pipeline used SVG-to-PNG rendering, GPT-5.6 Luna as a judge for 1–5 scoring across three criteria, and Gemini 3.1 Flash-Lite for feature extraction; there were 11 retries in total.
- •Before statistical analysis, the author reports that manual inspection did not show pelican-bicycle images looking noticeably better than the other prompt combinations.