July 21, 2026
Mona Lisa meets comment-section mayhem
"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok
AI art showdown turns messy as commenters roast the test, the hype, and especially Grok
TLDR: A drawing test had four AI systems try to sketch the Mona Lisa with digital pencils, and the results sparked more drama than awe. Commenters argued over whether the test was meaningful or just marketing, while Grok got openly laughed at for its rough performance.
Four big-name AI systems were handed a blank page, digital colored pencils, and one very intimidating assignment: recreate the Mona Lisa and other famous-style scenes stroke by stroke. The experiment tracked every move, every peek at the canvas, and even how much money each attempt cost. On paper, it was a neat test of whether these tools can slowly improve a drawing like a human sketching. In the comments, though, the real masterpiece was the chaos.
The strongest reaction was a split between curious fascination and total eye-rolling. One camp loved the spectacle of pitting models against each other and watching some clearly flop under pressure. The other camp was not buying the premise at all, with one blunt commenter declaring, "I think this is just an ad." Another dismissed the whole thing with "Useless", arguing these systems were never meant to draw this way in the first place. Ouch.
And then there was Grok, which became the thread’s designated comic relief. People weren’t just disappointed — they were bewildered. The loudest laugh came from "Grok! LOL!", followed by genuine confusion over why it looked so different from the others. Meanwhile, the funniest chaos-gremlin energy came from a commenter suggesting the models should try drawing LeBron James or Heisenberg, joking that "everyone will refuse." Even a broken signup error page got dragged into the discourse, giving the whole conversation that classic internet vibe: half product debate, half dumpster fire, fully entertaining.
Key Points
- •The article describes a drawing arena where AI models use colored-pencil-style tools on a blank canvas and can inspect their work during execution.
- •Four models—GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash—were tested across two target reproductions and five prompt-based drawing tasks, totaling 28 drawings.
- •Target-image tasks used the Mona Lisa and Van Gogh's Starry Night, and outputs were scored with structural similarity (SSIM).
- •The article says Claude Fable 5 generally took longer and cost more than other models while producing worse output on this drawing task.
- •The drawing harness used in the experiments is open source at github.com/hershalb/canvas-arena.