anagnorisis.cloudSign in

← Hourlies

Hourly ·

No, AI Labs Aren't Pelicanmaxxing

A systematic experiment across seven frontier models finds no evidence of "pelicanmaxxing" — labs aren't secretly training on Simon Willison's famous benchmark.

No, AI Labs Aren't Pelicanmaxxing

For years, Simon Willison has tested every major LLM release with the same prompt: "Generate an SVG of a pelican riding a bicycle." What started as a joke became one of AI's most famous informal benchmarks — and sparked a natural question: are labs secretly optimizing for it?

Dylan Castillo put that question to the test. He generated 1,008 SVGs across seven frontier models — GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro — using 48 animal-vehicle combinations and had GPT-5.6 Luna score them all.

The verdict: no pelicanmaxxing detected. Pelicans aren't drawn any better than other animals. Bicycles aren't drawn any better than other vehicles. And no lab draws the famous combination better than its component scores would predict.

GLM-5.2 came closest with a slight boost on the pelican-bicycle cell and an eye-catching first sample, but the effect was too small to be significant. One oddity: all 21 pelican-bicycle images across every model faced right — the only combination with unanimous direction. But with 48 combinations, one reaching 21 out of 21 isn't that surprising.

The more plausible story, Castillo notes, is "SVGmaxxing" — labs quietly improving SVG generation in general, rather than gaming one specific prompt. The full dataset and pipeline are open-source on GitHub.

Sources: Dylan Castillo / Iwana Labs, Hacker News discussion

More Hourlies Stories

Content on Anagnorisis is summarized, paraphrased, and editorialized from publicly available sources for length and clarity. Original sources are linked where available. All trademarks belong to their respective owners.

More from Anagnorisis