Back to List
Investigating AI Model Performance: Are Frontier Labs Optimizing for the Famous Pelican Benchmark?
Industry NewsArtificial IntelligenceLLM BenchmarksSVG Generation

Investigating AI Model Performance: Are Frontier Labs Optimizing for the Famous Pelican Benchmark?

Researcher Dylan Castillo has launched an investigation into whether major AI labs are 'pelicanmaxxing'—specifically optimizing their large language models (LLMs) to excel at the famous 'pelican riding a bicycle' SVG generation prompt. Originally popularized by Simon Willison, this informal benchmark has become a staple of AI model releases on platforms like Hacker News. Castillo’s experiment involved generating 1,008 SVGs across seven frontier models, including GPT-5.6 Terra and Claude Sonnet 5, using a grid of 48 prompt variations. By testing different animals and vehicles, such as flamingos on scooters or whales on planes, the study aims to determine if models perform disproportionately better on the original pelican prompt compared to similar tasks. The analysis, conducted using an LLM judge and Claude Fable 5, explores the integrity of informal benchmarks in an industry with billions of dollars at stake.

Hacker News

Key Takeaways

  • The Pelican Benchmark Origin: Simon Willison’s prompt, "Generate an SVG of a pelican riding a bicycle," has evolved from a tongue-in-cheek test into a globally recognized informal benchmark for LLM releases.
  • The 'Pelicanmaxxing' Hypothesis: There is growing concern that AI labs may be "benchmaxxing"—tuning models specifically to perform well on famous prompts to influence user perception and market value.
  • Large-Scale Testing: Dylan Castillo conducted an experiment generating 1,008 SVGs to test this theory, utilizing seven frontier models through OpenRouter.
  • Methodological Depth: The test used a 48-prompt grid (8 animals × 6 vehicles) to compare performance on the original pelican prompt against variations of varying difficulty and similarity.
  • Advanced Analysis: The results were evaluated using an LLM judge and analyzed with Claude Fable 5 to identify patterns of specific optimization.

In-Depth Analysis

The Rise of Informal Benchmarking in AI

For several years, the AI community has looked toward Simon Willison’s "pelican riding a bicycle" prompt as a quick litmus test for the spatial reasoning and code-generation capabilities of new models. What started as a simple, humorous request has gained significant traction, often appearing as a top-voted comment in Hacker News threads whenever a new model is announced. Because this prompt is so widely known, it has raised questions about the validity of model performance. In an industry where billions or even trillions of dollars in valuation can hinge on public perception of model "intelligence," the temptation for labs to specifically optimize for these viral benchmarks—a practice referred to as "pelicanmaxxing"—is substantial.

Experimental Design and Methodology

To investigate whether models are being unfairly tuned for the pelican prompt, Dylan Castillo designed a controlled experiment. He moved beyond the single prompt to create a grid of 48 distinct combinations, featuring eight different animals and six different vehicles. The animals selected included the pelican, flamingo, heron, otter, raccoon, antelope, whale, and cat. The vehicles included the bicycle, unicycle, skateboard, scooter, plane, and boat.

This selection was intentional; animals like the flamingo and heron were chosen for their physical similarity to the pelican, while others like the cat or raccoon were considered "easy" cases. The antelope was categorized as a "hard" case, and the whale was chosen as a radical departure from the original subject. By maintaining nearly identical phrasing to Willison’s original prompt while swapping the subjects, Castillo created a baseline to see if a model that can draw a perfect pelican on a bike fails significantly when asked to draw a flamingo on a scooter.

Model Selection and Execution

The experiment utilized seven of the most prominent frontier models available via OpenRouter: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro. To ensure consistency, Castillo requested the same level of reasoning effort from every model and set the temperature to 1.0. Each of the 48 prompts was run three times per model, resulting in a dataset of 1,008 generated SVGs. This volume of data allows for a statistical look at whether the "pelican" and "bicycle" tokens trigger higher quality outputs than their counterparts, which would suggest targeted optimization by the developers.

Industry Impact

The implications of "pelicanmaxxing" extend far beyond a simple SVG of a bird. If AI labs are indeed optimizing for specific, famous prompts, it suggests a shift from general-purpose intelligence toward "benchmark gaming." This practice can mislead researchers, investors, and end-users about a model's true capabilities. As informal benchmarks continue to carry weight in the community, the need for more robust, varied, and randomized testing—like the animal-vehicle grid used in this study—becomes critical. It highlights a growing tension in the AI industry: the balance between genuine architectural improvements and the marketing necessity of performing well on the internet's favorite tests.

Frequently Asked Questions

Question: What exactly is "pelicanmaxxing" in the context of AI?

"Pelicanmaxxing" refers to the hypothetical practice where AI labs specifically fine-tune or optimize their models to perform exceptionally well on the "pelican riding a bicycle" SVG prompt. This is done because the prompt is a famous informal benchmark, and a good result can generate positive PR and user trust.

Question: Which models were tested in Dylan Castillo's experiment?

The experiment tested seven frontier models: GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2, and DeepSeek V4 Pro.

Question: How were the 1,008 generated SVGs evaluated?

The SVGs were processed through a three-stage evaluation system involving an LLM judge to score the images, with the final analysis of the data performed by Claude Fable 5.

Related News

OpenAI Halts Specific Astra Model Development Phases Citing Critical Cybersecurity Prowess Concerns
Industry News

OpenAI Halts Specific Astra Model Development Phases Citing Critical Cybersecurity Prowess Concerns

OpenAI has officially announced a strategic slowdown in the development of its upcoming AI model, Astra. This decision involves the suspension of work on specific aspects of the model, primarily driven by internal concerns regarding its cybersecurity prowess. The move highlights a cautious approach by OpenAI as it navigates the complexities of developing advanced artificial intelligence that may possess dual-use capabilities. By pausing these specific development tracks, the company is prioritizing the mitigation of potential security risks over the speed of deployment. This development marks a significant moment for the Astra project, reflecting the rigorous safety and security evaluations that upcoming models must undergo before further progression or public release.

Fenix Flexin Admits to Using AI for 'Rubberz' Following Exposure by Producer Medasin and Treblo
Industry News

Fenix Flexin Admits to Using AI for 'Rubberz' Following Exposure by Producer Medasin and Treblo

LA rapper Fenix Flexin has officially acknowledged the use of artificial intelligence in the production of his 80s synth-pop-themed track, 'Rubberz.' This admission comes after a period of speculation and public claims made by producer Medasin, who utilized social media to demonstrate that the song was created using an AI tool called Treblo (formerly known as Sonauto). The situation reached a turning point when Treblo released its own AI detection software, which specifically identified 'Rubberz' as a product of its platform. This case marks a significant moment in the music industry, highlighting the increasing transparency—or lack thereof—surrounding AI-generated content and the emerging role of detection technology in verifying artistic authenticity.

Roku Launches Experimental AI-Generated FAST Channel Shifting Focus from Traditional Classic Content to Constant AI Streams
Industry News

Roku Launches Experimental AI-Generated FAST Channel Shifting Focus from Traditional Classic Content to Constant AI Streams

Roku has introduced a new experiment within the Free Ad-supported Streaming Television (FAST) sector, moving away from the traditional model of rediscovering classic films and series. This new initiative, titled "Fairground," focuses on providing viewers with a continuous stream of AI-generated content. Unlike conventional FAST channels that curate professionally produced entertainment, Roku's latest venture represents a pivot toward automated media consumption. The move has sparked discussions regarding the quality and nature of such content, with early critiques comparing the viewing experience to "eating from a trough." This comparison highlights a potential shift in how streaming platforms approach content volume versus traditional production values, signaling a significant experimental phase for Roku as it explores the intersection of artificial intelligence and ad-supported streaming media.