@reach_vb: hey @simonw - given pelican on a bike svg is now saturated (?), what's your go-to vibe check for a model/ harness these…
Summary
A tweet from @reach_vb asking @simonw about methods to evaluate AI models or harnesses, referencing a saturated 'pelican on a bike svg' benchmark.
Similar Articles
@karpathy: More on the pelican on the bicycle test from @simonw: https://simonwillison.net/2025/Jun/6/six-months-in-llms/… I uploa…
Simon Willison's keynote at AI Engineer World's Fair reviews the last six months in LLMs, highlighting over 30 significant model releases and his 'pelican on a bicycle' SVG benchmark as a practical evaluation tool.
Are AI Labs Pelicanmaxxing?
An analysis investigates whether AI labs are optimizing their models for the popular 'pelican riding a bicycle' SVG benchmark, testing seven frontier models across 48 prompts with varied animals and vehicles, finding no strong evidence of overfitting.
scosman/pelicans_riding_bicycles
Simon Willison's link post highlights a dataset or project titled 'pelicans_riding_bicycles', likely used for LLM training or generative AI experimentation.
@omarsar0: Important discussion. Measuring models against harnesses is completely broken. I prefer to test model quality against m…
The author criticizes the broken system of measuring AI models against test harnesses due to biases and lack of standardization, while speculating that models like Claude may soon generate harnesses dynamically.
Are AI labs pelicanmaxxing?
Dylan Castillo conducted a rigorous investigation to determine if AI labs have been secretly training models to draw pelicans riding bicycles. Testing multiple models with various animal-vehicle combinations, he found no evidence of 'pelicanmaxxing'.