Tag
An analysis investigates whether AI labs are optimizing their models for the popular 'pelican riding a bicycle' SVG benchmark, testing seven frontier models across 48 prompts with varied animals and vehicles, finding no strong evidence of overfitting.