@charles_irl: hate to see my favorite benchmark was actually in the training data all along
Summary
A tweet discusses concerns about an AI benchmark being included in training data, with Simon Willison sharing a related movie trivia fact.
View Cached Full Text
Cached at: 09/20/26, 09:14 AM
hate to see my favorite benchmark was actually in the training data all along
Simon Willison (@simonw): TIL that 1988 classic Who Framed Roger Rabbit features a pelican riding a bicycle!
Similar Articles
@svpino: I don't trust benchmarks. We've all seen this movie: New model beats everyone else on a benchmark. People hype it. Then…
The tweet critiques traditional AI benchmarks and introduces TRACES, a new benchmark that evaluates AI's discovery process by focusing on how models reach answers, including tool usage, error correction, and evidence tracing.
The benchmarks the big labs don't want you to see
The article discusses undisclosed benchmarks used by major AI labs, highlighting issues with transparency in the AI industry.
@charles_irl: new benchmark just dropped
Andon Labs released a new benchmark testing whether AI models refuse to play a Nazi marching song, finding that Claude Opus 4.8 and GPT 5.5 always refused, Gemini 3.5 Flash refused half the time, and Grok 4.3 almost always played it.
@thsottiaux: Do you still trust benchmarks or do you just listen to your friends? What makes you try a new model?
A tweet questioning the trustworthiness of benchmarks and asking what drives users to try new AI models.
Does anyone else feel like AI benchmarks are becoming less useful for predicting real-world performance?
The article discusses the growing disconnect between high AI benchmark scores and actual real-world performance, highlighting issues like consistency, latency, and context handling.