GPT-5.5 with tools now surpasses the 10-year-old level on the BabyVision benchmark
Summary
GPT-5.5, when equipped with tools, has surpassed the performance of a 10-year-old child on the BabyVision benchmark, marking a notable advancement in AI visual reasoning.
Similar Articles
GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]
A new arXiv paper introduces the ActiveVision benchmark designed to test repeated visual perception, finding that frontier vision models like GPT-5.5 and Claude Fable 5 score only 10.6% and 3.5% respectively, while humans achieve 96.1%.
Introducing GPT-5.2
OpenAI introduces GPT-5.2, the most capable model series yet, with significant improvements in knowledge work, code generation, image perception, long-context understanding, and tool-calling. The GPT-5.2 Thinking variant achieves state-of-the-art performance on professional benchmarks, outperforming human experts on 70.9% of GDPval tasks across 44 occupations.
Introducing GPT-5
OpenAI introduces GPT-5, a significant leap in AI intelligence featuring state-of-the-art performance across coding, math, writing, health, and visual perception. The unified system includes a smart efficient model, a deeper reasoning model (GPT-5 thinking), and a real-time router for optimal response selection.
GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B and most today’s low-tier models
GPT-5, which was the best model just a year ago, is now outperformed by Qwen3.6 27B and other current low-tier models, highlighting the rapid pace of AI advancement.
Introducing GPT-5.1 for developers
OpenAI releases GPT-5.1, a new model in the GPT-5 series that dynamically adapts thinking time based on task complexity, offering 2-3x faster performance than GPT-5 while maintaining frontier intelligence. The release includes extended prompt caching (24-hour retention), new coding tools (apply_patch and shell), and a 'no reasoning' mode for latency-sensitive applications.