GPT-5.5 with tools now surpasses the 10-year-old level on the BabyVision benchmark

Reddit r/singularity Models

Summary

GPT-5.5, when equipped with tools, has surpassed the performance of a 10-year-old child on the BabyVision benchmark, marking a notable advancement in AI visual reasoning.

No content available
Original Article

Similar Articles

GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]

Reddit r/MachineLearning

A new arXiv paper introduces the ActiveVision benchmark designed to test repeated visual perception, finding that frontier vision models like GPT-5.5 and Claude Fable 5 score only 10.6% and 3.5% respectively, while humans achieve 96.1%.

Introducing GPT-5.2

OpenAI Blog

OpenAI introduces GPT-5.2, the most capable model series yet, with significant improvements in knowledge work, code generation, image perception, long-context understanding, and tool-calling. The GPT-5.2 Thinking variant achieves state-of-the-art performance on professional benchmarks, outperforming human experts on 70.9% of GDPval tasks across 44 occupations.

Introducing GPT-5

OpenAI Blog

OpenAI introduces GPT-5, a significant leap in AI intelligence featuring state-of-the-art performance across coding, math, writing, health, and visual perception. The unified system includes a smart efficient model, a deeper reasoning model (GPT-5 thinking), and a real-time router for optimal response selection.

Introducing GPT-5.1 for developers

OpenAI Blog

OpenAI releases GPT-5.1, a new model in the GPT-5 series that dynamically adapts thinking time based on task complexity, offering 2-3x faster performance than GPT-5 while maintaining frontier intelligence. The release includes extended prompt caching (24-hour retention), new coding tools (apply_patch and shell), and a 'no reasoning' mode for latency-sensitive applications.