opus 4.8 is still very much blind - EyeBench-V3 visual benchmark (similar to IBench)
Summary
EyeBench-V3 visual benchmark evaluates Claude Opus 4.8, finding it still fails basic vision tasks, similar to IBench. The benchmark is introduced via a Twitter thread by Adonis Singh.
Similar Articles
OpenAI's new chip is better than Vera rubin on benchmark
OpenAI has unveiled a new chip that surpasses the Vera Rubin benchmark, indicating progress in AI hardware development.
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?
The paper introduces GameXpert-Bench, a benchmark with three tracks to evaluate coding agents' game development capabilities across creation, bug repair, and optimization. It finds that current agents are better at initial generation than at defect discovery and multi-turn optimization.
Measuring Activation Control in Large Language Models
This paper introduces the Activation Controllability Benchmark to measure how well large language models can modulate their residual stream via natural-language instructions, finding that most models can do so to some extent, which could evade activation-based monitoring methods and pose risks for AI safety.
K-Bench: measuring model performance on real scientific agent requests
The paper introduces K-Bench 01, a benchmark for evaluating AI agents on real scientific requests, revealing that no model consistently meets the threshold for acceptable performance, with overclaiming as a common failure.
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
The paper introduces GUI-Primitives, a benchmark of 994 contrastive instruction pairs to diagnose spatial reasoning failures in vision-language models for GUI grounding, revealing that most failures stem from candidate localization rather than relation understanding.