opus 4.8 is still very much blind - EyeBench-V3 visual benchmark (similar to IBench)

Reddit r/singularity News

Summary

EyeBench-V3 visual benchmark evaluates Claude Opus 4.8, finding it still fails basic vision tasks, similar to IBench. The benchmark is introduced via a Twitter thread by Adonis Singh.

https://preview.redd.it/22texjo58l4h1.png?width=3340&format=png&auto=webp&s=73039f304a4ee253ca214b3378cc14a83909fc62 [https://x.com/adonis\_singh/status/2060133072482324521](https://x.com/adonis_singh/status/2060133072482324521) [https://x.com/search?q=eyebench-v3%20(from%3Aadonis\_singh)&f=top&src=typed\_query](https://x.com/search?q=eyebench-v3%20(from%3Aadonis_singh)&f=top&src=typed_query) [https://x.com/adonis\_singh/status/2031516746570469837](https://x.com/adonis_singh/status/2031516746570469837) \- benchmark introduction post
Original Article

Similar Articles

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

arXiv cs.AI

The paper introduces GameXpert-Bench, a benchmark with three tracks to evaluate coding agents' game development capabilities across creation, bug repair, and optimization. It finds that current agents are better at initial generation than at defect discovery and multi-turn optimization.

Measuring Activation Control in Large Language Models

arXiv cs.AI

This paper introduces the Activation Controllability Benchmark to measure how well large language models can modulate their residual stream via natural-language instructions, finding that most models can do so to some extent, which could evade activation-based monitoring methods and pose risks for AI safety.