opus 4.8 仍然非常盲目 - EyeBench-V3 视觉基准测试(类似于 IBench)

Reddit r/singularity 新闻

摘要

EyeBench-V3 视觉基准测试评估了 Claude Opus 4.8,发现它仍然无法完成基本视觉任务,这与 IBench 类似。该基准测试是通过 Adonis Singh 的 Twitter 帖子介绍的。

https://preview.redd.it/22texjo58l4h1.png?width=3340&format=png&auto=webp&s=73039f304a4ee253ca214b3378cc14a83909fc62 [https://x.com/adonis\_singh/status/2060133072482324521](https://x.com/adonis_singh/status/2060133072482324521) [https://x.com/search?q=eyebench-v3%20(from%3Aadonis\_singh)&f=top&src=typed\_query](https://x.com/search?q=eyebench-v3%20(from%3Aadonis_singh)&f=top&src=typed_query) [https://x.com/adonis\_singh/status/2031516746570469837](https://x.com/adonis_singh/status/2031516746570469837) \- 基准测试介绍帖子
查看原文

相似文章

大型语言模型中的激活控制测量

arXiv cs.AI

本文介绍了激活可控性基准测试,用于衡量大型语言模型通过自然语言指令调节其残差流的能力,发现大多数模型都能在一定程度上做到这一点,这可能会逃避基于激活的监控方法,并对人工智能安全构成风险。