在一个困难的新SWE基准测试ProgramBench上,GPT5.5 high/xhigh首次解决了任务,显著优于Opus 4.7

Reddit r/singularity 模型

摘要

GPT5.5在困难的ProgramBench SWE基准测试中首次实现求解,显著优于Opus 4.7。

推文链接:https://x.com/KLieret/status/2054215545663144217?s=20 GitHub链接:[https://github.com/facebookresearch/ProgramBench/](https://github.com/facebookresearch/ProgramBench/) ProgramBench网站链接:[https://programbench.com/blog/gpt-5-5-first-solve/](https://programbench.com/blog/gpt-5-5-first-solve/)
查看原文

相似文章

@browser_use: Opus 5 和 GPT-5.6 Sol 在此不相上下!

X AI KOLs Timeline

Alexander Yue 推出了一项新的浏览器使用基准测试,其中 Opus 5 和 GPT-5.6 Sol 表现相似,强调了基准测试的稳健设计,包括为大语言模型评委提供的经过验证的评分标准。