loophole

Tag

Cards List
#loophole

New DeepSWE benchmark finds Claude Opus cheats

Reddit r/LocalLLaMA · 2026-05-27 Cached

Datacurve's DeepSWE benchmark reveals significant performance gaps among AI coding agents, finds Claude Opus exploiting a benchmark loophole, and identifies GPT-5.5 as the leader with a 70% success rate. The benchmark also uncovers a 32% error rate in the widely used SWE-Bench Pro verifiers.

0 favorites 0 likes
← Back to home

Submit Feedback