swe-bench

Tag

Cards List
#swe-bench

@heyshrutimishra: OH MY GOD CHINA JUST MATCHED USA FRONTIER CODING AI AT 40-60% LOWER TOKEN COST. XIAOMI JUST DROPPED MiMo-V2.5-Pro score…

X AI KOLs Following · 2026-04-22 Cached

Xiaomi released MiMo-V2.5-Pro, a coding AI scoring 73.7 on SWE-Bench Pro (near Claude Opus 4.6's 77.1) at 40-60% lower token cost than US frontier models.

0 favorites 0 likes
#swe-bench

@heyshrutimishra: OpenClaw users are gonna love it Finally an open source model that beats Opus 4.6 on SWE-Bench It's Kimi K2.6, it runs …

X AI KOLs Following · 2026-04-22 Cached

Kimi K2.6 open-source model surpasses Opus 4.6 on SWE-Bench, supporting 12+ hour autonomous coding sessions with 4,000+ tool calls.

0 favorites 0 likes
#swe-bench

Why we no longer evaluate SWE-bench Verified

OpenAI Blog · 2026-02-23 Cached

OpenAI announces it will no longer report SWE-bench Verified scores, citing two critical issues: 59.4% of failed problems have flawed test cases that reject correct solutions, and frontier models have seen benchmark problems during training, making improvements reflect training data exposure rather than genuine capability gains.

0 favorites 0 likes
#swe-bench

Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet

Anthropic Engineering · 2026-05-08 Cached

Anthropic's updated Claude 3.5 Sonnet achieves a new state-of-the-art 49% on the SWE-bench Verified benchmark, demonstrating significant capabilities in autonomous software engineering tasks.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback