Scores in the currently ongoing AtCoder heuristics finals.(Will be ongoing till 8th july 19:00 JST)
Summary
Scores update from the ongoing AtCoder heuristics finals, which will continue until July 8, 19:00 JST.
Similar Articles
How can Deepseek v4 top the coding leaderboards and still sit 8 months behind the frontier?
Analysis of DeepSeek V4's top coding scores versus its reported 8-month gap behind the frontier, highlighting differences between narrow benchmark optimization and broader reasoning tests, plus the practical performance hit when running quantized local versions.
humanity's last exam current benchmarks thoughts?
Discussion of recent AI model scores on the 'humanity's last exam' benchmark, noting improvement from GPT-4o's 2.7% in May 2024 to around 45% by June 2026, questioning the exam's difficulty.
@akshdeeps_001: New leaderboard just dropped
A new leaderboard related to AI or technology benchmarks has been announced in a tweet by @akshdeeps_001.
What are you doing this weekend?
A developer shares their weekend project of building a low-level infix language that compiles to WebAssembly, and offers a personal ranking of AI coding tools from contextual autocomplete to frontier models.
@_philschmid: https://x.com/_philschmid/status/2081744861829414977
EvoCode-Bench is a multi-turn coding benchmark with 26 tasks across 5 domains, designed to evaluate AI agents on evolving specifications and cumulative testing in a persistent workspace, revealing that single-turn scores dramatically overstate reliability.