iterative-development

Tag

Cards List
#iterative-development

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

arXiv cs.AI · 2026-05-26 Cached

Introduces EvoCode-Bench, a benchmark of 26 stateful coding tasks across 227 rounds that evaluates coding agents in multi-turn iterative interactions, revealing that single-round performance overestimates multi-round capabilities by 22–40 points.

0 favorites 0 likes
← Back to home

Submit Feedback