Tag
PhysMent introduces an interactive benchmark that evaluates LLM physical reasoning via iterative tool-mediated experimentation with a MuJoCo physics simulator, revealing model weaknesses in multi-step procedural tasks.
SWE-Interact is a new testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering tasks, revealing that strong single-turn benchmark performance does not reliably transfer to interactive, iterative workflows where agents must discover user intent and adapt to evolving requirements.