Tag
This paper proposes environment evolution to incrementally increase task difficulty for terminal agents off-policy, enhancing continuous learning signals and improving performance on benchmarks like Terminal-Bench 2.1.
Introduces OpenART, a large-scale arena for red-teaming AI agents via open-ended environment evolution, with over 10K stateful scenarios across 50 domains, plus EMHA, a black-box hypergraph attack achieving 85% ASR across 75 agent-model configurations.