Tag
StateM is a runtime system that enhances long-horizon agent execution through durable states and recoverable runbooks, achieving 95.3% accuracy on Terminal-Bench 2.1 and significantly reducing API costs.
This article discusses the differences between Agent execution and general code execution and recommends two sandbox solutions: E2B (based on Firecracker) and OpenSandbox (based on Docker), which are suitable for production-grade and private deployment scenarios, respectively.
This paper introduces OpenClawBench, a large-scale dataset for benchmarking process-side anomalies in real-world AI agent execution trajectories. It reveals that task success can hide process failures, with 9.33% of oracle-passing executions containing anomalies, and provides structured supervision via a novel taxonomy.