agent-execution

Tag

Cards List
#agent-execution

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Hugging Face Daily Papers · 2026-08-15 Cached

StateM is a runtime system that enhances long-horizon agent execution through durable states and recoverable runbooks, achieving 95.3% accuracy on Terminal-Bench 2.1 and significantly reducing API costs.

0 favorites 0 likes
#agent-execution

@seclink: The biggest difference between Agent execution and general code execution is: Agent execution requires extremely low cold start time (millisecond-level response), frequent file system state synchronization (Agent needs to read and write intermediate code, output files), and flexible API/network access control. Below are two condensed core recommended solutions: Solution 1…

X AI KOLs Timeline · 2026-07-08 Cached

This article discusses the differences between Agent execution and general code execution and recommends two sandbox solutions: E2B (based on Firecracker) and OpenSandbox (based on Docker), which are suitable for production-grade and private deployment scenarios, respectively.

0 favorites 0 likes
#agent-execution

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

arXiv cs.AI · 2026-05-29 Cached

This paper introduces OpenClawBench, a large-scale dataset for benchmarking process-side anomalies in real-world AI agent execution trajectories. It reveals that task success can hide process failures, with 9.33% of oracle-passing executions containing anomalies, and provides structured supervision via a novel taxonomy.

0 favorites 0 likes
← Back to home

Submit Feedback