runtime-harness

Tag

Cards List
#runtime-harness

EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents

arXiv cs.LG · 2d ago Cached

Introduces EvoHarness-RL, a framework that learns runtime harness policies for long-horizon LLM agents, enabling them to construct and update external state (belief, progress, experience) during task execution. Using Qwen3-8B on ALFWorld, it achieves 96.9% success and reveals harness annealing and evolution dynamics.

0 favorites 0 likes
#runtime-harness

@gm8xx8: Harness-R1 keeps the target model frozen and instead learns to edit the runtime around it from failure trajectories, mo…

X AI KOLs Timeline · 5d ago Cached

Harness-R1 introduces a method that learns to edit executable runtime harnesses for LLM agents from failure trajectories, using a separate 9B engineer optimized with online GRPO while keeping the target model frozen. It improves success rates across WebShop, ALFWorld, and DBBench.

0 favorites 0 likes
#runtime-harness

@omarsar0: // Adapt the Interface, Not the Model // I am fascinated by the results across my cheap-model-plus-good-harness builds.…

X AI KOLs Following · 2026-05-23 Cached

Proposes Life-Harness, a method that improves frozen LLM agents by adapting the runtime interface instead of model weights, achieving an average 88.5% relative improvement across 126 settings and 18 backbones.

0 favorites 0 likes
← Back to home

Submit Feedback