Scaling Automatic Research Agents via World Models
Summary
This paper introduces World Model RL to scale automatic research agents by replacing environment execution with a learned world model, thereby accelerating post-training by 3-4x and enabling smaller agents to outperform larger ones on benchmarks.
View Cached Full Text
Cached at: 09/11/26, 02:14 AM
Paper page - Scaling Automatic Research Agents via World Models
Source: https://huggingface.co/papers/2608.12564
Abstract
World Model RL replaces costly environment execution with a learned world model and applies debiasing and denoising to accelerate post-training of autonomous research agents.
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of everyAutoResearchtrajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we proposeWorld Model RL(WMRL), which replaces environment execution with aworld modelto remove this bottleneck. Additionally, theworld modelcan be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations,Online DebiasingandInverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. BeyondAutoResearch, WMRL also transfers to post-trainingembodied VLA policies, which demonstrates the generalizability of our method.
View arXiv pageView PDFAdd to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.12564 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2608.12564 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.12564 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Scaling Automatic Research Agents via World Models
This paper identifies a scalability bottleneck in RL-trained automatic research agents—environment execution dominates training cost—and proposes World Model RL (WMRL) with online debiasing and inverse-variance denoising to replace real execution, achieving 3–4x training speedups and better generalization.
DSWorld: A Data Science World Model for Efficient Autonomous Agents
DSWorld introduces a Data Science World Model that predicts environment state transitions to reduce costly trial-and-error in autonomous agents, achieving 14x acceleration in RL training and 3-6x in inference while maintaining competitive performance.
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
Agent-World introduces a self-evolving training framework for general agent intelligence that autonomously discovers real-world environments and tasks via the Model Context Protocol, enabling continuous learning. Agent-World-8B and 14B models outperform strong proprietary models across 23 challenging agent benchmarks.
Policy and World Modeling Co-Training for Language Agents
This paper introduces PaW, a co-training framework that adds auxiliary world modeling supervision to policy learning during on-policy RL rollouts, improving language agent training without additional computational overhead.
Building World Models with Agent Swarms
The author describes building a world model harness that coordinates research agents to maximize evaluation metrics, achieving a 25x score improvement in a multimodal masked reconstruction task using geospatial inputs.