CompoWorld: Compositional Environment Scaling for General Agents
Summary
CompoWorld introduces a method to scale tasks by composing reusable services for training general agents, improving performance by 9.17 points on average across benchmarks and surpassing models like Claude Opus on AutomationBench.
View Cached Full Text
Cached at: 09/29/26, 04:11 AM
Paper page - CompoWorld: Compositional Environment Scaling for General Agents
Source: https://huggingface.co/papers/2609.33665 Published on Sep 27
#3 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
Automaticallygeneratedenvironmentsprovideascalablesourceofinteractiondatafortraininggeneralagents.However,existingapproachesmainlygeneratetaskswithinasingleenvironment,whilereal-worldworkflowsrequireagentstoconnectinformationandactionsacrossmultipleservices.WeintroduceCompositionalEnvironmentScaling(CompoWorld),whichexpandsthetaskspacebycomposingafinitelibraryofreusableservices.Codingagentsturntoolspecificationsintoverifiedserviceswithtypedstatesandsharedinterfaces,whileaworldmodelhandlestoolsthatcannotbereliablyimplemented.Arandom-walkprocedureconnectsservicesthroughdependencygraphs,enablingthegenerationandverificationoftasksthatrequireinformationtoflowacrossservices.Verifiedtrajectoriessupportsupervisedfine-tuning(SFT),whileourCompletion-FocusedRubricRewardguidesreinforcementlearning(RL)towardfulltaskcompletionbyemphasizingcriteriawithlowerpassrateswithineachrolloutgroup.Weconstruct448servicesexposing10,130toolsanduse3KSFTtrajectoriesand1KRLtaskstotrainQwen3.6-35B-A3B.ExperimentalresultsshowthatCompoWorldimprovesonitsbackboneby9.17pointsonaverageacrosseightbenchmarks.OnAutomationBench,itsurpassesfrontiermodelssuchasClaudeOpus4.6andleadsallcomparedagent-specialized35B-A3Bmodels.
View arXiv pageView PDFGitHub0Add to collection
Get this paper in your agent:
hf papers read 2609\.33665
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33665 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33665 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33665 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
Agent-World introduces a self-evolving training framework for general agent intelligence that autonomously discovers real-world environments and tasks via the Model Context Protocol, enabling continuous learning. Agent-World-8B and 14B models outperform strong proprietary models across 23 challenging agent benchmarks.
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
AgentCompass is an open-source, lightweight, and extensible evaluation infrastructure for LLM-based agents, decoupling benchmarks, harness, and environment for flexible configurations. It supports over 20 benchmarks across five capability dimensions and provides fault-tolerant runtime and trajectory analysis tools.
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
Apodex 1.1 improves sustained, verifiable progress on complex real-world tasks by scaling executable environments and training agents for long-horizon coordination, achieving leading performance with a smaller 35B-parameter model.
Scaling Automatic Research Agents via World Models
This paper introduces World Model RL to scale automatic research agents by replacing environment execution with a learned world model, thereby accelerating post-training by 3-4x and enabling smaller agents to outperform larger ones on benchmarks.
@0xCodez: https://x.com/0xCodez/status/2058513716509913581
A comprehensive walkthrough on building multi-agent teams with Claude Managed Agents, covering role design, model mixing, and parallel execution to scale from one to 20 agents.