CompoWorld: Compositional Environment Scaling for General Agents

Hugging Face Daily Papers Papers

Summary

CompoWorld introduces a method to scale tasks by composing reusable services for training general agents, improving performance by 9.17 points on average across benchmarks and surpassing models like Claude Opus on AutomationBench.

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (CompoWorld), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:11 AM

Paper page - CompoWorld: Compositional Environment Scaling for General Agents

Source: https://huggingface.co/papers/2609.33665 Published on Sep 27

#3 Paper of the day Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

Automaticallygeneratedenvironmentsprovideascalablesourceofinteractiondatafortraininggeneralagents.However,existingapproachesmainlygeneratetaskswithinasingleenvironment,whilereal-worldworkflowsrequireagentstoconnectinformationandactionsacrossmultipleservices.WeintroduceCompositionalEnvironmentScaling(CompoWorld),whichexpandsthetaskspacebycomposingafinitelibraryofreusableservices.Codingagentsturntoolspecificationsintoverifiedserviceswithtypedstatesandsharedinterfaces,whileaworldmodelhandlestoolsthatcannotbereliablyimplemented.Arandom-walkprocedureconnectsservicesthroughdependencygraphs,enablingthegenerationandverificationoftasksthatrequireinformationtoflowacrossservices.Verifiedtrajectoriessupportsupervisedfine-tuning(SFT),whileourCompletion-FocusedRubricRewardguidesreinforcementlearning(RL)towardfulltaskcompletionbyemphasizingcriteriawithlowerpassrateswithineachrolloutgroup.Weconstruct448servicesexposing10,130toolsanduse3KSFTtrajectoriesand1KRLtaskstotrainQwen3.6-35B-A3B.ExperimentalresultsshowthatCompoWorldimprovesonitsbackboneby9.17pointsonaverageacrosseightbenchmarks.OnAutomationBench,itsurpassesfrontiermodelssuchasClaudeOpus4.6andleadsallcomparedagent-specialized35B-A3Bmodels.

View arXiv pageView PDFGitHub0Add to collection

Get this paper in your agent:

hf papers read 2609\.33665

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33665 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33665 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33665 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

arXiv cs.AI

AgentCompass is an open-source, lightweight, and extensible evaluation infrastructure for LLM-based agents, decoupling benchmarks, harness, and environment for flexible configurations. It supports over 20 benchmarks across five capability dimensions and provides fault-tolerant runtime and trajectory analysis tools.

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Hugging Face Daily Papers

Apodex 1.1 improves sustained, verifiable progress on complex real-world tasks by scaling executable environments and training agents for long-horizon coordination, achieving leading performance with a smaller 35B-parameter model.

Scaling Automatic Research Agents via World Models

Hugging Face Daily Papers

This paper introduces World Model RL to scale automatic research agents by replacing environment execution with a learned world model, thereby accelerating post-training by 3-4x and enabling smaller agents to outperform larger ones on benchmarks.