Learning to Reason with Curriculum II: Compositional Generalization
Summary
This paper theoretically analyzes how curriculum learning, by decomposing complex problems into simpler sub-problems and composing solutions, can dramatically reduce the sample complexity of learning to simulate sequential computations (semiautomata) compared to direct methods, achieving subpolynomial supervision requirements in supervised fine-tuning and exponentially weaker coverage conditions in reinforcement learning with verifiable rewards.
View Cached Full Text
Cached at: 06/29/26, 05:25 AM
# Learning to Reason with Curriculum II: Compositional Generalization
Source: [https://arxiv.org/abs/2606.27721](https://arxiv.org/abs/2606.27721)
[View PDF](https://arxiv.org/pdf/2606.27721)
> Abstract:Compositional generalization, the ability to solve complex problems by combining solutions to simpler sub\-problems, is a fundamental capability of both natural and artificial intelligence, and a key mechanism underlying chain\-of\-thought reasoning\. However, the theoretical underpinnings of compositional generalization remain poorly understood: when and why does decomposing a problem into parts yield more efficient learning than solving it directly? We study this question through the canonical problem of learning to simulate semiautomata \(predicting the outcome of $T$ steps of sequential computation\), a model that captures state tracking, regular language recognition, and modular arithmetic\. We show that an autocurriculum\-based approach building on Part I of this series, recursively decomposing longer sequences into shorter sub\-problems, learning to solve them, and composing the solutions, achieves dramatically better statistical complexity than direct methods\. \(i\) For a setting inspired by supervised fine\-tuning \(SFT\) where the learner receives interactive feedback on intermediate states of the computation, curriculum facilitates learning from only $2^\{\\mathcal\{O\}\(\\sqrt\{\\log T\}\)\}$ tokens of supervision; i\.e\., subpolynomial in the sequence length $T$, overcoming the $\\Omega\(T\)$ token barrier required by direct simulation\. \(ii\) For a setting inspired by reinforcement learning with verifiable rewards \(RLVR\), where the learner improves a pre\-trained reference model using an outcome verifier, we show that curriculum reduces the requirement on the reference model from coverage at the full sequence length $T$ to coverage at a shorter block length $B \\ll T$, an exponentially weaker condition\.
## Submission history
From: Nived Rajaraman \[[view email](https://arxiv.org/show-email/ba6853cb/2606.27721)\] **\[v1\]**Fri, 26 Jun 2026 05:09:08 UTC \(119 KB\)Similar Articles
RL Post-Training Builds Compositional Reasoning Strategies
This paper investigates whether reinforcement learning post-training can compose primitive skills into higher-level compositional strategies, using a fully observable rewrite-grammar environment. The authors find that RL reorganizes primitive competence through phased compositional mechanisms, while rejection fine-tuning plateaus due to producing many invalid shortcut-like rewrites.
Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
The paper introduces RACES, a recursive automated composition framework that treats verifiable environments as composable building blocks to scale reinforcement learning for LLMs, enabling efficient reasoning generalization through compositional operators.
What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA
This paper empirically studies how the composition of training data (curriculum) affects the skills learned by RL-based memory agents in multi-session question answering. It finds that curriculum composition acts as a fine-grained lever on specialization, with mixed benchmarks yielding the best overall performance and narrow out-of-domain sets transferring targeted temporal reasoning skills.
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR
This paper proposes Transfer-Aware Curriculum (TAC), a bandit-style online curriculum for multi-domain RLVR that prioritizes domains whose updates benefit other domains using gradient-geometry alignment. TAC improves macro-averaged accuracy on Qwen3-1.7B and Llama3.2-3B over fixed and learnability-only curricula.
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
This paper introduces a framework that decomposes compound logical answer options into atomic judgments and uses an operator-constrained integer linear program to improve large language model reasoning over AND, OR, and NEITHER/NOR operators. It achieves significant F1 gains on LOGICAL-COMMONSENSEQA and a new benchmark LOGICAL-SATA.