Dyad: Extending Large Language Models with Native Typed Decision-Making

arXiv cs.LG Papers

Summary

The paper introduces Dyad, an architecture that augments LLMs with an environment-conditioned action encoder for native typed decision-making, showing that training the encoder alone enables modular adaptation and joint training outperforms conventional RL post-training, including a 3.80% gain on ALFWorld with a 9B model.

arXiv:2609.36116v1 Announce Type: new Abstract: We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows. We investigate two complementary reinforcement learning settings driven by environment interaction. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters. Jointly optimizing both components outperforms conventional RL post-training across diverse interactive tasks and model scales, including a 3.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding.
Original Article
View Cached Full Text

Cached at: 09/30/26, 09:46 AM

# Dyad: Extending Large Language Models with Native Typed Decision-Making
Source: [https://arxiv.org/html/2609.36116](https://arxiv.org/html/2609.36116)
Yundaichuan Zhan Weishi Wang Wenbiao Liu Daniel DahlmeierChengwei Qin Juncheng Li Fredrik D\. Johansson Zhongqi Yue††thanks:Corresponding author\.Affiliation:Zhejiang UniversityAffiliation:The Hong Kong University of Science and Technology \(Guangzhou\)Affiliation:Chalmers University of TechnologyAffiliation:Microsoft Research

###### Abstract

We study how to build more capable general\-purpose agents by extending large language models \(LLMs\) with native typed decision\-making\. We introduceDyad, an architecture that augments a pretrained LLM with an environment\-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM’s internal state to yield a distribution over typed actions\. By factorizing decision\-making into representations of the evolving interaction state and environment\-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows\. We investigate two complementary reinforcement learning settings driven by environment interaction\. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters\. Jointly optimizing both components outperforms conventional RL post\-training across diverse interactive tasks and model scales, including a 3\.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding\.

## 1Introduction

Solving tasks through interaction with an environment requires two fundamentally different forms of computation\([Yao et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib26)\): deliberate reasoning about goals or plans, and typed decision\-making over available actions, analogous to Kahneman’s System 2 and System 1, respectively\([Kahneman, 2011](https://arxiv.org/html/2609.36116#bib.bib1)\)\. For example, a shopping agent must reason about a customer’s requirements and compare products across web pages \(System 2\), while repeatedly selecting from the available actions in the active page to search, inspect products, choose options, and complete a purchase \(System 1\)\.

Today, LLM\-based agents typically handle both forms through autoregressive generation,i\.e\., reasoning in language and generating text to specify actions, as illustrated in the middle of Figure[1](https://arxiv.org/html/2609.36116#S1.F1)\. However, this overloads the language\-generation interface: the LLM assigns probabilities to vocabulary tokens rather than directly to permissible typed actions\([Yue et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib27)\)\. Recent dedicated System 1 models, such as Jev\([Almeida, 2026](https://arxiv.org/html/2609.36116#bib.bib2)\)and its open\-source alternatives\([Nandakishor, 2026](https://arxiv.org/html/2609.36116#bib.bib29);[Featherless AI, 2026](https://arxiv.org/html/2609.36116#bib.bib31)\), address this limitation by predicting typed decisions directly and in parallel, rather than generating their textual representations token by token, as shown on the left of Figure[1](https://arxiv.org/html/2609.36116#S1.F1)\.

Yet language reasoning and typed decision\-making need not be developed independently, as systems 1 and 2 are complementary modes of cognition rather than distinct physical systems\([Kahneman, 2011](https://arxiv.org/html/2609.36116#bib.bib1)\): reasoning guides decisions, while decision outcomes inform subsequent reasoning\. Recent works\([Yue et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib27);[Zhao et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib28)\)take a step toward this integration by extending LLMs with direct environment actions, enabling interleaved language reasoning and decision\-making\. However, they are limited to fixed, pre\-defined action spaces, and do not generalize to dynamically provided actions in unseen environments \(e\.g\., unseen web pages\)\. This raises a natural question: can pretrained LLMs support native typed decision\-making over dynamic action spaces, and can jointly learning reasoning and decision\-making yield more capable agents?

Figure 1:Conceptual comparison of System 1, System 2, and Dyad for reasoning and acting\.We introduceDyad, an architecture that extends pretrained LLMs with*native typed decision\-making*while retaining their autoregressive language reasoning capabilities, as illustrated on the right of Figure[1](https://arxiv.org/html/2609.36116#S1.F1)\. Specifically, we augment the pretrained LLM with an action encoder that embeds available action descriptions independently and in parallel, conditioned on the environment\. Because these action representations depend on the environment and action semantics, but not on the evolving interaction state, they can be precomputed and reused across decision steps whenever the action definitions remain unchanged\. When a decision is required, these embeddings are scored against the LLM’s internal state representation to produce a distribution directly over the available typed actions\. This factorization separates the evolving interaction state from environment\-specific action semantics, providing an inductive bias for learning reusable representations rather than an entangled state–action mapping\. It further amortizes action encoding across interactions, enabling efficient decision\-time scoring and adaptation to previously unseen action spaces\.

Dyad is trained in two complementary settings, both driven by end\-to\-end rewards from environment interaction\. First, we freeze the pretrained LLM and train only the action encoder through reinforcement learning\. The resulting agent consistently improves performance across four unseen environments, demonstrating that native typed decision\-making can enhance existing LLMs without modifying their pretrained parameters\. Second, we jointly optimize the LLM and action encoder on diverse interactive tasks\. Dyad outperforms agentic RL baselines across model scales\. On ALFWorld, it achieves a 3\.80% average absolute gain with Qwen3\.5\-9B\. Dyad also better preserves general knowledge, reasoning, and coding capabilities\.

In summary, our contributions are threefold:

- •We introduce Dyad, to our knowledge the*first*architecture to extend pretrained LLMs with native, parallel typed decision\-making over dynamically provided action spaces\.
- •We develop an end\-to\-end inference and training pipeline integrated with vLLM\([Kwon et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib3)\)and verl\([Sheng et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib4)\), which we plan to open\-source\.
- •We show that Dyad improves agent performance across both frozen\-LLM adaptation and joint optimization, generalizes across environments and model scales, and yields broad gains in general knowledge, reasoning, and coding\.

## 2Dyad: Native Typed Decision\-Making in Pretrained LLMs

### 2\.1Problem Formulation

#### Setting\.

We consider an LLM agent that interleaves language generation with direct interaction in an external environmentee\. At steptt, the environment exposes a set of permissible typed actions𝒜t\\mathcal\{A\}\_\{t\}, which may change over time\. Eacha∈𝒜ta\\in\\mathcal\{A\}\_\{t\}is an environment\-defined operation, potentially with arguments, that can be executed directly\. This setting encompasses both direct decisions, such as routing a support request to a team, and tasks combining language reasoning with action, such as navigating unfamiliar software\.

#### Policy\.

The agent observes a language historyhth\_\{t\}, comprising previously generated tokens and textual observations from its interactions\. Its policyπθ\\pi\_\{\\theta\}supports two modes: language generation, with a distributionπθLM​\(v∣ht,e\)\\pi\_\{\\theta\}^\{\\mathrm\{LM\}\}\(v\\mid h\_\{t\},e\)over the vocabulary𝒱\\mathcal\{V\}, and typed decision\-making, with a distributionπθact​\(a∣ht,e,𝒜t\)\\pi\_\{\\theta\}^\{\\mathrm\{act\}\}\(a\\mid h\_\{t\},e,\\mathcal\{A\}\_\{t\}\)over the currently permissible actions𝒜t\\mathcal\{A\}\_\{t\}\.

#### Interaction\.

Generating a tokenv∈𝒱v\\in\\mathcal\{V\}extends the language history, whereas selecting a typed actiona∈𝒜ta\\in\\mathcal\{A\}\_\{t\}invokes the corresponding environment operation\. The resulting textual observation is appended to the history; for example, anopen\(fridge\)action may return a description of the fridge’s contents\. The environment provides rewards based on task outcomes, such as successfully retrieving a requested item from the fridge\. The agent’s objective is to maximize the expected cumulative reward from environment interaction\.

To realize such a policy, existing approaches often jointly encode the interaction state and available actions, entangling action semantics with the evolving state\([Nandakishor, 2026](https://arxiv.org/html/2609.36116#bib.bib29);[Lee, 2026](https://arxiv.org/html/2609.36116#bib.bib30)\)\. We introduce Dyad, which instead factorizes typed decision\-making into pretrained LLM state representations within a single architecture \(Section[2\.2](https://arxiv.org/html/2609.36116#S2.SS2)\), followed by its learning procedure \(Section[3](https://arxiv.org/html/2609.36116#S3)\)\.

### 2\.2Architecture

![Refer to caption](https://arxiv.org/html/2609.36116v1/DynamicExpA.png)Figure 2:Dyad architecture\.State and environment\-conditioned action representations are computed independently\. Generating the interaction tokenvactv\_\{\\mathrm\{act\}\}triggers typed action selection; otherwise, the pretrained LM head generates language tokens\.Dyad extends a pretrained LLM with an action encoder, enabling language generation and typed decision\-making within a single policy\. Letffdenote the pretrained LLM’s transformer backbone,𝐖LM\\mathbf\{W\}\_\{\\mathrm\{LM\}\}its language modeling head, andggthe action encoder\. Given the interaction historyhth\_\{t\}and environment descriptiondesc⁡\(e\)\\mathrm\{desc\}\(e\), the two components produce a state representation𝐪t\\mathbf\{q\}\_\{t\}and an environment\-conditioned representation𝐳a\\mathbf\{z\}\_\{a\}for each available actiona∈𝒜ta\\in\\mathcal\{A\}\_\{t\}:

𝐪t=f⁡\(desc⁡\(e\)⊕ht\),𝐳a=g⁡\(desc⁡\(e\)⊕desc⁡\(a\)\),\\mathbf\{q\}\_\{t\}=f\(\\mathrm\{desc\}\(e\)\\oplus h\_\{t\}\),\\qquad\\mathbf\{z\}\_\{a\}=g\(\\mathrm\{desc\}\(e\)\\oplus\\mathrm\{desc\}\(a\)\),\(1\)
wheredesc⁡\(⋅\)\\mathrm\{desc\}\(\\cdot\)denotes textual description and⊕\\oplusdenotes concatenation\. Dyad retains the pretrained language modeling head to generate tokens according to

πθLM​\(v∣ht,e\)=softmax⁡\(𝐖LM​𝐪t\)v\.\\pi\_\{\\theta\}^\{\\mathrm\{LM\}\}\(v\\mid h\_\{t\},e\)=\\operatorname\{softmax\}\\bigl\(\\mathbf\{W\}\_\{\\mathrm\{LM\}\}\\mathbf\{q\}\_\{t\}\\bigr\)\_\{v\}\.\(2\)Upon generating a dedicated interaction tokenvact∈𝒱v\_\{\\mathrm\{act\}\}\\in\\mathcal\{V\}, Dyad selects a typed action by scoring each available action against the state representation:

πθact​\(a∣ht,e,𝒜t\)=softmax⁡\(𝐙t​𝐪t\)a,𝐙t=\[𝐳a1⋯𝐳a\|𝒜t\|\]⊤\.\\pi\_\{\\theta\}^\{\\mathrm\{act\}\}\(a\\mid h\_\{t\},e,\\mathcal\{A\}\_\{t\}\)=\\operatorname\{softmax\}\\bigl\(\\mathbf\{Z\}\_\{t\}\\mathbf\{q\}\_\{t\}\\bigr\)\_\{a\},\\qquad\\mathbf\{Z\}\_\{t\}=\\begin\{bmatrix\}\\mathbf\{z\}\_\{a\_\{1\}\}&\\cdots&\\mathbf\{z\}\_\{a\_\{\|\\mathcal\{A\}\_\{t\}\|\}\}\\end\{bmatrix\}^\{\\top\}\.\(3\)The key design choice in Dyad is to factorize the computation ofπθact​\(a∣ht,e,𝒜t\)\\pi\_\{\\theta\}^\{\\mathrm\{act\}\}\(a\\mid h\_\{t\},e,\\mathcal\{A\}\_\{t\}\)into a state representation𝐪t\\mathbf\{q\}\_\{t\}and an environment\-conditioned action representation𝐳a\\mathbf\{z\}\_\{a\}with Eq\.[1](https://arxiv.org/html/2609.36116#S2.E1)\. This encodes the inductive bias that the evolving interaction state can be represented independently of the currently available actions, while an action’s meaning depends on its environment rather than the current state\. The action encoder therefore learns to express action semantics in the pretrained LLM’s representation space, enabling their compatibility to be measured directly\. Moreover, independent action encoding makes the decision policy permutation\-equivariant: reordering the available actions simply reorders their probabilities\. We detail the two representations below\.

State Representation\.We take the pretrained LLM’s hidden state at the current decision position as𝐪t\\mathbf\{q\}\_\{t\}\. Pretrained for next\-token prediction, the backbone provides contextual features that support both language generation and typed decision\-making\.

Action Representation\.We implementggas a transformer that maps the concatenated environment and action descriptions to an embedding compatible with𝐪t\\mathbf\{q\}\_\{t\}\. Action embeddings are computed independently and in parallel, and can be reused across decision steps as long as their descriptions and environment context remain unchanged\.

Amortized Action Encoding\.Action representations can be computed once and reused whenever their action descriptions and environment context remain unchanged\. Newly introduced actions need only be encoded when first encountered\. Consequently, steady\-state decision\-time inference requires only the pretrained LLM and cached action representations\.

## 3Learning Dyad from Environment Interaction

We first initialize the action encoder to align its representations𝐳a\\mathbf\{z\}\_\{a\}with the LLM’s state representations𝐪t\\mathbf\{q\}\_\{t\}, then optimize Dyad through reinforcement learning \(RL\) from environment interaction\.

### 3\.1Action Encoder Initialization

The action encoderggconsists of a pretrained transformer backbone, separate fromff, and a projector that maps its outputs to action representations𝐳a\\mathbf\{z\}\_\{a\}with the same dimensionality as𝐪t\\mathbf\{q\}\_\{t\}\. To initialize the projector, we synthesize a small collection of action\-selection examples𝒟=\{\(hi,ei,𝒜i,ai⋆\)\}i=1N\\mathcal\{D\}=\\\{\(h\_\{i\},e\_\{i\},\\mathcal\{A\}\_\{i\},a\_\{i\}^\{\\star\}\)\\\}\_\{i=1\}^\{N\}\. For each example, we sample a permissible action set𝒜i\\mathcal\{A\}\_\{i\}for environmenteie\_\{i\}and select a target actionai⋆∈𝒜ia\_\{i\}^\{\\star\}\\in\\mathcal\{A\}\_\{i\}\. A teacher LLM then generates task context and reasoninghih\_\{i\}that support selectingai⋆a\_\{i\}^\{\\star\}\. We initialize the projector by minimizing the cross\-entropy loss, while keepingff,𝐖LM\\mathbf\{W\}\_\{\\mathrm\{LM\}\}, and the action encoder backbone frozen:

ℒalign=−𝔼\(h,e,𝒜,a⋆\)∼𝒟​\[log⁡πθact​\(a⋆∣h,e,𝒜\)\]\.\\mathcal\{L\}\_\{\\mathrm\{align\}\}=\-\\mathbb\{E\}\_\{\(h,e,\\mathcal\{A\},a^\{\\star\}\)\\sim\\mathcal\{D\}\}\\left\[\\log\\pi\_\{\\theta\}^\{\\mathrm\{act\}\}\\bigl\(a^\{\\star\}\\mid h,e,\\mathcal\{A\}\\bigr\)\\right\]\.\(4\)This lightweight initialization aligns the two representation spaces before learning from environment rewards\. Further details are provided in Appendix[A\.2](https://arxiv.org/html/2609.36116#A1.SS2)\.

### 3\.2Agentic Reinforcement Learning

Starting from the initialized model, we optimize Dyad using rewards from environment interaction\. Trajectories interleave language generation and typed decision\-making under the policy defined in Eqs\.[2](https://arxiv.org/html/2609.36116#S2.E2)and[3](https://arxiv.org/html/2609.36116#S2.E3)\. We apply GRPO or GiGPO using task\-outcome rewards, treating sampled vocabulary tokens and typed actions as decisions under their respective policy distributions\. We consider two training settings\. In joint optimization, we update both the pretrained LLM and the entire action encoder\. In frozen\-LLM adaptation, we keepffand𝐖LM\\mathbf\{W\}\_\{\\mathrm\{LM\}\}fixed and update onlygg, including its backbone and projector\. Further optimization details are provided in Appendix[A\.3](https://arxiv.org/html/2609.36116#A1.SS3)\.

## 4Related Work

#### Language\-Based Action Interfaces\.

LLM agents commonly interact with external environments by generating textual action specifications, as in reasoning and tool\-use agents\([Yao et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib26);[Schick et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib23)\)\. Program\-guided methods use generated code to invoke external computation\([Gao et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib25);[Gou et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib20)\), while interactive agents extend this interface to web navigation and multi\-step environment control\([Zhou et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib21);[Liu et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib22)\)\. These approaches offer a flexible interface, but express environment actions through vocabulary\-token prediction rather than a policy directly over permissible typed actions\.

#### Typed Decision\-Making\.

Dedicated decision models such as Jev\([Almeida, 2026](https://arxiv.org/html/2609.36116#bib.bib2)\)predict distributions directly over provided choices\. As summarized in Table[1](https://arxiv.org/html/2609.36116#S4.T1), such models admit several architectural parameterizations\. Joint action\-set encoders, as in Laya\([Nandakishor, 2026](https://arxiv.org/html/2609.36116#bib.bib29)\), process the interaction state and available actions together, producing decision scores for all actions in parallel\. AR encoders instead process each state–action pair to assign a scalar score, as exemplified by the reranker\-based approach in SemIf\([Lee, 2026](https://arxiv.org/html/2609.36116#bib.bib30)\)\. Single\-token AR methods assign vocabulary identifiers to available choices and score them using the LLM’s language modeling head, as in Simple Jev\([Featherless AI, 2026](https://arxiv.org/html/2609.36116#bib.bib31)\)\. However, these approaches entangle action scoring with the evolving interaction state or a fixed vocabulary basis\.

Table 1:Conceptual comparison of methods for typed decision\-making\.FFdenotes a model jointly processing its specified inputs, andid⁡\(a\)\\operatorname\{id\}\(a\)maps an available action to a vocabulary identifier\.
#### Native Actions in Pretrained LLMs\.

Beyond dedicated decision models, recent work integrates language reasoning and environment actions within pretrained LLMs\. ToolkenGPT\([Hao et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib24)\)introduces learned tool embeddings into the language modeling head, enabling tool invocation through autoregressive generation of special tokens\. ExpA\([Yue et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib27)\)goes beyond the vocabulary\-token interface by extending LLMs with directly executable actions, enabling interleaved language reasoning and native decision\-making\. However, both rely on predefined action spaces\. Dyad extends this direction by factorizing typed decision\-making into pretrained LLM state representations and independently encoded, environment\-conditioned action representations, constructing a native decision head over dynamically provided actions\.

## 5Experiments

### 5\.1Experimental Setup

Benchmarks\.We perform agentic RL separately on DIVE\([Chen et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib5)\)and CodeGym\([Du et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib6)\), and evaluate the resulting agents on ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2609.36116#bib.bib7)\), WebShop\([Yao et al\., 2022](https://arxiv.org/html/2609.36116#bib.bib8)\),τ2\\tau^\{2\}\-bench\([Barres et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib9)\), and SWE\-bench Verified \(SWE\)\([Jimenez et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib10)\)without further training on these environments\. We assess general knowledge, mathematical reasoning, and coding with MMLU\-Pro\([Wang et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib11)\),[HMMT February 2026](https://huggingface.co/datasets/MathArena/hmmt_feb_2026), and LiveCodeBench v6\([Jain et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib12)\), respectively\. We also train agents separately on ALFWorld and WebShop to evaluate Dyad’s in\-domain gains over the corresponding RL baselines\.

Baselines\.We compare Dyad\-GRPO with GRPO\([Shao et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib14)\)and Dyad\-GiGPO with GiGPO\([Feng et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib15)\), using the same LLM and RL training environment for each pair\. We also evaluate each model without additional training as a reference for agent performance and general capabilities\.

Implementation Details\.For joint training, we use Qwen3\.5\-4B, Qwen3\.5\-9B, and Qwen3\.5\-27B\([Qwen Team, 2026](https://arxiv.org/html/2609.36116#bib.bib16)\)\. For frozen\-LLM adaptation, we use the 4B and 9B Qwen models, Llama\-3\.2\-3B\-Instruct\([Meta, 2024](https://arxiv.org/html/2609.36116#bib.bib17)\), and Gemma\-4\-E2B\-it\([Abd et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib18)\)\. Dyad uses attention pooling as the default projector and follows the alignment and agentic RL procedure in Section[3](https://arxiv.org/html/2609.36116#S3)\. We implement the training and inference pipeline with verl\([Sheng et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib4)\)and vLLM\([Kwon et al\., 2023](https://arxiv.org/html/2609.36116#bib.bib3)\)\. Further implementation details are provided in Appendices[A](https://arxiv.org/html/2609.36116#A1)and[B](https://arxiv.org/html/2609.36116#A2)\.

### 5\.2Agentic Post\-training

Table 2:In\-domain performance of Qwen3\.5\-9B\.Evaluations use 3 seeds\.In\-Domain Performance\.When training and evaluation use the same environment, Dyad improves both GRPO and GiGPO on Qwen3\.5\-9B across ALFWorld and WebShop, outperforming the corresponding baseline in all four comparisons \(Table[2](https://arxiv.org/html/2609.36116#S5.T2)\)\. The largest absolute gain is 5\.89% on ALFWorld with GRPO\. GiGPO already outperforms GRPO in both environments, and Dyad provides further improvements over this stronger baseline\.

Cross\-Environment Performance\.When trained on DIVE or CodeGym and evaluated without further adaptation, Dyad outperforms the corresponding RL baseline in 41 of 48 model–algorithm–environment comparisons \(Table[3](https://arxiv.org/html/2609.36116#S5.T3)\)\. It also improves over the original pretrained model in 47 of 48 evaluations, with gains spanning both training environments and all three model scales\. These results show that the benefits of Dyad are not confined to the environment used for agentic RL\.

Table 3:Cross\-environment performance\.Agents are evaluated without further training\. Superscripts show absolute changes \(%\) from the corresponding RL baseline\. ALF denotes ALFWorld\. Evaluations use 3 seeds\.Figure 3:General capabilities\.Dashed lines indicate zero change relative to each model without additional training\. Evaluations are on 3 seeds\.General\-Capability Retention\.Learning both reasoning and environment control through vocabulary generation can create interference between the two objectives during agentic RL\([Yue et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib27)\)\. We therefore measure changes in general knowledge, mathematical reasoning, and coding relative to the original pretrained models \(Figure[3](https://arxiv.org/html/2609.36116#S5.F3)\)\. Dyad generally outperforms the corresponding GRPO and GiGPO baselines, both by recovering capability losses after DIVE training and by amplifying gains after CodeGym training\. Notably, both Dyad variants exceed the original models on every CodeGym capability evaluation, while the RL baselines still regress on general knowledge or coding in some settings\. Together with the agent\-performance gains, these results suggest that the action encoder provides an additional route for task adaptation, reducing interference with pretrained language capabilities\.

Training Dynamics\.Both Dyad variants attain higher final validation rewards than their corresponding RL baselines \(Figure[4](https://arxiv.org/html/2609.36116#S5.F4)\)\. As training progresses, Dyad produces substantially shorter responses, suggesting more concise and efficient reasoning, while increasing the number of environment interactions per trajectory\. This shift toward less generation and more interaction is especially beneficial for multi\-step tasks, where progress depends on repeatedly acting on and incorporating feedback from the environment\.

Figure 4:Training dynamics\.Curves show validation means for Qwen3\.5\-9B on CodeGym; shading indicates±1\\pm 1standard deviation\.
### 5\.3Modular Adaptation with a Frozen LLM

Table 4:Frozen\-LLM adaptation\.Dyad is trained with GRPO on DIVE\. Evaluations use 3 seeds\.Frozen\-LLM Adaptation\.With the LLM frozen, training the action encoder improves agent performance over the original model in most evaluated settings \(Table[4](https://arxiv.org/html/2609.36116#S5.T4)\)\. Both Qwen models improve on all four environments, and the gains extend to other model families, with improvements on WebShop for every model\. Joint training generally achieves higher scores, but the gap varies by model and environment: for both Qwen models and Gemma, frozen adaptation is close onτ2\\tau^\{2\}\-bench, whereas larger gaps remain on ALFWorld\. For both Qwen models, the gap on SWE is also small\. These results show that training only the action encoder can improve agent performance beyond the RL training environment without updating the LLM\.

### 5\.4System 1 Evaluation

Table 5:JevBench results\.Accuracy \(%\) on the public subset\. Best accuracy and latency \(s\) values are bold\.†\\daggerJev latency comes from published API measurements\. Benchmark and measurement details are provided in Appendix[B\.3](https://arxiv.org/html/2609.36116#A2.SS3)\.Comparison with Dedicated System 1 Models\.We evaluate Qwen3\.5\-4B \+ Dyad, trained on CodeGym using GRPO, against Jev on both direct decision\-making and agentic interaction\. On JevBench, Dyad with reasoning reaches 86\.18% overall accuracy, slightly exceeding Jev’s 85\.01%, with the largest advantage on the hard subset \(Table[5](https://arxiv.org/html/2609.36116#S5.T5)\)\. Without reasoning, Dyad remains competitive while reducing latency to 0\.31 s, compared with 0\.67 s reported for Jev\. On the more agentic ALFWorld benchmark, Dyad achieves higher success rates than both Jev and zero\-shot Qwen3\.5\-4B across all six task types, while requiring fewer interaction steps overall and on most task types \(Figure[5](https://arxiv.org/html/2609.36116#S5.F5)\)\. Together, these results show that Dyad can serve as an effective typed decision model while retaining autoregressive reasoning within the same policy\.

Figure 5:Comparison with Jev\.Qwen3\.5\-4B serves as the zero\-shot reference on ALFWorld\.
### 5\.5Ablation Studies

Figure 6:Ablation studies\.CodeGym validation reward with Qwen3\.5\-4B, GRPO, and attention pooling\. Both panels reuse the Dyad \(full\) curve; the horizontal dashed line denotes the alignment\-only reference\.Trainable Components\.Updating the full action encoder yields the highest final validation reward, with the LLM trainable in all variants \(Figure[6](https://arxiv.org/html/2609.36116#S5.F6)\)\. Training the projector while freezing the encoder backbone is competitive early on, but reward declines later and ends below the variant with the entire encoder frozen\. Even with fixed action representations, LLM updates improve reward over the shared aligned initialization\.

Alignment and Reinforcement Learning\.Alignment improves initial performance, and the full model outperforms RL\-only training throughout the evaluated budget \(Figure[6](https://arxiv.org/html/2609.36116#S5.F6)\)\. Although RL without alignment improves substantially, it does not close the gap\. Continuing RL from the aligned checkpoint also raises reward above the alignment\-only reference\. These comparisons support complementary roles for the two stages: supervised alignment provides a useful initialization, while agentic RL further improves task performance\.

Table 6:Action pooling ablation\.Projector Architecture\.Validation reward is highest with attention pooling, followed by MLP\-weighted pooling and mean pooling \(Table[6](https://arxiv.org/html/2609.36116#S5.T6)\)\. In Dyad, each pooled action vector forms a row of the action\-scoring matrix𝐙t\\mathbf\{Z\}\_\{t\}and is compared directly with the LLM state through a dot product \(Equation[3](https://arxiv.org/html/2609.36116#S2.E3)\)\. The projector therefore needs to produce representations suitable for this comparison from a separate encoder’s token features\. MLP\-weighted pooling learns token weights, but its output remains a weighted average of the encoder states\. Attention pooling additionally learns value and output projections, allowing the projector to transform the features used for action scoring \(Appendix[A\.1](https://arxiv.org/html/2609.36116#A1.SS1)\)\. This gives it a direct way to adapt the pooled representation to the LLM’s state representation space, in addition to learning token weights\. Its advantage over MLP\-weighted pooling is consistent with a benefit from combining token aggregation and feature transformation\.

## 6Conclusion

We introducedDyad, an architecture for extending pretrained LLMs with native typed decision\-making over dynamically provided action spaces\. Dyad factorizes action selection into a representation of the evolving interaction state and independently encoded, environment\-conditioned action representations, allowing the same pretrained LLM state to support both autoregressive language reasoning and direct typed decisions\. This design provides a reusable decision interface without requiring actions to be represented through the language\-model vocabulary\. Across agentic post\-training experiments, Dyad improves both in\-domain and cross\-environment performance over conventional RL baselines across model scales\. These gains are accompanied by shorter reasoning traces, more frequent environment interaction, and generally stronger performance on the evaluated knowledge, reasoning, and coding benchmarks\. Dyad also supports modular adaptation: training only the action encoder improves agent performance without modifying the pretrained LLM\. Finally, on dedicated System 1 evaluations, Dyad is competitive with specialized decision models while retaining language reasoning within the same policy\. Together, these results suggest that explicitly separating language reasoning from typed decision\-making is a promising foundation for general\-purpose agents that must reason and act across diverse, previously unseen environments\.

### AI use statement

We used generative AI to assist with synthetic data generation and literature search\. The authors take responsibility for the final content\.

### Ethics statement

This work studies decision\-making in language model agents using public benchmarks\. We encourage responsible use of the proposed method\.

### Reproducibility statement

Implementation details, data construction, experimental settings, hyperparameters, and prompt templates are provided in Appendices A and B to support reproducibility\.

## References

- Abdet al\.\(2026\)S\. E\. Abd, V\. Aggarwal, R\. Algayres, A\. Andreev, O\. Bachem, I\. Ballantyne, C\. Brick, V\. Carbune, M\. Casbon, M\. Chaturvedi, A\. Chawla, V\. Cotruta, A\. Coucke, P\. Culliton, R\. Dadashi, L\. Dixon, M\. Elhawaty, U\. Evci, C\. Farabet, J\. Ferret, F\. Galgani, S\. Girgin, J\. Grill, M\. Grootendorst, J\. Guo, C\. Hardin, Y\. He, S\. M\. Hernandez, O\. Homburger, L\. Hussenot, J\. Ji, A\. Joulin, A\. Kamath, P\. Kassraie, O\. Lacombe, P\. Lahoti, G\. Liu, G\. Martins, L\. Martins, T\. Matejovicova, R\. Merhej, N\. Momchev, S\. Mondal, R\. Mullins, S\. R\. Panyam, S\. Pathak, S\. Perrin, A\. S\. Pinto, E\. Pot, A\. Pouget, A\. Ramé, S\. Ramos, D\. Reid, D\. Rim, M\. Rivière, K\. Roth, L\. Rouillard, O\. Sanseviero, P\. G\. Sessa, S\. Settle, D\. Sinopalnikov, S\. Smoot, P\. Stanczyk, A\. Steiner, L\. Stewart, I\. O\. Tolstikhin, M\. Tschannen, A\. Tsitsulin, N\. Vieillard, R\. Wu, P\. Xu, H\. Yang, E\. Yvinec, B\. Zhang, L\. Zhang, J\. Zou, N\. Aagnes, A\. Abdelhamed, J\. Adámek, S\. Agrawal, S\. Agrawal, I\. Alabdulmohsin, J\. Alayrac, U\. Alon, C\. Amarnath, A\. Anand, C\. Anastasiou, S\. Ariafar, F\. Aubet, K\. Axiotis, F\. Barbero, J\. K\. Barral, A\. Bendebury, U\. Bergmann, S\. Bileschi, K\. Black, M\. Blondel, S\. Borgeaud, and A\. BrazinskasGemma 4 technical report\.CoRRabs/2607\.02770\.External Links:[Link](https://doi.org/10.48550/arXiv.2607.02770),[Document](https://dx.doi.org/10.48550/ARXIV.2607.02770),2607\.02770Cited by:[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p3.1)\.
- Almeida \(2026\)D\. AlmeidaIntroducing system one models & Jev\.Note:TypeSafe AI blogPublished September 15, 2026External Links:[Link](https://typesafe.ai/blog/introducing-system-one-models-and-jev)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p2.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px2.p1.1)\.
- Barreset al\.\(2026\)V\. Barres, H\. Dong, S\. Ray, X\. Si, and K\. R\. Narasimhanτ2\\tau^\{2\}\-Bench: evaluating conversational agents in a dual\-control environment\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=OC2z7iSQKa)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px5),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Chenet al\.\(2026\)A\. Chen, C\. Zhang, J\. Liu, J\. Chen, C\. Du, Y\. Li, M\. Zhong, Q\. Wang, Z\. Zhu, J\. Song, K\. Ji, J\. He, P\. Zhao, and Y\. XiaoDIVE: scaling diversity in agentic task synthesis for generalizable tool use\.InForty\-third International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=4uM4P5NaLU)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px1),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Dekonincket al\.\(2026\)J\. Dekoninck, N\. Jovanović, T\. Gehrunger, K\. Rögnvaldsson, I\. Petrov, C\. Sun, and M\. VechevBeyond benchmarks: MathArena as an evaluation platform for mathematics with LLMs\.In3rd AI for Math Workshop: Toward Self\-Evolving Scientific Agents,Note:ICML 2026 WorkshopExternal Links:[Link](https://openreview.net/forum?id=DmPE4byHuN)Cited by:[§B\.2](https://arxiv.org/html/2609.36116#A2.SS2.SSS0.Px2.p1.1)\.
- Duet al\.\(2026\)W\. Du, H\. Gong, Z\. Ling, K\. Liu, L\. Shen, X\. Yao, Y\. Xu, D\. Shi, Y\. Yang, and J\. ChenGeneralizable end\-to\-end tool\-use RL with synthetic CodeGym\.InInternational Conference on Learning Representations,pp\. 17733–17756\.External Links:[Link](https://openreview.net/forum?id=QRSeFZfu8E)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px2),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Featherless AI \(2026\)Featherless AISimple Jev: structured classification and scoring from next\-token logits\.Note:GitHub repositoryRevision c077d5df, accessed September 26, 2026External Links:[Link](https://github.com/featherless-ai/simple-jev)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p2.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px2.p1.1)\.
- Fenget al\.\(2025\)L\. Feng, Z\. Xue, T\. Liu, and B\. AnGroup\-in\-group policy optimization for LLM agent training\.InAdvances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2\-7, 2025 / Mexico City, Mexico, November 30 \- December 5, 2025,D\. Belgrave, C\. Zhang, L\. N\. Montoya, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, N\. Chen, I\. V\. M\. Ruíz, and A\. Loaiza\-Bonilla \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2025/hash/420c9f777c0b4f78d515e53cf74d58b2-Abstract-Conference.html)Cited by:[§A\.3](https://arxiv.org/html/2609.36116#A1.SS3.SSS0.Px1.p1.3),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p2.1)\.
- Gaoet al\.\(2023\)L\. Gao, A\. Madaan, S\. Zhou, U\. Alon, P\. Liu, Y\. Yang, J\. Callan, and G\. NeubigPAL: Program\-aided Language Models\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 10764–10799\.External Links:[Link](https://proceedings.mlr.press/v202/gao23f.html)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.
- Gouet al\.\(2024\)Z\. Gou, Z\. Shao, Y\. Gong, Y\. Shen, Y\. Yang, N\. Duan, and W\. ChenCRITIC: Large Language Models Can Self\-Correct with Tool\-Interactive Critiquing\.InInternational Conference on Learning Representations,pp\. 57734–57811\.External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2024/hash/fef126561bbf9d4467dbb8d27334b8fe-Abstract-Conference.html)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.
- Haoet al\.\(2023\)S\. Hao, T\. Liu, Z\. Wang, and Z\. HuToolkenGPT: augmenting frozen language models with massive tools via tool embeddings\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2023/hash/8fd1a81c882cd45f64958da6284f4a3f-Abstract-Conference.html)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px3.p1.1)\.
- Jainet al\.\(2025\)N\. Jain, K\. Han, A\. Gu, W\. Li, F\. Yan, T\. Zhang, S\. Wang, A\. Solar\-Lezama, K\. Sen, and I\. StoicaLiveCodeBench: holistic and contamination free evaluation of large language models for code\.InInternational Conference on Learning Representations,pp\. 58791–58831\.External Links:[Link](https://openreview.net/forum?id=chfJJYC3iL)Cited by:[§B\.2](https://arxiv.org/html/2609.36116#A2.SS2.SSS0.Px3),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Jimenezet al\.\(2024\)C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. NarasimhanSWE\-bench: can language models resolve real\-world GitHub issues?\.InInternational Conference on Learning Representations,pp\. 54107–54157\.External Links:[Link](https://openreview.net/forum?id=VTF8yNQM66)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px6.p1.1),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Kahneman \(2011\)D\. KahnemanThinking, fast and slow\.Farrar, Straus and Giroux\.Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p1.1),[§1](https://arxiv.org/html/2609.36116#S1.p3.1)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with pagedattention\.InProceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23\-26, 2023,J\. Flinn, M\. I\. Seltzer, P\. Druschel, A\. Kaufmann, and J\. Mace \(Eds\.\),pp\. 611–626\.External Links:[Link](https://doi.org/10.1145/3600006.3613165),[Document](https://dx.doi.org/10.1145/3600006.3613165)Cited by:[2nd item](https://arxiv.org/html/2609.36116#S1.I1.i2.p1.1),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p3.1)\.
- Lee \(2026\)T\. LeeSemIf: semantic ifs from open models\.Note:GitHub repositoryRevision 23cf1f39, accessed September 26, 2026External Links:[Link](https://github.com/TheoLeeCJ/SemIf-OpenJev)Cited by:[§2\.1](https://arxiv.org/html/2609.36116#S2.SS1.SSS0.Px3.p2.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2024\)X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang, S\. Zhang, X\. Deng, A\. Zeng, Z\. Du, C\. Zhang, S\. Shen, T\. Zhang, Y\. Su, H\. Sun, M\. Huang, Y\. Dong, and J\. TangAgentBench: evaluating llms as agents\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=zAdUB0aCTQ)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.
- Meta \(2024\)MetaLlama\-3\.2\-3B\-Instruct\.Note:Model cardExternal Links:[Link](https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct)Cited by:[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p3.1)\.
- Nandakishor \(2026\)NandakishorLaya: non\-autoregressive system 1 decision engine\.Note:GitHub repositoryRevision 4066d5d5, accessed September 26, 2026External Links:[Link](https://github.com/NandhaKishorM/laya)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p2.1),[§2\.1](https://arxiv.org/html/2609.36116#S2.SS1.SSS0.Px3.p2.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px2.p1.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.5 model\.Note:OpenAI API documentationAccessed September 25, 2026External Links:[Link](https://developers.openai.com/api/docs/models/gpt-5.5)Cited by:[§A\.2](https://arxiv.org/html/2609.36116#A1.SS2.SSS0.Px1.p1.1),[§A\.2](https://arxiv.org/html/2609.36116#A1.SS2.SSS0.Px1.p2.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.5: towards native multimodal agents\.Note:Qwen blogExternal Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p3.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906-Abstract-Conference.html)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. K\. Li, Y\. Wu, and D\. GuoDeepSeekMath: pushing the limits of mathematical reasoning in open language models\.Note:arXiv preprint arXiv:2402\.03300External Links:[Link](https://arxiv.org/abs/2402.03300)Cited by:[§A\.3](https://arxiv.org/html/2609.36116#A1.SS3.SSS0.Px1.p1.3),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p2.1)\.
- Shenget al\.\(2025\)G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu, W\. Zhang, R\. Zhang, Y\. Peng, H\. Lin, and C\. WuHybridFlow: A flexible and efficient RLHF framework\.InProceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 \- 3 April 2025,pp\. 1279–1297\.External Links:[Link](https://doi.org/10.1145/3689031.3696075),[Document](https://dx.doi.org/10.1145/3689031.3696075)Cited by:[2nd item](https://arxiv.org/html/2609.36116#S1.I1.i2.p1.1),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p3.1)\.
- Shridharet al\.\(2021\)M\. Shridhar, X\. Yuan, M\. Côté, Y\. Bisk, A\. Trischler, and M\. HausknechtALFWorld: aligning text and embodied environments for interactive learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0IOX0YcCdTn)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px3),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Wanget al\.\(2024\)Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang, T\. Li, M\. Ku, K\. Wang, A\. Zhuang, R\. Fan, X\. Yue, and W\. ChenMMLU\-pro: A more robust and challenging multi\-task language understanding benchmark\.InAdvances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 \- 15, 2024,A\. Globersons, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. M\. Tomczak, and C\. Zhang \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2024/hash/ad236edc564f3e3156e1b2feafb99a24-Abstract-Datasets/_and/_Benchmarks/_Track.html)Cited by:[§B\.2](https://arxiv.org/html/2609.36116#A2.SS2.SSS0.Px1),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Yaoet al\.\(2022\)S\. Yao, H\. Chen, J\. Yang, and K\. NarasimhanWebShop: towards scalable real\-world web interaction with grounded language agents\.InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 \- December 9, 2022,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper/_files/paper/2022/hash/82ad13ec01f9fe44c01cb91814fd7b8c-Abstract-Conference.html)Cited by:[§B\.1](https://arxiv.org/html/2609.36116#A2.SS1.SSS0.Px4),[§5\.1](https://arxiv.org/html/2609.36116#S5.SS1.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\-5, 2023,External Links:[Link](https://openreview.net/forum?id=WE\_vluYUL-X)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p1.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.
- Yueet al\.\(2025\)Z\. Yue, W\. Wang, Y\. Zhan, J\. Li, D\. Dahlmeier, and F\. D\. JohanssonExpanding the Action Space of LLMs to Reason Beyond Language\.Note:arXiv preprint arXiv:2510\.07581External Links:[Link](https://arxiv.org/abs/2510.07581)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p2.1),[§1](https://arxiv.org/html/2609.36116#S1.p3.1),[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px3.p1.1),[§5\.2](https://arxiv.org/html/2609.36116#S5.SS2.p3.1)\.
- Zhaoet al\.\(2026\)K\. Zhao, B\. Zhu, J\. Zhou, X\. Zhu, Z\. Yue, and H\. ZhangThinking with Images as Continuous Actions: Numerical Visual Chain\-of\-Thought\.Note:arXiv preprint arXiv:2602\.23959External Links:[Link](https://arxiv.org/abs/2602.23959)Cited by:[§1](https://arxiv.org/html/2609.36116#S1.p3.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. NeubigWebArena: A realistic web environment for building autonomous agents\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=oKn9c6ytLx)Cited by:[§4](https://arxiv.org/html/2609.36116#S4.SS0.SSS0.Px1.p1.1)\.

## Appendix AImplementation Details

### A\.1Action Encoding

The action encodergguses a pretrained transformer backbone separate fromff, followed by a shared projector\. Alignment trains only the projector\. Agentic RL updates both the encoder backbone and projector, including when the LLM is frozen\. The projector pools the encoder’s token representations into an action representation with the same dimensionality as𝐪t\\mathbf\{q\}\_\{t\}\.

Each candidate is encoded with the environment’s action descriptions and its own definition\. These descriptions supply the action context indesc⁡\(e\)\\mathrm\{desc\}\(e\)\. The permissible set𝒜t\\mathcal\{A\}\_\{t\}determines which candidates participate in the action distribution at a decision step\. The encoder input excludes the current task context and reasoning inhth\_\{t\}\. Action representations can be reused while their input descriptions and encoder parameters remain unchanged, and are recomputed when either changes\.

With these textual inputs fixed, permuting the rows of𝐙t\\mathbf\{Z\}\_\{t\}permutes the corresponding action probabilities\. This property concerns the order of candidates in the scoring matrix\. Reordering the action descriptions within the encoder or LLM input can change their representations and is not covered by this property\.

#### Tool Calls\.

Each environment interaction involves a single typed decision over𝒜t\\mathcal\{A\}\_\{t\}\. The language modeling head generates reasoning and any required arguments\. The selected action’s fixed textual form is appended tohth\_\{t\}, even when it spans multiple tokens\. Once complete, the call is executed and the resulting observation is appended to the history\.

#### Projector Architectures\.

Mean pooling uniformly averages encoder token representations and adds no trainable pooling parameters; this variant skips alignment\. MLP\-weighted pooling uses an MLP to score each token, then computes a softmax\-weighted average of the representations\. Attention pooling uses a learnable query to attend to projected keys and values, followed by an output projection\. Padding tokens are masked in all three variants\.

#### Action Encoder Input Formats\.

We use structured MCP\-style action descriptions as the default input to the action encoder\. Table[7](https://arxiv.org/html/2609.36116#A1.T7)compares this format with natural\-language descriptions for Qwen3\.5\-4B with Dyad\-GRPO trained on CodeGym\. MCP\-style inputs yield absolute gains of 2\.87% on ALFWorld and 4\.39% on WebShop\. The consistent advantage across these two evaluation environments supports structured descriptions as the encoder’s default input format beyond the environment used for agentic RL\.

Natural\-language descriptions are substantially shorter in both environments, with a particularly large length difference on ALFWorld\. The performance advantage of MCP\-style inputs therefore comes with additional description tokens\. This comparison favors the structured input in terms of task success, but does not isolate formatting from differences in description length and content\.

Table 7:Action encoder input formats\.Qwen3\.5\-4B with Dyad\-GRPO is trained on CodeGym and evaluated on ALFWorld and WebShop\. Length denotes the mean number of description tokens\.

### A\.2Action Encoder Initialization

#### Data Construction\.

We construct synthetic examples for action selection across five domains: spatial manipulation, state change, information, workflow and communication, and transformation\. Each domain contains ten training actions and four additional actions excluded from training, giving 50 training actions and 20 held\-out actions\. For each action, we obtain a structured definition specifying its name, purpose, and arguments, and a corresponding natural\-language description\. We construct synthetic action\-selection examples across five domains, with 50 training actions and 20 held\-out actions\. Each example contains 4–10 permissible actions𝒜i\\mathcal\{A\}\_\{i\}from one domain and a targetai⋆∈𝒜ia\_\{i\}^\{\\star\}\\in\\mathcal\{A\}\_\{i\}\. GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.36116#bib.bib19)\)generates task context and reasoninghih\_\{i\}conditioned on these preselected actions and the target\. Each case has paired MCP\-style and natural\-language action descriptions, sharing its context, reasoning, action order, and target\. Table[8](https://arxiv.org/html/2609.36116#A1.T8)reports the dataset sizes\.

For each example, we first fix the permissible actions𝒜t\\mathcal\{A\}\_\{t\}and target actiona∈𝒜ta\\in\\mathcal\{A\}\_\{t\}\. Each set contains 4–10 actions from one domain\. We balance set sizes and target frequencies within each domain, approximately balance the use of other candidates, and shuffle their order\. GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.36116#bib.bib19)\)then generates task context and reasoning conditioned on these actions and the target\. The action descriptions providedesc⁡\(e\)\\mathrm\{desc\}\(e\), while the generated context and reasoning enterhth\_\{t\}\. These examples supervise action selection without environment rollouts\.

#### Inputs and Supervision\.

The LLM input combines the action descriptions, task context, and reasoning with format instructions, an example call, and a prefix marking the action\-selection position\. The example call uses an action sampled uniformly from𝒜t\\mathcal\{A\}\_\{t\}, without conditioning on the target\. Each input toggcontains the same action descriptions, the definition of the action being encoded, and an instruction identifying that action\. It excludes the task context and reasoning and does not identify the supervised target\. All candidates are encoded as𝐳a\\mathbf\{z\}\_\{a\}and scored against𝐪t\\mathbf\{q\}\_\{t\}at the action\-selection position\. We minimize Equation[4](https://arxiv.org/html/2609.36116#S3.E4), averaged over training examples, with a softmax temperature of one\. Only the projector is updated, and no next\-token prediction loss is applied to the supplied text\.

#### Input Formats and Data Splits\.

Table[8](https://arxiv.org/html/2609.36116#A1.T8)summarizes the dataset sizes\. Each case has two versions, using structured MCP\-style tool definitions or natural\-language action descriptions\. They share the task context, reasoning, candidate order, and target action, and remain in the same split\.

Table 8:Alignment dataset\.Each case has paired MCP\-style and natural\-language records\.Validation and test are split by case, stratified by domain, target action, and number of permissible actions, using seed 42\. Their targets come from the same 20 actions excluded from all training candidate sets; other candidates may include both training and held\-out actions\. Thus, validation and test contain distinct cases within a shared action pool\. We select the projector checkpoint by validation negative log\-likelihood\. Because this evaluation pool had been used for model selection before subdivision, the test split provides a diagnostic rather than a historically independent evaluation\.

#### Supervision and Data Splits\.

Only the projector is trained using Equation[4](https://arxiv.org/html/2609.36116#S3.E4); the generated text is supplied as input, with supervision applied to the target action\. Validation and test targets come from actions absent from all training candidate sets\. The two splits contain distinct cases from the same held\-out action pool, with paired formats kept together\. We select checkpoints by validation cross\-entropy\. The evaluation pool was used in earlier model selection, so its test split is not a historically independent holdout\.

#### Data Checks\.

We check action\-definition structure, candidate coverage, paired\-format consistency, and split isolation\. Near\-duplicate contexts are identified by word\-set overlap within each domain and generation pool and regenerated\. These checks do not establish semantic equivalence of every description pair or disjointness from downstream benchmark actions\.

### A\.3Agentic Reinforcement Learning

#### Sampled Decisions and Loss Aggregation\.

Policy\-gradient terms are computed for sampled vocabulary tokens and typed actions\. The fixed text appended after an action selection records that decision and contributes no additional policy\-gradient terms\. Environment observations are also excluded\. Letiiindex a sampled decision andpi​\(θ\)p\_\{i\}\(\\theta\)denote its probability underπθLM\\pi\_\{\\theta\}^\{\\mathrm\{LM\}\}orπθact\\pi\_\{\\theta\}^\{\\mathrm\{act\}\}, evaluated at the recorded history and, for a typed action, its permissible action set\. The rollout parametersθold\\theta\_\{\\mathrm\{old\}\}and the advantageA^i\\widehat\{A\}\_\{i\}computed from environment rewards are held fixed during each update\. The standard clipped surrogate for either decision type is

ρi​\(θ\)\\displaystyle\\rho\_\{i\}\(\\theta\)=pi​\(θ\)pi​\(θold\),\\displaystyle=\\frac\{p\_\{i\}\(\\theta\)\}\{p\_\{i\}\(\\theta\_\{\\mathrm\{old\}\}\)\},\(5\)ℓi​\(θ\)\\displaystyle\\ell\_\{i\}\(\\theta\)=−min⁡\(ρi​\(θ\)​A^i,clip⁡\(ρi​\(θ\),1−ϵlow,1\+ϵhigh\)​A^i\)\.\\displaystyle=\-\\min\\\!\\left\(\\rho\_\{i\}\(\\theta\)\\widehat\{A\}\_\{i\},\\,\\operatorname\{clip\}\\\!\\left\(\\rho\_\{i\}\(\\theta\),1\-\\epsilon\_\{\\mathrm\{low\}\},1\+\\epsilon\_\{\\mathrm\{high\}\}\\right\)\\widehat\{A\}\_\{i\}\\right\)\.The clipping bounds follow Table[10](https://arxiv.org/html/2609.36116#A2.T10)\. Rollout and training probabilities use the same sampling temperature and permissible actions\. GRPO and GiGPO differ in how environment rewards are used to estimate advantages\([Shao et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib14);[Feng et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib15)\)\.

For aggregation groupbb, letℐbLM\\mathcal\{I\}\_\{b\}^\{\\mathrm\{LM\}\}andℐbact\\mathcal\{I\}\_\{b\}^\{\\mathrm\{act\}\}index sampled vocabulary tokens and typed actions, respectively\. LetNbN\_\{b\}be the number of response tokens in that group andcbc\_\{b\}the weight assigned to that group\. We aggregate the sampled\-decision losses into language and action contributions:

ℒLM\\displaystyle\\mathcal\{L\}^\{\\mathrm\{LM\}\}=∑bcbNb​∑i∈ℐbLMℓi​\(θ\),\\displaystyle=\\sum\_\{b\}\\frac\{c\_\{b\}\}\{N\_\{b\}\}\\sum\_\{i\\in\\mathcal\{I\}\_\{b\}^\{\\mathrm\{LM\}\}\}\\ell\_\{i\}\(\\theta\),\(6\)ℒact\\displaystyle\\mathcal\{L\}^\{\\mathrm\{act\}\}=∑bcbNb​∑i∈ℐbactℓi​\(θ\)\.\\displaystyle=\\sum\_\{b\}\\frac\{c\_\{b\}\}\{N\_\{b\}\}\\sum\_\{i\\in\\mathcal\{I\}\_\{b\}^\{\\mathrm\{act\}\}\}\\ell\_\{i\}\(\\theta\)\.Both contributions share the same denominator and group weights\. The response\-token countNbN\_\{b\}includes fixed action text but excludes padding and environment observations\. The action contribution is not separately normalized by the number of actions\.

In joint optimization, the policy\-gradient update toffand𝐖LM\\mathbf\{W\}\_\{\\mathrm\{LM\}\}usesℒLM\+ℒact\\mathcal\{L\}^\{\\mathrm\{LM\}\}\+\\mathcal\{L\}^\{\\mathrm\{act\}\}, with action representations𝐳a\\mathbf\{z\}\_\{a\}treated as constants\. The update toggusesℒact\\mathcal\{L\}^\{\\mathrm\{act\}\}, with state representations𝐪t\\mathbf\{q\}\_\{t\}treated as constants, and trains both its backbone and projector\. The action contribution has the same numerical value in both updates, but its gradients are taken with respect to different components\. In frozen\-LLM adaptation, only the update toggis applied\.

## Appendix BExperimental Details

### B\.1Agentic Environments

Table[9](https://arxiv.org/html/2609.36116#A2.T9)summarizes task counts and the average number of available tool or operation types per task\. For ALFWorld and WebShop, these counts exclude state\-dependent arguments and therefore do not measure the number of permissible choices at each decision step\. The substantially larger toolsets in DIVE contrast with CodeGym’s smaller sets of atomic functions\. SWE exposes only one Bash interface, but its command arguments support repository inspection, editing, and testing\. These statistics characterize the breadth of exposed interfaces; the number of interfaces alone does not determine the complexity of choosing and executing an action\.

Table 9:Environment statistics\.Environments cover training and transfer evaluation\.#### DIVE\([Chen et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib5)\)\.

The dataset contains diverse tool\-use tasks synthesized from execution traces of real\-world tools\. Each task pairs a query with an available toolset and a verifiable reference answer, requiring agents to gather and process information through multiple tool calls\.

#### CodeGym\([Du et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib6)\)\.

Coding problems are converted into interactive tool\-use environments by exposing atomic functions from their solutions as callable tools\. Agents solve tasks through sequential tool interactions, while the underlying code and test cases support automatic verification\. The counts in Table[9](https://arxiv.org/html/2609.36116#A2.T9)use the English\-task reproduction released as[VanishD/CodeGym](https://huggingface.co/datasets/VanishD/CodeGym/tree/85286359a342f7a288aea74273772b69b9b784c2)\.

#### ALFWorld\([Shridhar et al\., 2021](https://arxiv.org/html/2609.36116#bib.bib7)\)\.

This benchmark provides text\-based environments for household tasks, such as finding, placing, cleaning, heating, and cooling objects\. Agents navigate and manipulate objects through textual commands, using observations from the environment to track progress toward the task goal\.

#### WebShop\([Yao et al\., 2022](https://arxiv.org/html/2609.36116#bib.bib8)\)\.

In this simulated e\-commerce environment, agents follow natural\-language shopping instructions\. They search for products, inspect product details, select options, and make a purchase that satisfies the requested attributes and constraints\.

#### τ2\\tau^\{2\}\-bench\([Barres et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib9)\)\.

This benchmark evaluates conversational agents that use tools and communicate with users to complete tasks\. Its dual\-control telecom environment allows both the agent and the user to act on a shared environment state, testing coordination alongside tool use and reasoning\. The statistics use the Retail, Airline, and Telecom base splits at[repository revisiona2c0247](https://github.com/sierra-research/tau2-bench/tree/a2c024725189473d2d7cea3a5cfdbcc67478e41f), which includes subsequentτ3\\tau^\{3\}updates\.

#### SWE\-bench Verified\.

This human\-filtered subset of SWE\-bench\([Jimenez et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib10)\)evaluates the resolution of real\-world GitHub issues\.111[https://www\.swebench\.com/](https://www.swebench.com/)Given an issue description and a code repository, agents inspect and modify the code to produce a patch, whose correctness is assessed using the benchmark’s tests\. We evaluate a 100\-task subset of the 500 verified instances\. We use the Bash interface of mini\-swe\-agent\.222[https://github\.com/SWE\-agent/mini\-swe\-agent](https://github.com/SWE-agent/mini-swe-agent)The evaluated model directly generates shell commands for repository inspection, editing, and testing\. A GiGPO\-style prompt provides recent observation–action pairs and the current execution feedback at each step\. The same Bash interface submits the patch and terminates the episode\.

The cross\-environment evaluations use agents trained with RL on DIVE or CodeGym without further RL on ALFWorld, WebShop,τ2\\tau^\{2\}\-bench, or SWE\-bench Verified\. This separation refers to the RL training environments; the alignment\-data split is described in Appendix[A\.2](https://arxiv.org/html/2609.36116#A1.SS2)\.

### B\.2General Capability Benchmarks

#### MMLU\-Pro\([Wang et al\., 2024](https://arxiv.org/html/2609.36116#bib.bib11)\)\.

The benchmark extends MMLU with more challenging, reasoning\-focused multiple\-choice questions across academic subjects and up to ten answer choices per question\. We evaluate on the full test set of 12,032 questions to assess general knowledge and reasoning after agentic post\-training\.

#### HMMT February 2026\.

We use the[MathArena release](https://huggingface.co/datasets/MathArena/hmmt_feb_2026)\([Dekoninck et al\., 2026](https://arxiv.org/html/2609.36116#bib.bib13)\)of problems from the February 2026 Harvard–MIT Mathematics Tournament\. These competition problems assess mathematical reasoning across algebra, geometry, number theory, and combinatorics\. We evaluate on all 33 problems and report avg@4, averaging correctness over four generated answers per problem and then across problems\.

#### LiveCodeBench v6\([Jain et al\., 2025](https://arxiv.org/html/2609.36116#bib.bib12)\)\.

The benchmark evaluates coding capabilities using programming contest problems collected over time from LeetCode, AtCoder, and Codeforces\. We evaluate on the full version 6 test set of 1,055 problems, released from May 2023 to April 2025, with generated programs evaluated against the provided test cases\.

### B\.3JevBench

#### Tasks and Evaluation Subset\.

JevBench evaluates typed decision\-making given a state, instructions, decision criteria, and a finite set of candidate answers\.333[https://github\.com/fstandhartinger/jevbench](https://github.com/fstandhartinger/jevbench)Its tasks include categorical choices, yes/no judgments, and ordinal ratings, covering problems such as routing, policy compliance, information extraction, and numerical reasoning\. Each item evaluates a standalone decision without an environment rollout\. We use the public subset of 231 items: 48 Easy, 72 Standard, and 111 Hard\.

#### Metrics and Latency Source\.

Table[5](https://arxiv.org/html/2609.36116#S5.T5)reports accuracy within each difficulty tier; Overall is the average weighted by the number of items per tier\. Mean output tokens include reasoning and final answers, and mean latency measures time until the complete response is returned\. The Jev latency entry is recomputed from the benchmark’s[published per\-item API timings](https://github.com/fstandhartinger/jevbench/blob/1bcc55eb6c8cffde2306b3db03ede39b61c6152a/results/v1.2/jevbench-v1.2-per-task.json)for the same public subset\. These timings come from a separate evaluation and include network overhead\. They provide an API latency reference; the different deployments do not establish a matched\-hardware speed comparison\. A corresponding Jev output\-token measurement for this subset is unavailable\.

### B\.4Hyperparameters

Table[10](https://arxiv.org/html/2609.36116#A2.T10)summarizes the hyperparameters for agentic RL in the joint optimization setting\. The response limit bounds each turn, while the interaction limit bounds the trajectory\. Shorter individual responses can therefore coexist with longer interaction sequences, as observed on CodeGym in Appendix[C\.2](https://arxiv.org/html/2609.36116#A3.SS2)\. Since these budgets and rollout batch sizes differ across environments, we interpret training curves through comparisons within each environment and model scale\.

Table 10:Agentic RL hyperparameters\.Settings correspond to joint post\-training\.Notes\.Prompt and response limits apply per interaction; responses exclude tool outputs\. Mini\-batches count interaction steps\. Thinking mode is disabled\.

### B\.5Prompt Templates

Figures[7](https://arxiv.org/html/2609.36116#A2.F7)–[12](https://arxiv.org/html/2609.36116#A2.F12)summarize the environment prompt templates\. Named placeholders represent the task, available actions, observations, and interaction history\. The templates request explicit reasoning; disabling the model’s native thinking mode does not remove these instructions\.

For templates with explicit history fields, these contain recent observation–action pairs, while the current observation supplies the latest feedback\. When no history is included, the corresponding history text is omitted; DIVE, ALFWorld, and WebShop also omit step\-count text\. The response instructions remain unchanged\. WebShop additionally falls back to its no\-history template when a prompt with history exceeds 13,000 characters\. SWE\-bench Verified follows the same GiGPO\-style per\-step structure, with fixed repository and submission instructions in the system message\. The user message contains the task, recent history, current observation, and available tool\. We retain our<think\>and<action\>response format while using mini\-swe\-agent’s Bash interface\.

DIVE prompt templateUSERYou are an expert agent solving a task using the available tools\.Your task is to:\{task\_description\}Prior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your available tools are:\{available\_actions\}Use these tool schemas to construct calls; the interface specifies the native tool\-call syntax inside<action\>\.Reason about the next step within<think\></think\>tags\.Then enclose your action within<action\></action\>tags\.Inside<action\>, call the appropriate tools using the native tool\-call format and their declared arguments, then close</action\>and wait for their results\.When ready, put your final answer inside<action\></action\>instead of tool calls, in the format requested by the query\.An action without tool calls ends the task; do not output reasoning alone or text outside these two blocks\.Figure 7:Prompt template for DIVE\.It specifies how the agent calls tools and submits its final answer\.CodeGym prompt templateSYSTEM\{original\_system\_messages\}USERYou are an expert agent operating in an interactive function\-calling environment\.Your task is to:\{task\_description\}Prior to this step, you have already taken\{step\_count\}step\(s\)\.Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your available functions are: \[\{available\_actions\}\]\.Reason about the next step within<think\></think\>tags\.Then output exactly one function call using its declared arguments:<\|FunctionCallBegin\|\>\[\{"name": "<function name\>", "parameters": \{"<argument name\>": "<value\>"\}\}\]<\|FunctionCallEnd\|\>Use \{\} for no arguments; do not write executable code\.Wait for feedback after each call and follow the task’s submission rules\.Figure 8:Prompt template for CodeGym\.The agent issues one function call at a time and waits for feedback before continuing\.ALFWorld prompt templateUSERYou are an expert agent operating in an interactive household environment\. Your task is to:\{task\_description\}Prior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your admissible actions of the current situation are: \[\{admissible\_actions\}\]\.Now it’s your turn to take an action\.You should first reason step\-by\-step about the current situation\. This reasoning process MUST be enclosed within<think\></think\>tags\.Once you’ve finished your reasoning, you should choose an admissible action for current step and present it within<action\></action\>tags\.Figure 9:Prompt template for ALFWorld\.The agent selects an admissible action based on the current observation and interaction history\.WebShop prompt templateUSERYou are an expert autonomous agent operating in an online shopping environment\.Your task is to:\{task\_description\}\.Prior to this step, you have already taken\{step\_count\}step\(s\)\. Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}\.Your admissible actions of the current situation are:\[\{available\_actions\}\]\.Now it’s your turn to take one action for the current step\.You should first reason step\-by\-step about the current situation, then think carefully which admissible action best advances the shopping goal\. This reasoning process MUST be enclosed within<think\></think\>tags\.Once you’ve finished your reasoning, you should choose an admissible action for current step and present it within<action\></action\>tags\.Figure 10:Prompt template for WebShop\.The agent uses the shopping instruction and current observation to select its next action\.τ2\\tau^\{2\}\-bench prompt templateSYSTEM\{domain\_policy\}\#Available tools\{tool\_schemas\}\# InstructionFollow the policy above and the declared tool schemas\.Ask the customer for missing information or required confirmation; do not invent facts or tool results\.USERYou are an expert agent helping a customer in an interactive service environment\.Your task is to:\{task\_description\}Prior to this step, you have already taken\{step\_count\}step\(s\)\.Below are the most recent\{history\_length\}observations and the corresponding actions you took:\{action\_history\}You are now at step\{current\_step\}and your current observation is:\{current\_observation\}Your available actions are: \[\{available\_actions\}\]\.Reason about the next step within<think\></think\>tags\.Then output exactly one action as valid JSON enclosed within<action\></action\>tags:<action\>\{"name": "<action name\>", "arguments": \{"<argument name\>": "<value\>"\}\}</action\>To message the customer, use name="respond" and arguments=\{"content": "<message\>"\}; this does not end the conversation\.Stop after</action\>and wait for feedback; output no text outside the<think\>and<action\>blocks or Markdown fences\.Figure 11:Prompt template forτ2\\tau^\{2\}\-bench\.The agent follows the domain policy when using tools and communicating with the customer\.SWE\-bench Verified prompt templateSYSTEMFix the issue in /testbed and verify the fix\. Each command runs in a fresh shell\. Do not modify tests or configuration files, or commit changes\.Save a source\-onlygit difftopatch\.txt, usinggit add \-Nfor new files\. Exclude reproduction scripts, helpers, and binaries\. Inspect the patch in a separate call, then submit in a final bash call:echo COMPLETE\_TASK\_AND\_SUBMIT\_FINAL\_OUTPUT && cat patch\.txtExit status 0 completes submission\.USERTask:\{task\_description\}Recent observation–action history:\{action\_history\}Current step:\{current\_step\}Current observation:\{current\_observation\}Available tools: \[\{available\_actions\}\]\.Put reasoning inside<think\></think\>, followed by one<action\></action\>block with native bash calls\. Include at least one call; group only independent commands\. Output nothing outside these blocks; then stop and wait for results\.Figure 12:Prompt template for SWE\-bench Verified\.GiGPO\-style context with mini\-swe\-agent’s Bash interface\.

## Appendix CAdditional Results and Analyses

### C\.1Task\-Specific Post\-training on ALFWorld and WebShop

We train separate agents on ALFWorld and WebShop, jointly updating the LLM and action encoder for Dyad\. Table[11](https://arxiv.org/html/2609.36116#A3.T11)reports general capabilities for the same agents evaluated in Table[2](https://arxiv.org/html/2609.36116#S5.T2)\. These runs are separate from the DIVE\- and CodeGym\-trained agents used for cross\-environment evaluation\.

Table 11:General capability retention\.Models are post\-trained on ALFWorld or WebShop\.Comparison with RL Baselines\.Both Dyad variants outperform their corresponding RL baselines in all 24 comparisons across the two model scales, two training environments, and three capability benchmarks\. The advantage extends to general knowledge, mathematical reasoning, and coding\. Together with the consistently higher in\-domain scores in Table[2](https://arxiv.org/html/2609.36116#S5.T2), these results show that stronger task adaptation with Dyad can accompany better general capabilities than conventional RL post\-training\.

Retention Relative to the Original Models\.GRPO and GiGPO score below the original models in every reported capability evaluation\. Dyad reduces these losses, but recovery varies with the training environment and capability\. At 9B, both Dyad variants exceed the original model’s coding score after training on either environment\. Mathematical reasoning is more dependent on the training source: both 9B Dyad variants exceed the original HMMT score after ALFWorld training, whereas both remain below it after WebShop training\. At 4B, several scores also remain below their original levels despite improving over the RL baselines\. Thus, the paired advantage supports better capability retention during task\-specific joint training, although it does not eliminate degradation in every setting\.

### C\.2Additional Training Dynamics

CodeGym\.Figure[13](https://arxiv.org/html/2609.36116#A3.F13)extends the 9B results in Section[5\.2](https://arxiv.org/html/2609.36116#S5.SS2)to Qwen3\.5\-4B and Qwen3\.5\-27B\. At both scales, Dyad\-GRPO and Dyad\-GiGPO achieve higher final validation rewards than their corresponding baselines, with generally shorter responses and more environment interactions\. Despite greater fluctuations at 4B, the paired advantage remains visible late in training\.

Qwen3\.5\-4B

Qwen3\.5\-27B

Figure 13:Additional training dynamics on CodeGym\.Validation reward, response length, and environment interactions for Qwen3\.5\-4B \(top\) and Qwen3\.5\-27B \(bottom\)\.DIVE\.Figure[14](https://arxiv.org/html/2609.36116#A3.F14)reports training dynamics for Qwen3\.5\-4B, Qwen3\.5\-9B, and Qwen3\.5\-27B\. Both Dyad variants attain higher final validation rewards and shorter responses than their corresponding baselines\. Their interaction counts rise initially but decline later, finishing below the baselines\. The reward advantage is therefore consistent across the two environments, while the interaction pattern differs: more interactions on CodeGym and fewer on DIVE\.

Qwen3\.5\-4B

Qwen3\.5\-9B

Qwen3\.5\-27B

Figure 14:Training dynamics on DIVE\.Models are jointly post\-trained on DIVE\.

Similar Articles

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models

arXiv cs.LG

This paper introduces LDM-v0, a large decision model trained offline on trajectories from thousands of diverse reinforcement learning environments, demonstrating that a single transformer policy can match the performance of task-specific policies across robotics, autonomous driving, inventory management, cybersecurity, trading, and video games.

Language Acquisition Device in Large Language Models

arXiv cs.CL

This paper proposes LAD-inspired pre-pretraining using a formal language called MP-Struct that encodes natural-language-like structures. It shows that this approach improves token efficiency and imparts human-like resistance to structurally implausible languages, challenging prior hypotheses about effective pre-pretraining languages.