TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
Summary
TRACE presents a simulation-based method to train reasoning agents for causal exploration using synthesized rewards, improving reinforcement learning in diagnostic tasks. Experiments show it outperforms prompted baselines in a digital-advertising diagnostic environment.
View Cached Full Text
Cached at: 09/11/26, 08:45 AM
# TRACE: Training Reasoning Agents for Causal Exploration with Synthesized Rewards
Source: [https://arxiv.org/html/2609.10315](https://arxiv.org/html/2609.10315)
###### Abstract
Reinforcement learning with verifiable rewards \(RLVR\) has advanced language\-model reasoning in domains such as mathematics and code, where objective answers are inexpensive to check\. Diagnostic reasoning over complex data lacks this advantage: establishing the true cause of an anomaly often requires costly expert investigation and may remain ambiguous after the fact\. We ask whether this asymmetry of verification can instead be engineered\. We sample an intervention, inject it into a controlled simulator, and generate the observations it would produce\. The hidden intervention provides an oracle label and objective reward, while the agent must still investigate noisy, confounded, and distributed evidence\.
We instantiate this approach inTRACE, a digital\-advertising diagnostic environment with 12 root causes and fine\-grained segment attribution\. Agents investigate each episode using Python and SQL and must identify both the root cause and, when applicable, the affected segment assignment\. On a held\-out 235\-episode test set, the strongest prompted baseline, Claude Opus 5, reaches0\.6860\.686FullAttr@1\. Supervised fine\-tuning raises Qwen3\.5\-35B\-A3B from0\.1590\.159to0\.6370\.637, and subsequent RL with synthesized rewards reaches0\.7570\.757, outperforming all evaluated prompted baselines, including frontier closed\-source models and a prompted Qwen3\.5\-122B\-A10B model\. The resulting policy also uses substantially fewer tool calls than the prompted 35B base\. These results provide evidence that access to a scalable, objective training signal can be a more important constraint than model scale alone\. More broadly, simulation\-based verification can make otherwise ambiguous diagnostic reasoning tasks amenable to scalable reinforcement learning\.
## 1Introduction
Reinforcement learning with verifiable rewards \(RLVR\) has driven recent progress in LLM reasoning\. Work on mathematical reasoning\([Guo et al\., 2025](https://arxiv.org/html/2609.10315#bib.bib1);[Shao et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib2)\)and coding agents\([Wei et al\., 2026](https://arxiv.org/html/2609.10315#bib.bib3);[Luo et al\., 2025b](https://arxiv.org/html/2609.10315#bib.bib4)\)shows that policies optimized against objective verifiers can acquire multi\-step reasoning and tool\-use capabilities\. These domains provide a natural*asymmetry of verification*: checking a solution against a known answer or test suite is considerably easier than producing the solution itself\([Wei, 2025](https://arxiv.org/html/2609.10315#bib.bib15)\)\. RLVR turns this asymmetry into a scalable training signal\.
Many practical reasoning tasks lack such verifiers\. Diagnostic reasoning over complex data is a representative example: determining why a system changed can require expensive expert investigation, and the resulting attribution may remain uncertain even after the fact\. Multiple changes can occur concurrently, observations are noisy, and plausible causes may produce similar signatures\. LLM\-as\-judge approaches\([Zheng et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib10);[Bai et al\., 2022](https://arxiv.org/html/2609.10315#bib.bib11);[Lee et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib12)\)do not eliminate this bottleneck because the target itself remains ambiguous, while learned judges can introduce additional opportunities for reward hacking\([Gao et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib13);[Skalse et al\., 2022](https://arxiv.org/html/2609.10315#bib.bib14)\)\. The central obstacle to applying RLVR in these settings is therefore access to scalable, objective ground truth\.
In this paper, we ask whether the asymmetry of verification can be*engineered*\. Rather than observing data and attempting to determine an unknown cause, we first sample an intervention, inject it into a controlled simulator, and generate the observations that the intervention would produce\. The hidden intervention is retained as an oracle label, making the agent’s final attribution inexpensive and deterministic to verify\. The agent does not observe this label: it must still investigate noisy, confounded, and distributed evidence to recover the cause\. Simulation therefore separates the difficulty of solving a diagnostic task from the difficulty of verifying its answer\.
We instantiate this approach inTRACE\(Training Reasoning Agents for Causal Exploration\), a simulated diagnostic environment for digital advertising\. EachTRACEinstance presents an agent with a campaign\-performance anomaly and a multi\-table database\. The agent investigates through Python and SQL, iterating between hypotheses and evidence, and ultimately attributes the anomaly to one of the predefined cause types while ruling out realistic confounders\.TRACEbuilds a*simulator–oracle–RL pipeline*: the simulator injects a hidden intervention, the oracle verifies that the agent\-visible data contains sufficient evidence to recover it, and RL uses the resulting hidden label as a synthesized reward\. We refer to this setting as*RL with synthesized rewards*, where hidden intervention labels are converted into objective reward signals\.
We evaluate on a held\-out 235\-episode test set enriched for segment\-specific attribution\. The strongest prompted baseline, Claude Opus 5, reaches0\.6860\.686FullAttr@1\. On Qwen3\.5\-35B\-A3B, SFT raises FullAttr@1 from0\.1590\.159to0\.6370\.637, and subsequent RL with synthesized rewards improves it to0\.7570\.757, surpassing every evaluated prompted baseline\. The post\-trained 35B model also substantially outperforms the prompted Qwen3\.5\-122B\-A10B model, which reaches0\.2830\.283\. This result provides evidence that, for this diagnostic setting, the binding constraint is access to an effective post\-training signal—including a scalable, objective reward—rather than model scale alone\. The best trained policy also averages11\.7311\.73tool calls per trajectory, compared with22\.0522\.05for the prompted 35B base\.
Our contributions are threefold\. First, we introduce a simulator–oracle–RL methodology for synthesizing objective reward in domains where natural verifiers are scarce\. The key idea is to generate tasks from controlled interventions, so that the same process that creates a difficult reasoning problem also provides an objective answer\. Second, we instantiate this methodology inTRACE, a tool\-using diagnostic environment with configurable causes, realistic confounders, and multi\-table evidence\.TRACEsupports held\-out evaluation and post\-training without human attribution labels or an LLM judge\. Third, we show that SFT followed by RL withTRACE’s synthesized rewards substantially improves an open\-weight 35B reasoning agent, enabling it to outperform frontier closed\-source models on diagnostic attribution tasks\. We also conduct ablation experiments to study the roles of supervised initialization and RL reward design in producing these gains\. Together, these results show how simulation\-based verification can make otherwise ambiguous diagnostic reasoning tasks amenable to scalable reinforcement learning\.
### 1\.1Related Work
#### RLVR and reasoning training\.
RLVR methods optimize policies against verifiable outcomes, including mathematical answer checkers\([Shao et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib2);[Guo et al\., 2025](https://arxiv.org/html/2609.10315#bib.bib1)\), formal proof validators\([Kim and Yun, 2026](https://arxiv.org/html/2609.10315#bib.bib5)\), and compilation or execution tests for code\([Wei et al\., 2026](https://arxiv.org/html/2609.10315#bib.bib3);[Luo et al\., 2025b](https://arxiv.org/html/2609.10315#bib.bib4)\)\. Recent work has broadened RLVR by improving the underlying policy optimization algorithms\([Shao et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib2);[Yu et al\., 2026](https://arxiv.org/html/2609.10315#bib.bib6);[Liu et al\., 2025](https://arxiv.org/html/2609.10315#bib.bib7);[Zheng et al\., 2025](https://arxiv.org/html/2609.10315#bib.bib8)\)and by extending the training recipe beyond math and code, for example to medical question answering\([Chen et al\., 2025](https://arxiv.org/html/2609.10315#bib.bib9)\)\. However, these advances still assume that verifiable targets already exist\. Our work instead studies how to construct such targets for domains where they are not naturally available\.
#### Synthetic data, simulators, and reward synthesis\.
Synthetic generation has been widely used to create instruction\-following data and reasoning trajectories for supervised fine\-tuning\([Wang et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib16);[Luo et al\., 2025a](https://arxiv.org/html/2609.10315#bib.bib17);[Yu et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib18)\)\. In these settings, the generator primarily supplies examples to imitate\. A related line of work develops simulated and interactive environments for evaluating tool\-using agents in science, web, and machine\-learning tasks\([Wang et al\., 2022](https://arxiv.org/html/2609.10315#bib.bib19);[Zhou et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib20);[Huang et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib21)\)\. These environments use interaction to test whether agents can complete tasks\.TRACEuses simulation differently: it samples a hidden root cause, generates the diagnostic data it would produce, and uses the sampled cause as the oracle label for reward calculation\.
#### Tool\-using data\-analysis and diagnostic agents\.
A growing line of work evaluates whether LLM agents can analyze structured data with external tools\. Text\-to\-SQL benchmarks provide an early testbed for this setting by asking models to translate natural\-language questions into executable database queries\([Yu et al\., 2018](https://arxiv.org/html/2609.10315#bib.bib22);[Zhong et al\., 2017](https://arxiv.org/html/2609.10315#bib.bib23);[Li et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib24);[Chang et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib25)\)\. More recent data\-analysis benchmarks move toward multi\-step tool use over tables, files, and code execution\([Huang et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib21);[Zhang et al\., 2026](https://arxiv.org/html/2609.10315#bib.bib26)\)\. These tasks capture important components of diagnostic investigation, but they generally score query correctness, code outputs, or final reports rather than root\-cause attribution under controlled confounders\.TRACEtargets this gap by constructing diagnostic tasks with simulated interventions, agent\-visible evidence, and verifiable rewards\.
#### Causal reasoning and root\-cause attribution\.
A parallel line of work evaluates whether LLMs can answer causal questions, reason over causal graphs, or recover causal structure from data\([Jin et al\., 2023](https://arxiv.org/html/2609.10315#bib.bib27);[Jin et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib28);[Kiciman et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib29);[Wang, 2024](https://arxiv.org/html/2609.10315#bib.bib30)\)\. These benchmarks probe causal knowledge and formal causal reasoning, often through static questions, symbolic graphs, or fixed datasets\.TRACEstudies a different setting: interactive root\-cause attribution in a simulated environment where the true cause is known by construction\. The agent must gather evidence through SQL and Python, distinguish the sampled cause from realistic confounders, and produce an attribution that can be verified against ground\-truth labels\.
## 2TheTRACEEnvironment
This section describesTRACE, a digital advertising environment that implements the simulator–oracle components of our methodology\.TRACEhas three parts: a stochastic simulator that generates multi\-table advertising data with known injected causes, a task interface through which an agent investigates each generated diagnostic instance, and an oracle verifier that certifies solvability and produces reward labels\. Figure[1](https://arxiv.org/html/2609.10315#S2.F1)summarizes the overall workflow\. We refer to each generated diagnostic instance as an*episode*: the observable history of a single campaign over a fixed time window, together with a hidden injected intervention that defines the ground\-truth cause\.
Figure 1:Overview of theTRACEworkflow\.### 2\.1Preliminaries: Campaign Metrics and Diagnosis
Digital advertising campaigns serve ads to users through online auctions: when a user loads a page with an ad slot, advertisers bid for the slot, and the winning ad is shown as an*impression*\. Under the commonly adopted pay\-per\-click model, the advertiser is charged only when the user*clicks*, and a click that leads to a purchase of the advertised product is recorded as an*order*, or*conversion*\. Campaign performance is usually summarized by metrics such as impressions, click\-through rate \(CTR, clicks/impressions\), conversion rate \(CVR, orders/clicks\), and cost\-per\-click \(CPC, total advertiser spend/clicks\)\.111In general second\-price auction mechanisms, an advertiser’s paid price depends on its competing bids, not the advertiser’s own\. We model the resulting gap between bid and CPC as multiplicative noise on advertiser spend\.These metrics are reported at the campaign level and across*segments*: combinations of categorical attributes such as placement, device, geography or user group\.
Segment\-level reporting is what makes diagnosis both possible and difficult\. The distribution of impressions across segments is shaped by interacting factors, including campaign configuration, competing campaigns, auction dynamics, and user behavior\. These factors also produce ordinary day\-to\-day metric noise\. Telling a real performance shift apart from this noise requires resolving three questions\. First, the cause is ambiguous: one metric movement can have several explanations \(*what*\)\. Second, the cause is localized: a segment\-level effect can be diluted in the campaign aggregate and surface only under the right split \(*where*\)\. Third, the timing is unknown: gradual onsets are easily missed by simple period comparisons \(*when*\)\. Resolving all three at once is the core reasoning patternTRACEis designed to elicit\.
### 2\.2Data Generation Pipeline
For eachTRACEepisode, the sampled intervention determines three hidden labels: the cause type, the affected segment slice, and the onset timing\.
Generation proceeds in three stages: building a campaign population \(Stage 1\), sampling an intervention \(Stage 2\), and simulating the resulting metrics \(Stage 3\)\.
#### Stage 1: Campaign population\.
We first construct a static universe of campaigns\. Each is assigned a vertical category \(e\.g\. electronics, fashion\), a bidding strategy \(e\.g\. search\-heavy, acquisition\), a baseline daily budget, and a baseline impression volume\. A campaign is active in a sampled subset of its possible segments\. Per\-campaign variation in click\-through and conversion rates is captured by a quality multiplier\.
#### Stage 2: Intervention sampling\.
Each episode spans a contiguous window split into two equal parts: a*baseline period*, in which the campaign runs normally, and an*intervention period*, in which a sampled cause is active\. We draw four components: the*injected cause*, one of the twelve predefined cause types \(§[2\.3](https://arxiv.org/html/2609.10315#S2.SS3)\); the*driver slice*, which specifies the affected segments using one or two attributes, such asplacement=TOP\_OF\_SEARCHorgeo=US∧\\wedgedevice=mobile; the*signal strength*, which sets the effect magnitude; and the*onset profile*, which determines whether the effect appears immediately, after a delay, or gradually over time\.
#### Stage 3: Metric simulation\.
The simulator renders daily metrics in three steps\. First, it draws the campaign’s total daily impressions from a baseline volume, a seasonality factor, and any volume change induced by active causes\. Second, it distributes these impressions across segments\. The base distribution reflects the bidding strategy: a search\-heavy campaign, for instance, weights search placements more heavily\. Active causes then reweight this distribution, shifting impressions between segments\. Third, it computes CTR, CVR, and CPC per segment from base rates, adjusted by the campaign’s quality multiplier and by any cause effects applied to the driver slice\. Clicks and orders are sampled from these rates, and spend follows from clicks and CPC\.
Equation \([1](https://arxiv.org/html/2609.10315#S2.E1)\) summarizes the simulator\. Active causes can affect the generated data through three channels: total impression volume, segment allocation, and per\-segment rates\.
impressions:Ni,t\\displaystyle\\text\{impressions:\}\\quad N\_\{i,t\}∼Poisson\(μiηtVi,t\),\\displaystyle\\sim\\mathrm\{Poisson\}\\\!\\left\(\\mu\_\{i\}\\,\\eta\_\{t\}\\,V\_\{i,t\}\\right\),\(1\)segment mix:𝝅i,t\\displaystyle\\text\{segment mix:\}\\quad\\boldsymbol\{\\pi\}\_\{i,t\}=Normalize\(𝐰i⊙𝐑i,t\),\\displaystyle=\\mathrm\{Normalize\}\\\!\\left\(\\mathbf\{w\}\_\{i\}\\odot\\mathbf\{R\}\_\{i,t\}\\right\),𝐍i,t\\displaystyle\\phantom\{\\text\{segment mix:\}\\quad\}\\mathbf\{N\}\_\{i,t\}∼Dirichlet\-Multinomial\(Ni,t,α𝝅i,t\),\\displaystyle\\sim\\mathrm\{Dirichlet\\text\{\-\}Multinomial\}\\\!\\left\(N\_\{i,t\},\\alpha\\,\\boldsymbol\{\\pi\}\_\{i,t\}\\right\),segment rates:θmi,g,t\\displaystyle\\text\{segment rates:\}\\quad\\theta^\{m\}\_\{i,g,t\}=θ¯mi,g⋅Qmi⋅Fmi,g,t,m∈\{CTR,CVR,CPC\}\.\\displaystyle=\\bar\{\\theta\}^\{m\}\_\{i,g\}\\cdot Q^\{m\}\_\{i\}\\cdot F^\{m\}\_\{i,g,t\},\\;\\;m\\in\\\{\\mathrm\{CTR\},\\mathrm\{CVR\},\\mathrm\{CPC\}\\\}\.
Hereiiindexes campaigns,ttindexes days, andggindexes segments\. The total impressionsNi,tN\_\{i,t\}are sampled from the campaign baseline volumeμi\\mu\_\{i\}, a seasonality factorηt\\eta\_\{t\}, and an intervention\-induced volume multiplierVi,tV\_\{i,t\}\. Segment impressions𝐍i,t=\(Ni,g,t\)g\\mathbf\{N\}\_\{i,t\}=\(N\_\{i,g,t\}\)\_\{g\}are drawn from a multinomial distribution with probabilities𝝅i,t\\boldsymbol\{\\pi\}\_\{i,t\}, obtained by reweighting the campaign’s baseline segment weights𝐰i\\mathbf\{w\}\_\{i\}by an intervention\-induced segment\-mix multiplier𝐑i,t\\mathbf\{R\}\_\{i,t\}\. The concentration parameterα\\alphacontrols day\-to\-day variability in segment allocation\. Finally,θi,g,tm\\theta^\{m\}\_\{i,g,t\}denotes the segment\-level rate for metricm∈\{CTR,CVR,CPC\}m\\in\\\{\\mathrm\{CTR\},\\mathrm\{CVR\},\\mathrm\{CPC\}\\\}, with baseline rateθ¯i,gm\\bar\{\\theta\}^\{m\}\_\{i,g\}, campaign quality multiplierQimQ^\{m\}\_\{i\}, and intervention\-induced rate multiplierFi,g,tmF^\{m\}\_\{i,g,t\}\.
### 2\.3Cause Taxonomy
TRACEdefines twelve cause types in four categories \(Table[1](https://arxiv.org/html/2609.10315#S2.T1)\), organized by*where*the causal signal appears\. The categories require different diagnostic strategies, so no fixed query procedure is sufficient\.
*Campaign\-wide causes*produce effects visible in aggregate campaign metrics\.
*Segment\-mix causes*redistribute impressions across segments while leaving campaign totals roughly unchanged\. Their signature is a change in traffic composition across periods\.
*Segment\-specific causes*alter rates inside a localized*driver slice*, leaving other segments unaffected\. These effects can be diluted in aggregate metrics and surface only under the right split\. This category is the hardest: the agent must both localize the driver slice and match the multi\-metric signature that distinguishes one cause from another\.
Finally,*no\-signal*episodes inject no cause at all: every fluctuation is stochastic noise or a confounder\. These episodes are not trivial because confounders, such as seasonality, can appear regardless of the cause\. They test whether an agent abstains when evidence is insufficient, penalizing models that default to plausible\-sounding attributions\.
Table 1:Root\-cause signatures, grouped by*signal level*\.CauseMetricSignatureDriverDimensionCampaign\-widebid increaseCPC↑\\uparrow, CTR↓\\downarrow\(mild\)—budget capImpr↓\\downarrow, clicks↓\\downarrow, orders↓\\downarrow; rates≈\\approx—page degradationCVR↓\\downarrow, CTR≈\\approx, orders↓\\downarrow—out of stockstock↓\\downarrow, CVR↓\\downarrow, orders↓\\downarrow—Segment\-mixplacement shiftsegment share↕\\updownarrowplacementtargeting broadeningImpr↑\\uparrow, CTR↓\\downarrow, CVR↓\\downarrowaudiencetargeting narrowingImpr↓\\downarrow, CTR↑\\uparrow, CVR↑\\uparrowaudienceSegment\-specificcreative fatigueCTR↓\\downarrow\(gradual\), Impr≈\\approx, CVR≈\\approxplacement/device/geocompetitive pressureCPC↑\\uparrow, CTR↓\\downarrowgeoad quality dropImpr↓\\downarrow, CPC↑\\uparrow, CTR↓\\downarrowdeviceaudience saturationCTR↓\\downarrow\(gradual\), CVR↓\\downarrow\(mild\)audienceNo signalno signalall\|Δ\|\|\\Delta\|within noise—#### Complexity axes\.
Episode difficulty is controlled by two cause\-conditioned axes\.*Dimensional complexity*controls*where*: the signal’s driver slice may filter on one attribute or, in harder cases, the intersection of two\.*Temporal complexity*controls*when*: the signal’s onset profiles can be immediate, delayed, or ramping\. Together with the cause taxonomy, these axes test whether the agent can identify what changed and where the signal appears under different temporal patterns\.
### 2\.4Database and Task Interface
For each episode, the agent receives query access to*fact tables*that contain daily metrics at the campaign and segment level\. The ground\-truth labels are stored in separate tables that are used only for verification, scoring and training reward, and they are never exposed to the agent\. The full schema is given in Appendix[A](https://arxiv.org/html/2609.10315#A1)\.
The agent receives a system prompt specifying its role, the candidate cause types, the available tables, and the required final\-answer format\. The user prompt describes the observed performance change for a specific campaign\. The agent interacts through a Python tool backed by a sandboxed SQL engine, with state persisted across calls so that later analyses can build on earlier queries\.
The agent’s final answer contains a target cause and supporting evidence\. Segment\-specific causes additionally require an affected segment slice, such asplacement=TOP\_OF\_SEARCH\. In harder episodes, the slice may be an intersection of two attributes, such asplacement=TOP\_OF\_SEARCH, geo=US\. Campaign\-wide, segment\-mix, and no\-signal episodes require no separate segment field\.
### 2\.5Solvability and Verification
A useful benchmark here must meet three competing requirements\. Episodes must be*solvable*: the injected signal must be recoverable from the agent\-visible data\. They must be*non\-trivial*: simple signature matching should not be sufficient\. And they must be*realistic*: metrics should covary and confounders should be present, as in operational data\. These requirements are in tension\. A signal weak enough to avoid trivial detection may become unrecoverable, while a signal strong enough to guarantee recovery may collapse the task into pattern matching\.
TRACEresolves this tension by separating generation from acceptance\. The simulator first generates diverse episodes with varying causes, complexity axes, noise, and confounders\. An oracle verifier then admits only episodes that remain recoverable from the agent\-visible data\. The verifier knows the ground\-truth cause but reads only the agent\-visible data\. It does not solve the episode\. Instead, it verifies that the injected signal is detectable, stronger than the strongest confounder, and distinguishable from competing causes\. Accepted episodes therefore retain realistic ambiguity while providing objective labels for scoring and reward construction\.
## 3RL with Synthesized Rewards
TheTRACEenvironment turns each accepted episode into an agent\-visible diagnostic task paired with a hidden oracle label\. The task defines the data observed by the agent, while the label records the injected intervention\. Comparing a trajectory’s final attribution with this label produces an objective synthesized reward without human annotation or an LLM judge\. This setup enables us to train an open\-weight diagnostic agent under the same tool\-use interface used for evaluation\.
### 3\.1Reward Design
We optimize the policy with a GRPO objective\([Shao et al\., 2024](https://arxiv.org/html/2609.10315#bib.bib2)\)using stabilization refinements from recent RLVR recipes\([Yu et al\., 2026](https://arxiv.org/html/2609.10315#bib.bib6)\)\. For each training episode, the policy samples a group of trajectories and receives a synthesized reward computed against the hidden oracle label\.
The reward combines three terms: a graded attribution reward, a binary full\-attribution reward, and a small formatting reward:
r=wattrrattr\+wfullrfull\+wfmtrfmt\.r=w\_\{\\mathrm\{attr\}\}r\_\{\\mathrm\{attr\}\}\+w\_\{\\mathrm\{full\}\}r\_\{\\mathrm\{full\}\}\+w\_\{\\mathrm\{fmt\}\}r\_\{\\mathrm\{fmt\}\}\.\(2\)The weights are nonnegative and sum to one, sor∈\[0,1\]r\\in\[0,1\]\.
The attribution reward decomposes into cause correctness and slice specification:
rattr=𝟏\{c^=c⋆\}rslice,r\_\{\\mathrm\{attr\}\}=\\mathbf\{1\}\\\{\\hat\{c\}=c^\{\\star\}\\\}\\,r\_\{\\mathrm\{slice\}\},\(3\)wherec^\\hat\{c\}is the predicted cause,c⋆c^\{\\star\}is the oracle cause, andrslicer\_\{\\mathrm\{slice\}\}scores slice specification when an affected segment slice is required\. The reward termrattrr\_\{\\mathrm\{attr\}\}gates attribution by cause correctness: a wrong cause receives zero attribution reward\.
For segment\-specific episodes, letZ^\\widehat\{Z\}andZ⋆Z^\{\\star\}denote the predicted and oracle driver slices, represented as sets of dimension–value pairs\. We measure their agreement using Jaccard similarity:J\(Z^,Z⋆\)=\|Z^∩Z⋆\|\|Z^∪Z⋆\|J\(\\widehat\{Z\},Z^\{\\star\}\)=\\frac\{\|\\widehat\{Z\}\\cap Z^\{\\star\}\|\}\{\|\\widehat\{Z\}\\cup Z^\{\\star\}\|\}\. The graded slice reward is
rslice=α\+\(1−α\)J\(Z^,Z⋆\),r\_\{\\mathrm\{slice\}\}=\\alpha\+\(1\-\\alpha\)J\(\\widehat\{Z\},Z^\{\\star\}\),\(4\)which provides partial credit for identifying some, but not all, of the affected dimensions\. For campaign\-wide, segment\-mix, and no\-signal episodes, where no driver slice is required, we setrslice=1r\_\{\\mathrm\{slice\}\}=1\.
The graded attribution reward assigns substantial partial credit to an incomplete multi\-dimensional slice\. To provide a stronger incentive for recovering the complete driver slice, we introduce a binary full\-attribution rewardrfull∈\{0,1\}r\_\{\\mathrm\{full\}\}\\in\\\{0,1\\\}\. We setrfull=𝟏\{rattr=1\}r\_\{\\mathrm\{full\}\}=\\mathbf\{1\}\\\{r\_\{\\mathrm\{attr\}\}=1\\\}, which equals one only when the predicted cause is correct and the predicted driver slice exactly matches the oracle slice\. Thus,rattrr\_\{\\mathrm\{attr\}\}provides dense partial credit, whereasrfullr\_\{\\mathrm\{full\}\}rewards a complete attribution\.
The last termrfmt∈\{0,1\}r\_\{\\mathrm\{fmt\}\}\\in\\\{0,1\\\}is a format reward indicating a parseable final answer under the required schema\. Parse failures results inrattr=0r\_\{\\mathrm\{attr\}\}=0\. In our main results, we set\(wattr,wfull,wfmt\)=\(0\.65,0\.30,0\.05\)\(w\_\{\\mathrm\{attr\}\},w\_\{\\mathrm\{full\}\},w\_\{\\mathrm\{fmt\}\}\)=\(0\.65,0\.30,0\.05\)andα=0\.5\\alpha=0\.5\.
For reward computation, we parse the final answer to extract the decision fields: the root cause and, when required, the affected segment slice\. Evidence is collected for interpretability and error analysis, but it does not affect the reward\. This keeps training and evaluation focused on the same attribution target\.
### 3\.2Training Setup
Our trained agent is based onQwen3\.5\-35B\-A3B, a mixture\-of\-experts open\-weight model with 35B total parameters and 3B active parameters\. The RL dataset contains 5,000 oracle\-verified episodes, split into 4,472 training and 528 validation tasks, with stratification by cause, signal level, and 1D/2D driver slice complexity\. The evaluation dataset contains 235 held\-out tasks that are disjoint from the training and validation tasks at episode, campaign, and intervention levels\.
Training and evaluation use the same multi\-turn tool interface\. The agent executes Python containing SQL queries over agent\-visible fact tables and return a structured final answer for cause attribution\.
We compare two RL initializations, the base model and an SFT warm start checkpoint, and study the impact of reward shape and KL regularization with ablation experiments in §[4\.3](https://arxiv.org/html/2609.10315#S4.SS3)\. Full optimizer, rollout, tool\-execution, hardware, and parallelism details are provided in Appendix[B](https://arxiv.org/html/2609.10315#A2)\.
### 3\.3Supervised Warm Start
The supervised fine\-tuning \(SFT\) as a warm\-up stage before RL teaches the required answer format and multi\-turn tool\-use pattern, and exposes the model to valid diagnostic trajectories\. We construct 1,200 SFT examples by rejection\-sampling complete teacher trajectories and retaining only those whose final cause and driver slice match the oracle label\. Each example contains the task prompt, tool calls, tool outputs, and final structured answer\.
Candidate teacher trajectories are generated by Claude Opus 4\.8\. For difficult segment\-specific episodes, we also provide the teacher model with lightweight hints about the root\-cause signature, as a way of improving rejection\-sampling efficiency\. These hints are not included in the retained task prompts, RL or evaluation prompts\.
## 4Experiments
We useTRACEto evaluate whether synthesized rewards can train a stronger tool\-using diagnostic agent than prompting alone\. Our experiments ask four questions\. Is the held\-outTRACEbenchmark challenging for strong frontier models? Does RL with synthesized rewards improve full attribution when initialized from either the base model or an SFT warm start? How do reward design and policy initialization affect the gains from RL? Finally, how does post\-training change tool\-call efficiency?
### 4\.1Experimental Setup
#### Evaluation dataset\.
All reported evaluations use the same held\-out 235\-episodeTRACEtest dataset\. The dataset is deliberately enriched for segment\-specific cases: 164 episodes require identifying a driver slice, including 49 whose oracle slice intersects two dimensions\. This composition emphasizes exact attribution under fine\-grained segmentation rather than only campaign\-level cause identification\.
#### Models\.
We evaluate four variants of Qwen3\.5\-35B\-A3B: the base model, an SFT model, an RL model initialized from the base model, and an RL model initialized from the SFT checkpoint\. We compare these variants with Qwen3\.5\-122B\-A10B, Claude Opus 5, Claude Sonnet 5, GPT\-5\.5, and GPT\-5\.6 Sol baselines\.
#### Metrics\.
We evaluate diagnostic correctness using the same cause and driver\-slice criteria used by the attribution rewards in Section[3\.1](https://arxiv.org/html/2609.10315#S3.SS1)\. LetC=𝟏\{c^=c⋆\}C=\\mathbf\{1\}\\\{\\hat\{c\}=c^\{\\star\}\\\}denote root\-cause correctness\. For segment\-specific episodes, driver\-slice agreement is measured by the Jaccard similarityJ\(Z^,Z⋆\)J\(\\widehat\{Z\},Z^\{\\star\}\)\. A trajectory achieves*full attribution*whenrfull=1r\_\{\\mathrm\{full\}\}=1: the cause is correct and, when a driver slice is required, the predicted slice exactly matches the oracle slice\.
We reportCause@kk, the probability that at least one ofkksampled trajectories identifies the correct cause, andFullAttr@kk, the probability for full attribution\. We additionally report mean driver\-slice Jaccard similarity conditioned on a correct cause for segment\-specific episodes, no\-signal accuracy, decision\-parse rate, and mean executed tool calls\.
### 4\.2Main Results
Table[2](https://arxiv.org/html/2609.10315#S4.T2)summarizes our main experiment results and shows three major findings\.
First,TRACEishard and unsaturated: the strongest prompted baseline, Claude Opus 5, achieves only0\.6860\.686FullAttr@1, with substantial remaining errors on segment\-specific episodes \(Section[4\.4](https://arxiv.org/html/2609.10315#S4.SS4)\)\.
Second,RL with synthesized rewards is additive to SFT\. SFT alone reaches0\.6370\.637FullAttr@1, while SFT→\\rightarrowRL improves it by 12\.0 percentage points to0\.757\\mathbf\{0\.757\}, surpassing all prompted baselines\.
Third,post\-training can outweigh prompted model scale\. Within the Qwen3\.5 family, the post\-trained 35B model substantially outperforms the prompted 122B model \(0\.7570\.757versus0\.2830\.283\) and also exceeds every evaluated closed\-source baseline\. This suggests thatTRACErewards a learnable diagnostic procedure that prompting and additional model scale alone do not reliably elicit\.
RL initialized directly from the 35B base also improves FullAttr@1 substantially \(0\.159→0\.4340\.159\\rightarrow 0\.434\), but remains below both SFT and SFT→\\rightarrowRL\. Its decision\-parse rate is also only0\.680\.68, compared with1\.001\.00for both SFT\-initialized models, suggesting that the supervised warm start helps the policy produce scoreable decisions as well as improve attribution\.
Table 2:Performance on the 235\-episode held\-outTRACEtest set under reward\-aligned decision parsing\. Cause@1 evaluates root\-cause identification\. FullAttr@kkrequires the correct cause and, for segment\-specific episodes, an exact driver\-slice match\. Jaccard is mean Jaccard similarity on cause\-correct segment\-specific trajectories; No\-Signal is accuracy on no\-signal episodes; Parsed is the decision\-extraction rate\.ModelCause@1JaccardFullAttr@1FullAttr@5No\-SignalParsedQwen3\.5\-35B0\.1840\.5960\.1590\.3700\.410\.53\+ SFT0\.6850\.9580\.6370\.7620\.761\.00\+ RL0\.4710\.9130\.4340\.7110\.280\.68\+ SFT→\\rightarrowRL0\.8230\.9280\.7570\.8510\.491\.00Qwen3\.5\-122B0\.2960\.8420\.2830\.4340\.780\.88Claude Opus 50\.7640\.9130\.6860\.8090\.870\.99GPT\-5\.6 Sol0\.6350\.8520\.5650\.7190\.300\.99GPT\-5\.50\.5810\.8930\.5240\.6430\.610\.99Claude Sonnet 50\.4720\.8990\.4380\.6070\.851\.00
### 4\.3Training Ablations
Table[3](https://arxiv.org/html/2609.10315#S4.T3)examines the effects of the full\-attribution reward, policy initialization, and KL regularization\.
The full\-attribution reward primarily improves difficult slice attribution\.With SFT initialization, RL using only the graded attribution reward reaches0\.5960\.596FullAttr@1 and obtains no exact matches on episodes with two\-dimensional driver slices\. Addingrfullr\_\{\\mathrm\{full\}\}increases overall FullAttr@1 to0\.7570\.757, with gains from0\.670\.67to0\.920\.92on one\-dimensional slices and from0\.000\.00to0\.270\.27on two\-dimensional slices\. This pattern suggests that an explicit reward for complete attribution is especially important when success requires recovering multiple driver dimensions\.
SFT initialization remains important\.Under the same reward containingrfullr\_\{\\mathrm\{full\}\}, RL initialized from the base model reaches0\.4340\.434FullAttr@1, compared with0\.7570\.757when initialized from SFT\. Thus, RL improves the base policy, but does not recover the performance of the supervised warm start under the matched training recipe\.
KL regularization does not improve the observed result\.Adding a KL penalty to the SFT reference reduces FullAttr@1 from0\.7570\.757to0\.7240\.724\. Performance on two\-dimensional slices decreases from0\.270\.27to0\.160\.16, while one\-dimensional performance remains unchanged at0\.920\.92\. No\-signal accuracy also decreases from0\.490\.49to0\.300\.30\.
Table 3:Training ablations on the held\-outTRACEtest set\. All columns report FullAttr@1, either overall or on the indicated episode subset\. The full\-attribution condition adds the binaryrfullr\_\{\\mathrm\{full\}\}term to the graded attribution reward\.Training conditionOverall1D Slice2D SliceNo SignalSFT only0\.6370\.710\.040\.76SFT→\\rightarrowRL, graded reward0\.5960\.670\.000\.64SFT→\\rightarrowRL, \+ full\-attribution term0\.7570\.920\.270\.49SFT→\\rightarrowRL, \+ full\-attribution term \+ KL0\.7240\.920\.160\.30RL from base, \+ full\-attribution term0\.4340\.510\.030\.28
### 4\.4Performance Breakdown
Table[4](https://arxiv.org/html/2609.10315#S4.T4)decomposes FullAttr@1 by cause scope and driver\-slice dimensionality\. SFT and SFT→\\rightarrowRL approach ceiling on campaign\-wide and segment\-mix causes\. Post\-training provides its largest gain over SFT on one\-dimensional segment\-specific episodes, improving FullAttr@1 from0\.710\.71to0\.920\.92\.
Two\-dimensional attribution remains the principal challenge\. SFT→\\rightarrowRL improves FullAttr@1 from0\.040\.04to0\.270\.27on these episodes, but still performs below Claude Opus 5 at0\.330\.33\. RL also reduces no\-signal accuracy from0\.760\.76after SFT to0\.490\.49, revealing a trade\-off between stronger attribution and avoiding false\-positive diagnoses\.
Table 4:FullAttr@1 by episode type\.NNis the number of test episodes in each subset\. Two\-dimensional driver slices remain the most difficult attribution setting\.SubsetNNBaseSFTSFT→\\rightarrowRLOpus 5Campaign\-wide causes300\.460\.990\.950\.95Segment\-mix causes210\.480\.981\.000\.98Segment\-specific, 1D slice1150\.040\.710\.920\.68Segment\-specific, 2D slice490\.010\.040\.270\.33No\-signal episodes200\.410\.760\.490\.87To examine two\-dimensional errors more closely, we analyze all five saved trajectories per episode\. Among cause\-correct trajectories whose oracle slice contains adevicedimension, SFT→\\rightarrowRL recovers the correct device–value pair in66/9466/94cases, compared with14/3914/39for SFT and12/3112/31for RL from the base\. Thus, the strongest trained model is substantially better at recovering this second dimension, although incomplete slices remain common\.
The model also tends to over\-attribute changes when no signal is present\. SFT→\\rightarrowRL predictsno\_signalin only49/10049/100saved no\-signal trajectories, compared with76/10076/100after SFT\. Another recurring ambiguity occurs between segment\-specific causes with similar observable signatures\. Of the 46 SFT→\\rightarrowRL failures to identifyad\_quality\_degradation, 24 predictcompetitive\_pressure\. These results suggest that further gains require both more reliable multi\-dimensional slice recovery and better calibration among related causes\.
### 4\.5Tool\-Call Efficiency
Figure[2](https://arxiv.org/html/2609.10315#S4.F2)compares FullAttr@1 with the mean number of executed tool calls per evaluation trajectory\. SFT simultaneously improves attribution and reduces mean tool use relative to the prompted 35B base, from22\.0522\.05to10\.7510\.75calls\. RL after SFT increases mean tool use by only0\.980\.98calls, to11\.7311\.73, while improving FullAttr@1 by 12\.0 percentage points\. In contrast, RL from the base averages16\.0416\.04calls while remaining less accurate than SFT\. These results indicate that the post\-training gains are not explained by simply making more tool calls\. They also show that supervised warm start teaches a more effective investigation\-and\-stopping procedure, rather than merely encouraging additional exploration\.
Figure 2:FullAttr@1 versus mean executed tool calls per evaluation trajectory\. Each trajectory entry corresponds to one executed Python\-tool call, including unsuccessful calls\. Blue points denote fine\-tuned Qwen3\.5\-35B variants, and orange points denote prompted baselines\. Better performance lies toward the upper left\.
## 5Conclusion
We introduced a simulator–oracle–RL approach for training diagnostic\-reasoning agents when naturally occurring verifiers are scarce\. A controlled simulator samples a hidden intervention, generates the data resulting from that intervention, and retains it as an oracle label\. This construction makes the final attribution objectively verifiable without removing the noise, confounders, multi\-table evidence, and exploratory tool use that make the diagnostic task difficult\. We instantiate this approach inTRACE, a generative digital\-advertising environment containing campaign\-wide, segment\-mix, and segment\-specific root causes\.
On the held\-outTRACEbenchmark, the strongest prompted baseline reaches0\.6860\.686FullAttr@1\. SFT raises Qwen3\.5\-35B\-A3B from0\.1590\.159to0\.6370\.637, and subsequent RL with synthesized rewards further improves it to0\.7570\.757, surpassing every evaluated prompted baseline, including Qwen3\.5\-122B\-A10B\. This result provides evidence that, for this diagnostic setting, the binding constraint is access to an effective post\-training signal—including a scalable, objective reward—rather than model scale alone\.
Several directions could extend this approach\. Richer agent harnesses could provide explicit planning, memory, adaptive tool selection, and improved stopping mechanisms, allowing the policy to conduct longer and more structured investigations\. The simulator–oracle construction could also be expanded to broader intervention families, held\-out root causes, and domains such as software operations and data\-quality diagnosis, where failures can be injected and their downstream effects observed\. Finally, simulator control enables systematic curricula over noise, confounding, evidence availability, and attribution complexity, providing a way to study which diagnostic procedures transfer across environments\.
More broadly, our results demonstrate a route to constructing objective training signals for otherwise ambiguous diagnostic tasks: when a domain admits a controllable generative model, the intervention that creates a difficult problem can also provide its verifier\.
## References
- Baiet al\.\(2022\)Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.Constitutional ai: harmlessness from ai feedback\.arXiv preprint arXiv:2212\.08073\.Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p2.1)\.
- Changet al\.\(2023\)S\. Chang, J\. Wang, M\. Dong, L\. Pan, H\. Zhu, A\. H\. Li, W\. Lan, S\. Zhang, J\. Jiang, J\. Lilien,et al\.Dr\. spider: a diagnostic evaluation benchmark towards text\-to\-sql robustness\.arXiv preprint arXiv:2301\.08881\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Chenet al\.\(2025\)J\. Chen, Z\. Cai, K\. Ji, X\. Wang, W\. Liu, R\. Wang, and B\. WangTowards medical complex reasoning with llms through medical verifiable problems\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 14552–14573\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1)\.
- Gaoet al\.\(2023\)L\. Gao, J\. Schulman, and J\. HiltonScaling laws for reward model overoptimization\.InInternational Conference on Machine Learning,pp\. 10835–10866\.Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p2.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.DeepSeek\-r1 incentivizes reasoning in llms through reinforcement learning\.Nature645\(8081\),pp\. 633–638\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.10315#S1.p1.1)\.
- Huanget al\.\(2024\)Q\. Huang, J\. Vora, P\. Liang, and J\. LeskovecMLAgentBench: evaluating language agents on machine learning experimentation\.InInternational Conference on Machine Learning,pp\. 20271–20309\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1),[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Jinet al\.\(2023\)Z\. Jin, Y\. Chen, F\. Leeb, L\. Gresele, O\. Kamal, Z\. Lyu, K\. Blin, F\. Gonzalez Adauto, M\. Kleiman\-Weiner, M\. Sachan,et al\.Cladder: assessing causal reasoning in language models\.Advances in Neural Information Processing Systems36,pp\. 31038–31065\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px4.p1.1)\.
- Jinet al\.\(2024\)Z\. Jin, J\. Liu, Z\. Lyu, M\. Sachan, R\. Mihalcea, M\. Diab, D\. Ha,et al\.Can large language models infer causation from correlation?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 28663–28679\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px4.p1.1)\.
- Kicimanet al\.\(2024\)E\. Kiciman, R\. Ness, A\. Sharma, and C\. TanCausal reasoning and large language models: opening a new frontier for causality\.Transactions on machine learning research\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px4.p1.1)\.
- Kim and Yun \(2026\)M\. Kim and S\. YunProcess\-verified reinforcement learning for theorem proving via lean\.arXiv preprint arXiv:2606\.20068\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1)\.
- Leeet al\.\(2024\)H\. Lee, S\. Phatale, H\. Mansoor, T\. Mesnard, J\. Ferret, K\. R\. Lu, C\. Bishop, E\. Hall, V\. Carbune, A\. Rastogi,et al\.RLAIF vs\. rlhf: scaling reinforcement learning from human feedback with ai feedback\.InInternational Conference on Machine Learning,pp\. 26874–26901\.Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p2.1)\.
- Liet al\.\(2023\)J\. Li, B\. Hui, G\. Qu, J\. Yang, B\. Li, B\. Li, B\. Wang, B\. Qin, R\. Geng, N\. Huo,et al\.Can llm already serve as a database interface? a big bench for large\-scale database grounded text\-to\-sqls\.Advances in Neural Information Processing Systems36,pp\. 42330–42357\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Liuet al\.\(2025\)Z\. Liu, C\. Chen, W\. Li, P\. Qi, T\. Pang, C\. Du, W\. S\. Lee, and M\. LinUnderstanding r1\-zero\-like training: a critical perspective\.arXiv preprint arXiv:2503\.20783\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1)\.
- Luoet al\.\(2025a\)H\. Luo, Q\. Sun, C\. Xu, P\. Zhao, J\. Lou, C\. Tao, X\. Geng, Q\. Lin, S\. Chen, Y\. Tang,et al\.Wizardmath: empowering mathematical reasoning for large language models via reinforced evol\-instruct\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 49573–49609\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1)\.
- Luoet al\.\(2025b\)M\. Luo, N\. Jain, J\. Singh, S\. Tan, A\. Patel, Q\. Wu, A\. Ariyak, C\. Cai, S\. Z\. T\. Venkat, S\. Zhu, B\. Athiwaratkun, M\. Roongta, C\. Zhang, L\. E\. Li, R\. A\. Popa, K\. Sen, and I\. StoicaDeepSWE: training a fully open\-sourced, state\-of\-the\-art coding agent by scaling RL\.Note:Together AI / Agentica technical blogTechnical blogExternal Links:[Link](https://www.together.ai/blog/deepswe)Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.10315#S1.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.Deepseekmath: pushing the limits of mathematical reasoning in open language models\.arXiv preprint arXiv:2402\.03300\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.10315#S1.p1.1),[§3\.1](https://arxiv.org/html/2609.10315#S3.SS1.p1.1)\.
- Skalseet al\.\(2022\)J\. Skalse, N\. Howe, D\. Krasheninnikov, and D\. KruegerDefining and characterizing reward gaming\.Advances in Neural Information Processing Systems35,pp\. 9460–9471\.Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p2.1)\.
- Wanget al\.\(2022\)R\. Wang, P\. Jansen, M\. Côté, and P\. AmmanabroluScienceworld: is your agent smarter than a 5th grader?\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 11279–11298\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1)\.
- Wanget al\.\(2023\)Y\. Wang, Y\. Kordi, S\. Mishra, A\. Liu, N\. A\. Smith, D\. Khashabi, and H\. HajishirziSelf\-instruct: aligning language models with self\-generated instructions\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: long papers\),pp\. 13484–13508\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1)\.
- Wang \(2024\)Z\. WangCausalbench: a comprehensive benchmark for evaluating causal reasoning capabilities of large language models\.InProceedings of the 10th SIGHAN Workshop on Chinese Language Processing \(SIGHAN\-10\),pp\. 143–151\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px4.p1.1)\.
- Wei \(2025\)J\. WeiAsymmetry of verification and verifier’s law\.Note:Blog postAccessed 2026\-06\-16External Links:[Link](https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law)Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p1.1)\.
- Weiet al\.\(2026\)Y\. Wei, O\. Duchenne, J\. Copet, Q\. Carbonneaux, L\. Zhang, D\. Fried, G\. Synnaeve, R\. Singh, and S\. WangSwe\-rl: advancing llm reasoning via reinforcement learning on open software evolution\.Advances in Neural Information Processing Systems38,pp\. 78500–78525\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.10315#S1.p1.1)\.
- Yuet al\.\(2024\)L\. Yu, W\. Jiang, H\. Shi, J\. Yu, Z\. Liu, Y\. Zhang, J\. Kwok, Z\. Li, A\. Weller, and W\. LiuMetamath: bootstrap your own mathematical questions for large language models\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 45040–45061\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1)\.
- Yuet al\.\(2026\)Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.Dapo: an open\-source llm reinforcement learning system at scale\.Advances in Neural Information Processing Systems38,pp\. 113222–113244\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2609.10315#S3.SS1.p1.1)\.
- Yuet al\.\(2018\)T\. Yu, R\. Zhang, K\. Yang, M\. Yasunaga, D\. Wang, Z\. Li, J\. Ma, I\. Li, Q\. Yao, S\. Roman,et al\.Spider: a large\-scale human\-labeled dataset for complex and cross\-domain semantic parsing and text\-to\-sql task\.InProceedings of the 2018 conference on empirical methods in natural language processing,pp\. 3911–3921\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Zhanget al\.\(2026\)D\. Zhang, S\. Zhoubian, M\. Cai, F\. Li, L\. Yang, W\. Wang, T\. Dong, Z\. Hu, J\. Tang, and Y\. YueDatascibench: an llm agent benchmark for data science\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 3685–3728\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Zhenget al\.\(2025\)C\. Zheng, S\. Liu, M\. Li, X\. Chen, B\. Yu, C\. Gao, K\. Dang, Y\. Liu, R\. Men, A\. Yang,et al\.Group sequence policy optimization\.arXiv preprint arXiv:2507\.18071\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px1.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing,et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems36,pp\. 46595–46623\.Cited by:[§1](https://arxiv.org/html/2609.10315#S1.p2.1)\.
- Zhonget al\.\(2017\)V\. Zhong, C\. Xiong, and R\. SocherSeq2sql: generating structured queries from natural language using reinforcement learning\.arXiv preprint arXiv:1709\.00103\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px3.p1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§1\.1](https://arxiv.org/html/2609.10315#S1.SS1.SSS0.Px2.p1.1)\.
## Appendix ADatabase and Output Schemas
### A\.1Agent\-Visible Fact Tables
Each episode provides access to a DuckDB fact store containing the four tables summarized in Table[5](https://arxiv.org/html/2609.10315#A1.T5)\. Rows are keyed by campaign identifiers, and the task prompt specifies the campaign to investigate\. The agent can access these tables only through its Python/SQL tool; no ground\-truth field is included in the tool\-visible database\.
Table 5:Agent\-visible database schema\. All tables include a date fielddsand campaign identifiers as appropriate\.TableGrainFieldsdaily\_campaignCampaign\-dayCampaign identifiers and category; budget, spend, impressions, clicks, orders, sales, page views, and brand searches\.segment\_dailySegment\-dayCampaign identifiers and category; ad product, targeting type, match type, placement, price band, audience, device, and geography; delivery, engagement, and conversion metrics\.budget\_logCampaign\-dayAdvertiser and campaign identifiers, budget, and spend\.inventoryCampaign\-dayCampaign identifiers and category, together with the fraction of advertised products that are in stock\.
### A\.2Hidden Oracle Tables
The generation database additionally contains three oracle\-only tables\.gt\_episoderecords the injected root cause, affected segment assignment, intervention strength, temporal pattern, and number of affected dimensions\.gt\_evidencestores post\-computed evidence associated with the intervention, including its source table, metric, segment, direction, and change relative to the baseline period\.gt\_validationrecords the oracle verifier’s detectability and confounder\-elimination checks\. These tables are retained for dataset construction and scoring but are excluded when the agent\-visible database is created\.
The required final answer contains adecisionobject with a non\-emptyroot\_causeand an optionaldriver\_segmentslist, anevidencelist, and a textualexplanation\. For campaign\-wide, segment\-mix, and no\-signal episodes,driver\_segmentsis omitted or null\. For segment\-specific episodes, it contains one or more dimension–value assignments\.
### A\.3Oracle Verification
After simulation, the oracle verifier compares the episode window with an equal\-length preceding baseline window\. It first checks that the expected signature is present at the appropriate level: campaign metrics for campaign\-wide causes, impression\-share changes for segment\-mix causes, and changes within the injected segment for segment\-specific causes\. For delayed and ramping interventions, it also tests the second half of the episode window so that a recoverable late\-onset signal is not rejected solely because it is diluted in the full\-window average\. Segment\-specific signals must additionally exceed their own baseline temporal variation, using a minimum signal\-to\-noise ratio of 1\.
The verifier then tests each alternative cause against the same agent\-visible data\. An alternative is eliminated using differences in signal level, source table, affected dimension, primary metric, or companion\-metric direction\. Episodes are rejected when the injected signal is absent or when a matching alternative cannot be eliminated\. No\-signal episodes are retained only when observed fluctuations do not realize another cause’s signature\. Thus, acceptance uses the hidden intervention to define what must be verified, but all detectability and distinguishability checks operate on the same fact tables available to the agent\.
## Appendix BTraining Details
### B\.1Training Setup
We train Qwen3\.5\-35B\-A3B using aslimefork with Megatron\-LM for optimization, SGLang for rollout inference, and Ray for orchestration\. RL uses an asynchronous pipeline in which rollout generation for the next step overlaps policy optimization, with updated weights synchronized from Megatron\-LM to SGLang after each optimizer step\. Training uses bfloat16 parameters with float32 gradient accumulation, softmax, and router computation\.
### B\.2Supervised Fine\-Tuning
The SFT stage trains on 1,200 oracle\-filtered teacher trajectories for three epochs\. We use a learning rate of10−510^\{\-5\}with cosine decay and apply the language\-model loss only to assistant turns under the Qwen3\.5 chat template\. Each example retains the complete interaction, including prompts, tool calls, tool outputs, and the final structured answer\.
### B\.3Reinforcement Learning
Table[6](https://arxiv.org/html/2609.10315#A2.T6)gives the shared GRPO configuration\. Each optimizer step draws 32 prompts and samples eight trajectories per prompt\. The agent has a 32,768\-token context window, a 30\-turn backstop, and a maximum of 2,048 generated tokens per turn\. Training rollouts use temperature1\.01\.0to maintain within\-group exploration\.
Table 6:Shared hyperparameters for the GRPO runs\.HyperparameterValueTraining set4,472 promptsValidation set528 promptsOptimizer steps300 \(approximately two epochs\)Prompts per step32Samples per prompt8Effective trajectories per step256OptimizerAdamW,β=\(0\.9,0\.98\)\\beta=\(0\.9,0\.98\)Learning rate10−610^\{\-6\}, constantWeight decay0\.1PPO clipping range\[0\.8,1\.28\]\[0\.8,1\.28\]Gradient clipping1\.0Entropy coefficient0Rollout temperature1\.0Maximum turns30Context length32,768 tokensMaximum generation per turn2,048 tokensValidation frequencyEvery 20 optimizer stepsValidation samplingFive trajectories per prompt at temperature 0\.6The main RL condition uses the reward weights reported in Section[3\.1](https://arxiv.org/html/2609.10315#S3.SS1)\. The KL ablation adds a penalty to the SFT reference policy with coefficient0\.0050\.005, using the low\-variance k3 estimator\. All RL conditions use the same data split, rollout group size, optimizer settings, and nominal 300\-step budget\. The validation split is stratified by root cause, signal level, and one\- versus two\-dimensional segment attribution and is disjoint from both the RL training prompts and SFT demonstrations\.
## Appendix CAdditional Evaluation Details
### C\.1Evaluation Protocol
All models receive the same task specification, candidate\-cause definitions, final\-answer contract, and Python/SQL tool interface over the agent\-visible fact store\. The Python state persists across calls within an episode\. Tool calls have a 60\-second execution timeout, and tool outputs are truncated after 8,000 characters\. A final\-turn nudge asks the model to return its answer before exhausting the interaction budget\.
Open\-weight evaluations sample five trajectories per task at temperature0\.60\.6with a 60\-step backstop\. API\-model evaluations use provider\-supported stochastic sampling without an explicit temperature and a 30\-turn backstop\. The reported Claude and GPT baselines use thexhighreasoning\-effort setting\. Claude models receive a maximum output\-token budget of 32,768, while GPT models receive 16,384 output tokens\. Every reported FullAttr@5 value is computed from five distinct samples per task\. The interaction limits do not bind for the trained or frontier models; one Qwen3\.5\-35B base\-model trajectory is truncated\.
Fornnsampled trajectories containingccsuccessful trajectories, Cause@kkand FullAttr@kkuse the standard unbiased estimator
1−\(n−ck\)\(nk\)\.1\-\\frac\{\\binom\{n\-c\}\{k\}\}\{\\binom\{n\}\{k\}\}\.\(5\)Cause@kktreats root\-cause correctness as success, whereas FullAttr@kkadditionally requires an exact segment assignment when the episode is segment\-specific\. Atk=1k=1, the estimator is the mean single\-trajectory success probability over the test tasks\.
### C\.2Answer\-Parser Sensitivity
For the primary evaluation, we parse each model’s final answer to extract the decision fields needed to assess attribution: a non\-empty root cause and any predicted segment assignment\. Malformed report\-only evidence, an omitted explanation, or benign extra JSON keys do not invalidate an otherwise scoreable decision\. A malformed segment assignment is treated as absent and therefore cannot earn full attribution on a segment\-specific episode\. We report the decision\-parse rate separately from diagnostic accuracy\.
As a sensitivity analysis, we also use an exact\-schema parser that requires the complete prescribed output structure\. Table[7](https://arxiv.org/html/2609.10315#A3.T7)reports a paired re\-scoring of the same saved trajectories under both parsers\. The parser choice has no effect on the SFT or SFT→\\rightarrowRL results, but exact\-schema parsing underestimates FullAttr@1 when a model produces a semantically scoreable decision with malformed or missing report\-only fields\.
Table 7:FullAttr@1 obtained by re\-scoring the same trajectories with an exact\-schema parser and the decision parser used for the main results\.ModelExact SchemaDecision ParserQwen3\.5\-35B0\.1400\.159\+ SFT0\.6370\.637\+ RL0\.3170\.434\+ SFT→\\rightarrowRL0\.7570\.757Claude Opus 50\.6650\.686Similar Articles
Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short
Reasoning Arena improves reinforcement learning with verifiable rewards by using trace tournaments and Bradley-Terry models to generate meaningful gradients from non-diverse reward groups, resulting in faster training and better reasoning performance.
Inducing Reasoning Primitives from Agent Traces
Introduces Reasoning Primitive Induction, a method that mines successful ReAct traces to cluster recurrent reasoning moves into typed pseudo-tools, outperforming the original agent by tens of percentage points on benchmarks.
From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization
Introduces STRACE, a framework that performs structural trajectory analysis and causal extraction to construct high signal-to-noise optimization contexts for improving long-horizon agents, outperforming baselines on a formal verification task.
Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents
The paper introduces Trace, a credit-guided framework that compiles noisy interaction histories into executable walkthroughs to enhance long-horizon AI agent performance, showing significant improvements in experiments.
TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
TRACE is a unified rollout budget allocation framework that enhances reward contrast in multi-turn agentic reinforcement learning by dynamically distributing resources across tree-structured rollouts based on prefix-level informativeness. It improves efficiency and accuracy on agentic benchmarks like Multi-Hop QA.