AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems

arXiv cs.CL Papers

Summary

This paper introduces AgentForesight, a framework for online auditing and early failure prediction in LLM-based multi-agent systems. It presents a new dataset, AFTraj-22K, and a specialized model, AgentForesight-7B, which outperforms leading proprietary models in detecting decisive errors during trajectory execution.

arXiv:2605.08715v1 Announce Type: new Abstract: LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level failure. Existing work frames this as \emph{post-hoc failure attribution}, diagnosing the responsible agent and step after the trajectory has ended. However, this paradigm forfeits any opportunity to intervene while trajectory is still unfolding. In this work, we introduce AgentForesight, a framework that reframes this problem as online auditing: at each step of an unfolding trajectory, an auditor observes only the current prefix and must either continue the run or alarm at the earliest decisive error, without access to future steps. To this end, we curate AFTraj-2K, a corpus of agentic trajectories across Coding, Math, and Agentic domains, in which safe trajectories are retained under a strict curation pipeline and unsafe trajectories are annotated at the step of their decisive error via consensus among multiple LLM judges. Built on that, we develop AgentForesight-7B, a compact online auditor trained with a coarse-to-fine reinforcement learning recipe that first equips it with a risk-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs, then sharpens this prior into precise step-level localization under a three-axis reward jointly targeting the what, where, and who of an audit verdict. Across AFTraj-2K and an external Who\&When benchmark, AgentForesight-7B outperforms leading proprietary models, including GPT-4.1 and DeepSeek-V4-Pro, achieving up to +19.9% performance gain and 3$\times$ lower step localization error, opening the loop from post-hoc failures detection to enabling deployment-time intervention. Project page: https://zbox1005.github.io/agent-foresight/
Original Article
View Cached Full Text

Cached at: 05/12/26, 06:55 AM

# AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Source: [https://arxiv.org/html/2605.08715](https://arxiv.org/html/2605.08715)
Boxuan Zhang1Jianing Zhu2Zeru Shi1Dongfang Liu3Ruixiang Tang1 1Rutgers University2The University of Texas at Austin3Purdue University \{bz362, rt836\}@scarletmail\.rutgers\.edu

###### Abstract

LLM\-based multi\-agent systems are increasingly deployed on long\-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory\-level failure\. Existing work frames this as*post\-hoc failure attribution*, diagnosing the responsible agent and step after the trajectory has ended\. However, this paradigm forfeits any opportunity to intervene while trajectory is still unfolding\. In this work, we introduceAgentForesight, a framework that reframes this problem as*online auditing*: at each step of an unfolding trajectory, an auditor observes only the current prefix and must either continue the run or alarm at the earliest decisive error without access to future steps\. To this end, we curateAFTraj\-22K, a corpus of agentic trajectories across Coding, Math, and Agentic domains, in which safe trajectories are retained under a strict curation pipeline and unsafe trajectories are annotated at the step of their decisive error via consensus among multiple LLM judges\. Built on that, we develop*AgentForesight*\-7B, a compact online auditor trained with a*coarse\-to\-fine*reinforcement learning recipe that first equips it with a risk\-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs, then sharpens this prior into precise step\-level localization under a three\-axis reward jointly targeting the*what*,*where*, and*who*of an audit verdict\. AcrossAFTraj\-22Kand an external Who&When benchmark,*AgentForesight*\-7B outperforms leading proprietary models, including GPT\-4\.1 and DeepSeek\-V4\-Pro, achieving up to\+19\.9%performance gain and3×\\timeslower step localization error, opening the loop from post\-hoc failure detection to enabling deployment\-time intervention\.

## 1Introduction

Large language models \(LLMs\) have rapidly evolved into agentic systems that plan, reason, and act across long\-horizon tasks through coordinated tool use and inter\-agent communication\[[59](https://arxiv.org/html/2605.08715#bib.bib59),[53](https://arxiv.org/html/2605.08715#bib.bib53),[17](https://arxiv.org/html/2605.08715#bib.bib17),[28](https://arxiv.org/html/2605.08715#bib.bib28)\]\. By decomposing complex objectives into specialized sub\-tasks, these systems now tackle problems once considered out of reach, spanning software development\[[20](https://arxiv.org/html/2605.08715#bib.bib20),[51](https://arxiv.org/html/2605.08715#bib.bib51)\], scientific discovery\[[11](https://arxiv.org/html/2605.08715#bib.bib11),[12](https://arxiv.org/html/2605.08715#bib.bib12)\], and open\-ended web navigation\[[66](https://arxiv.org/html/2605.08715#bib.bib66),[33](https://arxiv.org/html/2605.08715#bib.bib33)\]\. However, such gains in capability come with a structural cost\. Since each step is conditioned on earlier outputs, a single*decisive error*, e\.g\., a malformed tool call or a flawed intermediate deduction, is easily accepted by downstream agents and cascades into a full\-trajectory failure\[[3](https://arxiv.org/html/2605.08715#bib.bib3),[63](https://arxiv.org/html/2605.08715#bib.bib63),[25](https://arxiv.org/html/2605.08715#bib.bib25)\]\. Once deployed in real\-world environments with access to APIs and external services, such failures extend beyond benchmark accuracy into unanticipated operational risks\[[60](https://arxiv.org/html/2605.08715#bib.bib60),[45](https://arxiv.org/html/2605.08715#bib.bib45)\], making reliability a central bottleneck for the deployment of LLM multi\-agent systems\.

Although prior work has recognized failure analysis as a central concern for reliable LLM multi\-agent systems, existing approaches predominantly frame it as*post\-hoc failure attribution*, asking which agent or step is responsible once the trajectory has already failed\[[63](https://arxiv.org/html/2605.08715#bib.bib63),[62](https://arxiv.org/html/2605.08715#bib.bib62),[67](https://arxiv.org/html/2605.08715#bib.bib67)\], as illustrated in Figure[1](https://arxiv.org/html/2605.08715#S1.F1)\(a\)\. For instance, Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]and AgenTracer\[[62](https://arxiv.org/html/2605.08715#bib.bib62)\]curate failed trajectories and train or prompt models to pinpoint the decisive error step after the run has ended, while AgentDebug\[[67](https://arxiv.org/html/2605.08715#bib.bib67)\]and related debugging frameworks\[[49](https://arxiv.org/html/2605.08715#bib.bib49),[19](https://arxiv.org/html/2605.08715#bib.bib19)\]analyze full trajectories to taxonomize failures and supply corrective feedback for subsequent retries\. However, confining failure analysis to the post\-hoc regime forgoes any opportunity to act while the trajectory is still unfolding\. Before a diagnosis is available, agents have already consumed further tool calls and external resources, and in deployment settings may have triggered irreversible side effects\. This naturally motivates a fundamental research question:

> Can we audit unfolding prefixes rather than completed trajectories to catch decisive errors before propagation locks in failure?

To answer this question, we introduce*online auditing*, where a dedicated auditor commits a continue\-or\-alarm verdict at every step of an unfolding trajectory, as illustrated in Figure[1](https://arxiv.org/html/2605.08715#S1.F1)\(b\)\. Concretely, instead of inspecting a completed trajectory with full hindsight, the auditor sees only the current*prefix*at each step and must judge it without access to future steps, tool responses, or the eventual outcome\. This reframe turns failure analysis from a passive post\-hoc diagnosis of completed runs into an active safeguard that can intervene before downstream propagation locks in the failure\. Operationalizing it places two new demands on the auditor:①it must reliably separate prefixes that are still safe from those already past a*decisive error*, and②it must commit at the very step the error occurs, not in hindsight\. Both demands exceed what existing failure\-attribution data or models can provide, motivating the creation of both a new dataset and a dedicated training recipe\.

![Refer to caption](https://arxiv.org/html/2605.08715v1/x1.png)Figure 1:Comparison of\(a\)*post\-hoc failure attribution*and\(b\)*online auditing*on the same multi\-agent task\.\(a\)Post\-hoc failure attribution inspects the trajectory only*after*it has failed and identifies the decisive error retrospectively, by which point downstream propagation has already locked in the failure\.\(b\)OurAgentForesightinstead evaluates each*prefix*as the trajectory unfolds and flags the decisive error at the very step it commits, opening an intervention window before the failure is locked in \(see Section[2](https://arxiv.org/html/2605.08715#S2)\)\.To instantiate this formulation, we developAgentForesight, a framework that addresses these two demands through a dedicated dataset and a*coarse\-to\-fine*training recipe\. We first constructAFTraj\-22K, a curated corpus of agentic trajectories spanning Coding, Math, and Agentic domains, pairing safe trajectories retained under a strict filtering pipeline with failure trajectories annotated at their*decisive error*step under multi\-judge voting verification\. Building on the curated dataset, we fine\-tune Qwen2\.5\-7B\-Instruct via reinforcement learning to obtain*AgentForesight*\-7B, a compact online auditor first equipped with a risk\-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs, then sharpened into precise step\-level localization under a three\-axis reward jointly targeting the structure of verdict \(*what*\), the timing of alarm \(*where*\), and the responsible agent \(*who*\)\. Together,*AgentForesight*\-7B runs alongside off\-the\-shelf multi\-agent systems and issues step\-level continue\-or\-alarm verdicts on unfolding trajectories, without retraining the underlying agentic system\.

We extensively evaluate*AgentForesight*\-7B onAFTraj\-22Kand the external Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]benchmark, where it surpasses both its Qwen2\.5\-7B\-Instruct base model and leading proprietary judges including GPT\-4\.1 and DeepSeek\-V4\-Pro, achieving\+19\.9%\+19\.9\\%higher Exact\-F1 and3×3\\timeslower step localization error than the strongest proprietary baseline\. These gains confirm that our coarse\-to\-fine recipe yields a compact online auditor that outperforms much larger proprietary judges under the prefix\-restricted online setting\. We summarize our contributions as follows:

- •We introduce*online auditing*, a deployment\-time reframing of agentic failure analysis that audits unfolding trajectories step by step rather than diagnosing them after failure \(Section[2](https://arxiv.org/html/2605.08715#S2)\)\.
- •We constructAFTraj\-22K, a curated corpus of agentic trajectories spanning Coding, Math, and Agentic domains, pairing strictly filtered safe runs with multi\-judge verified failure runs annotated at their*decisive error*step \(Section[3\.1](https://arxiv.org/html/2605.08715#S3.SS1)\)\.
- •We develop*AgentForesight*\-7B, a compact online auditor trained via a*coarse\-to\-fine*RL recipe that first equips it with a risk\-anticipation prior at the failure boundary, then sharpens this prior into precise step\-level localization under the structure, timing, and attribution optimization \(Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)\)\.
- •We empirically show that*AgentForesight*\-7B surpasses its base model and leading proprietary judges onAFTraj\-22Kand Who&When benchmark \(Section[4](https://arxiv.org/html/2605.08715#S4)\)\.

## 2Problem Formulation

We formalize the problem of monitoring multi\-agent failures under two settings: ①*post\-hoc failure attribution*, the prevailing setup in prior work\[[63](https://arxiv.org/html/2605.08715#bib.bib63),[62](https://arxiv.org/html/2605.08715#bib.bib62),[67](https://arxiv.org/html/2605.08715#bib.bib67)\], and ②*online auditing*, the deployment\-time formulation we introduce\. We first define the shared trajectory model and*decisive error*, then specify the formal setup for each setting, and close with a contrast clarifying the scope of our contribution\.

##### Multi\-Agent Trajectory\.

We model a multi\-agent execution as a turn\-based systemℳ=\(𝒮,𝒩,πsys,Ψ,Ω\)\\mathcal\{M\}=\(\\mathcal\{S\},\\mathcal\{N\},\\pi\_\{\\text\{sys\}\},\\Psi,\\Omega\), where𝒮\\mathcal\{S\}is the set of system states,𝒩\\mathcal\{N\}is the finite set of agent roles \(e\.g\.,Planner,WebAgent,CodeWriter\),πsys\\pi\_\{\\text\{sys\}\}is the system policy that produces the next turn given the current state,Ψ\\Psiis the state\-update function, andΩ:𝒯→\{0,1\}\\Omega:\\mathcal\{T\}\\to\\\{0,1\\\}is the binary outcome function that judges a completed trajectory against the task specification \(Ω​\(τ\)=1\\Omega\(\\tau\)=1for success,0for failure\), with𝒯\\mathcal\{T\}denoting the space of finite trajectories\. The observed trajectory ofℳ\\mathcal\{M\}is a sequence of turns,

τ=\(t0,t1,…,tN−1\),ti=\(rolei,actioni,contenti\),\\tau=\(t\_\{0\},t\_\{1\},\\ldots,t\_\{N\-1\}\),\\qquad t\_\{i\}=\(\\mathrm\{role\}\_\{i\},\\mathrm\{action\}\_\{i\},\\mathrm\{content\}\_\{i\}\),\(1\)whereNNis the trajectory length,rolei∈𝒩\\mathrm\{role\}\_\{i\}\\in\\mathcal\{N\}identifies the agent at turntit\_\{i\}, and the pair\(actioni,contenti\)\(\\mathrm\{action\}\_\{i\},\\mathrm\{content\}\_\{i\}\)records its action together with the resulting observable content\.

##### Decisive Error\.

Following\[[63](https://arxiv.org/html/2605.08715#bib.bib63),[62](https://arxiv.org/html/2605.08715#bib.bib62)\], we adopt the*decisive error*, whose correction would have flipped the trajectory outcome from failure to success, as the operational unit of failure analysis\.

###### Definition 2\.1\(Decisive error\)

For a failure trajectoryτ\\tauwithΩ​\(τ\)=0\\Omega\(\\tau\)=0, letτ\+:=τ0:k−1⊕t~\\tau^\{\+\}:=\\tau\_\{0:k\-1\}\\oplus\\tilde\{t\}denote the prefix with stepkkreplaced by an admissible correctiont~\\tilde\{t\}\. The*decisive error step*is

k∗=min⁡\{k∈\[T\]:∃t~∈𝒯k​\(τ0:k−1\),τ~∈ℛℳ​\(τ\+\),Ω​\(τ\+⊕τ~\)=1\},k^\{\*\}\\;=\\;\\min\\bigl\\\{\\,k\\in\[T\]\\,:\\,\\exists\\,\\tilde\{t\}\\in\\mathcal\{T\}\_\{k\}\(\\tau\_\{0:k\-1\}\),\\;\\tilde\{\\tau\}\\in\\mathcal\{R\}\_\{\\mathcal\{M\}\}\(\\tau^\{\+\}\),\\;\\Omega\(\\tau^\{\+\}\\oplus\\tilde\{\\tau\}\)=1\\,\\bigr\\\},\(2\)where𝒯k​\(τ0:k−1\)\\mathcal\{T\}\_\{k\}\(\\tau\_\{0:k\-1\}\)is the set of admissible correct turns at positionkk, andℛℳ​\(⋅\)\\mathcal\{R\}\_\{\\mathcal\{M\}\}\(\\cdot\)is the set of suffix trajectories reachable from a corrected prefix under the system policyπsys\\pi\_\{\\text\{sys\}\}\. Intuitively,k∗k^\{\*\}is the earliest step whose error cannot be recovered by any downstream rollout underπsys\\pi\_\{\\text\{sys\}\}, so that an oracle correction atk∗k^\{\*\}is both necessary and sufficient to salvage the trajectory\. We calla∗=rolek∗a^\{\*\}=\\mathrm\{role\}\_\{k^\{\*\}\}the*responsible agent*, and annotate failed trajectory with\(k∗,a∗\)\(k^\{\*\},a^\{\*\}\), while successful ones with\(SAFE,∅\)\(\\textsc\{SAFE\},\\emptyset\)\.

##### Post\-hoc Failure Attribution\.

Prior methods\[[63](https://arxiv.org/html/2605.08715#bib.bib63),[62](https://arxiv.org/html/2605.08715#bib.bib62),[67](https://arxiv.org/html/2605.08715#bib.bib67)\]take a completed failure trajectoryτ\\tautogether with its terminal outcomeΩ​\(τ\)=0\\Omega\(\\tau\)=0as input, and emit a single retrospective prediction:

y^post=fpost​\(τ\)=\(k^,a^\)∈\{0,…,N−1\}×𝒩\.\\hat\{y\}\_\{\\text\{post\}\}=f\_\{\\text\{post\}\}\(\\tau\)=\(\\hat\{k\},\\hat\{a\}\)\\in\\\{0,\\ldots,N\{\-\}1\\\}\\times\\mathcal\{N\}\.\(3\)Three properties characterise this setup:\(i\)*full hindsight*overτ\\tauandΩ​\(τ\)\\Omega\(\\tau\);\(ii\)*single\-shot*output;\(iii\)prediction occurs*after*the failure has materialized, leaving no intervention window\.

##### Online Auditing\.

Online auditing reframes failure analysis as a deployment\-time decision, where an auditor runs alongside the multi\-agent system at every step and decides, on prefix evidence alone, whether to allow execution to continue\.

###### Definition 2\.2\(Online auditing\)

Letτ0:k=\(t0,…,tk\)\\tau\_\{0:k\}=\(t\_\{0\},\\ldots,t\_\{k\}\)denote the prefix ofτ\\tauup to turntkt\_\{k\}\. An online auditor is a function

y^k=fonline​\(τ0:k\)∈\{Continue\}∪\(\{Alarm\}×\{0,…,k\}×𝒩\),\\hat\{y\}\_\{k\}=f\_\{\\text\{online\}\}\(\\tau\_\{0:k\}\)\\;\\in\\;\\\{\\textsc\{Continue\}\\\}\\;\\cup\\;\\bigl\(\\\{\\textsc\{Alarm\}\\\}\\times\\\{0,\\ldots,k\\\}\\times\\mathcal\{N\}\\bigr\),\(4\)applied at each stepk=0,…,N−1k=0,\\ldots,N\{\-\}1\. AContinueverdict signals that no decisive error has yet been observed in the visible window, while anAlarmverdict halts execution and reports a predicted decisive error stepk^∈\{0,…,k\}\\hat\{k\}\\in\\\{0,\\ldots,k\\\}together with the predicted responsible agenta^∈𝒩\\hat\{a\}\\in\\mathcal\{N\}\.

The setup inverts the three post\-hoc properties:\(i\)only*prefix\-restricted*information, with no access totk\+1:N−1t\_\{k\+1:N\-1\}or the terminal label;\(ii\)*per\-step*output,NNverdicts per trajectory;\(iii\)anAlarmat stepkkcreates an*intervention window*beforetk\+1t\_\{k\+1\}is committed\. Directly applyingfpostf\_\{\\text\{post\}\}to each prefix is ill\-posed, since they are trained assuming thatΩ​\(τ\)=0\\Omega\(\\tau\)=0is observed, which fails on a live prefix\.

## 3Methodology

In this section, we presentAgentForesight, a framework that operationalizes the demands of online auditing through \(1\) a curated corpusAFTraj\-22Ksupplying prefix\-level supervision \(Section[3\.1](https://arxiv.org/html/2605.08715#S3.SS1)\), and \(2\) a coarse\-to\-fine training recipe producing the compact online auditor*AgentForesight*\-7B \(Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)\)\. Detailed pseudocode for both components is provided in Appendix[A](https://arxiv.org/html/2605.08715#A1)\.

### 3\.1AFTraj\-22K: A Curated Corpus for Online Agentic Auditing

The online\-auditing setup of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2)demands training data with three properties absent from existing failure\-attribution corpora: \(i\) per\-step ground truth\(k∗,a∗\)\(k^\{\*\},a^\{\*\}\)for unsafe trajectories, \(ii\) verified safe trajectories that admit prefix\-restricted supervision at every step, and \(iii\) coverage across heterogeneous multi\-agent frameworks and task domains\. Existing open\-source benchmarks fall short on at least one of these axes\. Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]provides step\-level decisive\-error annotations but contains only failed trajectories, leaving the safe regime unsupervised; ATBench\[[25](https://arxiv.org/html/2605.08715#bib.bib25)\]includes both safe and unsafe trajectories but focuses on safety\-specific tasks and supplies only trajectory\-level labels\. We therefore constructAFTraj\-22K, a unified corpus of multi\-agent trajectories collected, filtered, and annotated for online auditing\. Figure[2](https://arxiv.org/html/2605.08715#S3.F2)\(a\) illustrates the construction pipeline\.

##### Trajectory Collection\.

We instantiate multi\-agent systems on a suite of off\-the\-shelf frameworks\[[53](https://arxiv.org/html/2605.08715#bib.bib53),[17](https://arxiv.org/html/2605.08715#bib.bib17),[41](https://arxiv.org/html/2605.08715#bib.bib41)\]and run them on tasks spanning mathematical reasoning\[[16](https://arxiv.org/html/2605.08715#bib.bib16)\], code generation\[[29](https://arxiv.org/html/2605.08715#bib.bib29)\], and open\-ended agentic problem solving\[[57](https://arxiv.org/html/2605.08715#bib.bib57),[33](https://arxiv.org/html/2605.08715#bib.bib33)\]\. This diversity in role decompositions, tool stacks, and task structure promotes broad coverage of multi\-agent dynamics rather than the idiosyncrasies of any single system\. Each rollout yields a turn\-level trajectoryτ∈𝒯\\tau\\in\\mathcal\{T\}as defined in Eq\.[1](https://arxiv.org/html/2605.08715#S2.E1), scored by the outcome functionΩ:𝒯→\{0,1\}\\Omega:\\mathcal\{T\}\\to\\\{0,1\\\}against the reference solution\. The raw pool of collected trajectories then partitions into two disjoint subsets,

𝒟succ=\{τ∣Ω​\(τ\)=1\},𝒟fail=\{τ∣Ω​\(τ\)=0\},\\mathcal\{D\}\_\{\\text\{succ\}\}=\\\{\\tau\\mid\\Omega\(\\tau\)=1\\\},\\quad\\mathcal\{D\}\_\{\\text\{fail\}\}=\\\{\\tau\\mid\\Omega\(\\tau\)=0\\\},\(5\)which feed the two parallel branches of the construction pipeline:𝒟succ\\mathcal\{D\}\_\{\\text\{succ\}\}supplies the source for verified safe trajectories, while𝒟fail\\mathcal\{D\}\_\{\\text\{fail\}\}together with controlled error injection on𝒟succ\\mathcal\{D\}\_\{\\text\{succ\}\}yields failure trajectories with decisive\-error annotations\. Source\-level details are deferred to Appendix[B\.3](https://arxiv.org/html/2605.08715#A2.SS3)\.

![Refer to caption](https://arxiv.org/html/2605.08715v1/x2.png)Figure 2:Overview ofAgentForesight\.\(a\)TheAFTraj\-22Kconstruction pipeline collects trajectories from off\-the\-shelf multi\-agent systems across multiple domains, retains successful runs through a strict filtering pipeline, and produces failure runs via decisive\-error injection and multi\-judge voting verification\.\(b\)A*coarse\-to\-fine*training recipe that first equips the auditor with a risk\-anticipation prior on adjacent safe/unsafe prefix pairs, then sharpens it into precise step\-level localization under a three\-axis reward targeting*what*,*where*,*who*of an audit verdict\. The resulting auditor issues per\-stepContinueorAlarmverdicts on prefix evidence\.
##### Curating Verified Safe Trajectories\.

A trajectoryτ∈𝒟succ\\tau\\in\\mathcal\{D\}\_\{\\text\{succ\}\}is not automatically safe111We use safe to refer to trajectories that complete successfully without containing any step whose correction would have changed the outcome \(Definition[2\.1](https://arxiv.org/html/2605.08715#S2.Thmtheorem1)\), which is distinct from the safety/alignment usage in RLHF literature\.at every step in the sense of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2), since a silent intermediate error may be masked by a downstream agent’s recovery, or by permissive evaluation criteria that flipΩ​\(τ\)\\Omega\(\\tau\)to11despite locally degenerate turns\. Treating such trajectories as positive supervision would teach the auditor to issueContinueon prefixes that contain warning signs it should learn to flag, directly undermining the prefix\-restricted supervision online auditing demands\. We therefore apply a three\-stage filtering pipeline of binary predicatesϕj:𝒯→\{0,1\}\\phi\_\{j\}:\\mathcal\{T\}\\to\\\{0,1\\\}to retain only trajectories that are safe at every prefix,

𝒟safe=\{τ∈𝒟succ\|ϕj​\(τ\)=1,∀j∈ℱ\},ℱ=\{outcome,integrity,coherence\},\\mathcal\{D\}\_\{\\text\{safe\}\}=\\bigl\\\{\\tau\\in\\mathcal\{D\}\_\{\\text\{succ\}\}\\;\\bigm\|\\;\\phi\_\{j\}\(\\tau\)=1,\\;\\forall j\\in\\mathcal\{F\}\\bigr\\\},\\quad\\mathcal\{F\}=\\\{\\text\{outcome\},\\,\\text\{integrity\},\\,\\text\{coherence\}\\\},\(6\)whereϕoutcome\\phi\_\{\\text\{outcome\}\}enforces strict outcome equivalence against the reference,ϕintegrity\\phi\_\{\\text\{integrity\}\}rejects trajectories with any invalid tool invocation, andϕcoherence\\phi\_\{\\text\{coherence\}\}verifies that each turn remains aligned with the declared sub\-goal under an LLM judge\. Eachτ∈𝒟safe\\tau\\in\\mathcal\{D\}\_\{\\text\{safe\}\}is treated as carrying the label\(Safe,∅\)\(\\textsc\{Safe\},\\emptyset\)at every prefixτ0:k\\tau\_\{0:k\}, providing the positive\-class supervision absent from prior failure\-attribution corpora\.

##### Constructing Failure Trajectories with Decisive Error Annotations\.

The training signal\(τ,k∗,a∗\)\(\\tau,k^\{\*\},a^\{\*\}\)required by online auditing demands both the existence of a verified failure and step\-level localization of its decisive error, neither of which is reliably extractable from naive sources\. We obtain this signal from two complementary streams that together cover distinct failure distributions\. The*constructive stream*operates on safe trajectories with by\-construction ground truth, while the*diagnostic stream*operates on naturally\-failed trajectories whose decisive step must be discovered\. Building on the paradigm of\[[62](https://arxiv.org/html/2605.08715#bib.bib62)\], the*constructive stream*applies controlled*decisive error injection*to verified safe trajectories, mirroring the counterfactual structure of Definition[2\.1](https://arxiv.org/html/2605.08715#S2.Thmtheorem1)\. Starting fromτ∈𝒟safe\\tau\\in\\mathcal\{D\}\_\{\\text\{safe\}\}, we sample an injection stepkinj∈\{1,…,\|τ\|−2\}k\_\{\\text\{inj\}\}\\in\\\{1,\\ldots,\|\\tau\|\{\-\}2\\\}and a fault categoryc∈𝒞c\\in\\mathcal\{C\}, generate a faulty turnt~kinj∼πfault\(⋅∣τ0:kinj−1,c\)\\tilde\{t\}\_\{k\_\{\\text\{inj\}\}\}\\sim\\pi\_\{\\text\{fault\}\}\(\\cdot\\mid\\tau\_\{0:k\_\{\\text\{inj\}\}\-1\},c\), and re\-roll the systemℳ\\mathcal\{M\}forward to obtain

τ~=τ0:kinj−1⊕t~kinj⊕τ~\>kinj,τ~\>kinj∼Roll​\(ℳ,τ0:kinj−1⊕t~kinj\),\\tilde\{\\tau\}\\;=\\;\\tau\_\{0:k\_\{\\text\{inj\}\}\-1\}\\,\\oplus\\,\\tilde\{t\}\_\{k\_\{\\text\{inj\}\}\}\\,\\oplus\\,\\tilde\{\\tau\}\_\{\>k\_\{\\text\{inj\}\}\},\\qquad\\tilde\{\\tau\}\_\{\>k\_\{\\text\{inj\}\}\}\\sim\\mathrm\{Roll\}\\bigl\(\\mathcal\{M\},\\,\\tau\_\{0:k\_\{\\text\{inj\}\}\-1\}\\oplus\\tilde\{t\}\_\{k\_\{\\text\{inj\}\}\}\\bigr\),\(7\)whereπfault\\pi\_\{\\text\{fault\}\}is realized by complementary turn\-rewriting and live\-replay variants suited to short\-horizon and tool\-augmented domains respectively\. A post\-injection check rejects candidates whoseΩ​\(τ~\)=1\\Omega\(\\tilde\{\\tau\}\)=1\(downstream agents recovered\) or whose targeted turn was not actually modified, after which each accepted sample is admitted to𝒟failinj\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}with verified label\(k∗,a∗\)=\(kinj,akinj\)\(k^\{\*\},a^\{\*\}\)=\(k\_\{\\text\{inj\}\},a\_\{k\_\{\\text\{inj\}\}\}\)\. The*diagnostic stream*operates onτ∈𝒟fail\\tau\\in\\mathcal\{D\}\_\{\\text\{fail\}\}, where the decisive error occurs at some unknown step inτ\\taubut must be localized\. We adopt a propose\-and\-verify ensemble designed to be strictly more conservative than single\-round majority voting\. A pool ofPPproposer calls returns candidate steps and their responsible agents, and each unique candidate is then re\-checked byVVverifier calls along four binary criteria\(sexists,ssubstantive,sdecisive,searliest\)\(s\_\{\\text\{exists\}\},s\_\{\\text\{substantive\}\},s\_\{\\text\{decisive\}\},s\_\{\\text\{earliest\}\}\)\. A candidate is admitted if and only if its support count, i\.e\., the number of verifiers under which all four criteria hold, exceeds the majority threshold,

𝒟failnat=\{\(τ,kcand,akcand\)\|∑j=1V∏rsr\(j\)≥⌊V/2⌋\+1\},\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{nat\}\}=\\Bigl\\\{\(\\tau,\\,k\_\{\\text\{cand\}\},\\,a\_\{k\_\{\\text\{cand\}\}\}\)\\;\\Big\|\\;\\textstyle\\sum\_\{j=1\}^\{V\}\\prod\_\{r\}s\_\{r\}^\{\(j\)\}\\,\\geq\\,\\lfloor V/2\\rfloor\+1\\Bigr\\\},\(8\)whererrranges over the four criteria above; the highest\-strict\-support candidate is then selected perτ\\tau, with ties broken by verifier confidence\. The final unsafe pool combines the two streams,𝒟unsafe=𝒟failinj∪𝒟failnat\\mathcal\{D\}\_\{\\text\{unsafe\}\}=\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}\\,\\cup\\,\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{nat\}\}, providing the step\-level decisive\-error supervision required by online auditing\.

##### Curated Dataset\.

Pooling the verified\-safe and verified\-unsafe streams constructed above yields a unified corpus that supplies\(Safe,∅\)\(\\textsc\{Safe\},\\emptyset\)labels on every prefix of safe trajectories𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}, and\(k∗,a∗\)\(k^\{\*\},a^\{\*\}\)labels at the decisive step of unsafe trajectories𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}\. We refer to this corpus asAFTraj\-22K, comprising∼\\sim2\.3K high\-fidelity annotated safe and unsafe trajectories, formally𝒟AFTraj=𝒟safe∪𝒟unsafe\\mathcal\{D\}\_\{\\text\{\{AFTraj\}\}\}\\;=\\;\\mathcal\{D\}\_\{\\text\{safe\}\}\\,\\cup\\,\\mathcal\{D\}\_\{\\text\{unsafe\}\}\. Detailed composition statistics and qualitative samples are presented in Appendix[B\.1](https://arxiv.org/html/2605.08715#A2.SS1)and[F](https://arxiv.org/html/2605.08715#A6)\.

### 3\.2Training*AgentForesight*\-7B: A Coarse\-to\-Fine Recipe

AlthoughAFTraj\-22Ksupplies the per\-step labels\(k∗,a∗\)\(k^\{\*\},a^\{\*\}\), training a base LLMπθ0\\pi\_\{\\theta\_\{0\}\}to act as an online auditorfonlinef\_\{\\text\{online\}\}faces two coupled obstacles:πθ0\\pi\_\{\\theta\_\{0\}\}has no internal sense of the safe\-versus\-unsafe boundary, and even with that boundary, it still needs to localize the decisive step and responsible agent within the unsafe regime\. A single\-stage policy\-gradient attempt collapses to predictingSafeon every prefix, since the precision\-targeting reward signal is too sparse to establish either capability from scratch\. We therefore train Qwen2\.5\-7B\-Instruct with a*coarse\-to\-fine*recipe that decouples the two: Stage 1 \(BPPO\) equips the auditor with a risk\-anticipation prior at the failure boundary, and Stage 2 sharpens this prior into precise step\-level localization under a three\-axis reward optimized by Group Relative Policy Optimization \(GRPO\)\[[15](https://arxiv.org/html/2605.08715#bib.bib15)\]\. Together the two stages operationalize the prefix\-restricted discrimination and step\-level timeliness demands of online auditing in Section[2](https://arxiv.org/html/2605.08715#S2)\.

##### Stage 1: Failure\-Boundary Alignment\.

For every unsafe trajectory\(τ,k∗,a∗\)∈𝒟unsafe\(\\tau,k^\{\*\},a^\{\*\}\)\\in\\mathcal\{D\}\_\{\\text\{unsafe\}\}, we construct two*boundary\-pair*prompts that differ by exactly one turn at the decisive step: the pre\-boundary promptτ0:k∗−1\\tau\_\{0:k^\{\*\}\-1\}with optimal verdictContinue, and the post\-boundary promptτ0:k∗\\tau\_\{0:k^\{\*\}\}with optimal verdictAlarmon stepk∗k^\{\*\}with responsible agenta∗a^\{\*\}\. The two prompts share a similar form but demand logically reversed verdicts, isolating the failure boundary as the salient signal an auditor must learn\. By learning this sharp transition, the auditor acquires an*implicit risk\-anticipation prior*at the failure boundary: training instills the discriminative signal that separates prefixes immediately preceding a decisive error from those still in the safe regime\. To turn this paired\-prompt contrast into a learning signal, we proposeBoundary\-Pair Preference Optimization \(BPPO\), a preference\-optimization\[[40](https://arxiv.org/html/2605.08715#bib.bib40)\]variant tailored to the boundary\-pair structure with two designs: \(i\) chosen and rejected responses are sampled from base\-policy rollouts and classified by their parsed verdicts, \(ii\) the data are partitioned𝒟pair=𝒟BS∪𝒟BE\\mathcal\{D\}\_\{\\text\{pair\}\}=\\mathcal\{D\}\_\{\\text\{BS\}\}\\cup\\mathcal\{D\}\_\{\\text\{BE\}\}by prompt position and two subsets are optimized jointly,

ℒBPPO​\(πθ;πref\)=−∑c∈\{BS,BE\}𝔼\(x,v∗,v\)∼𝒟c​\[log⁡σ​\(β​Δθ​\(x,v∗,v\)\)\],\\mathcal\{L\}\_\{\\text\{BPPO\}\}\(\\pi\_\{\\theta\};\\pi\_\{\\text\{ref\}\}\)=\-\\\!\\\!\\\!\\sum\_\{c\\in\\\{\\text\{BS\},\\,\\text\{BE\}\\\}\}\\\!\\\!\\\!\\mathbb\{E\}\_\{\(x,\\,v^\{\*\},\\,v\)\\sim\\mathcal\{D\}\_\{c\}\}\\\!\\Bigl\[\\,\\log\\sigma\\\!\\bigl\(\\beta\\,\\Delta\_\{\\theta\}\(x,v^\{\*\},v\)\\bigr\)\\,\\Bigr\],\(9\)whereΔθ​\(x,v∗,v\)=log⁡πθ​\(v∗∣x\)πref​\(v∗∣x\)−log⁡πθ​\(v∣x\)πref​\(v∣x\)\\Delta\_\{\\theta\}\(x,v^\{\*\},v\)=\\log\\\!\\frac\{\\pi\_\{\\theta\}\(v^\{\*\}\\mid x\)\}\{\\pi\_\{\\text\{ref\}\}\(v^\{\*\}\\mid x\)\}\-\\log\\\!\\frac\{\\pi\_\{\\theta\}\(v\\mid x\)\}\{\\pi\_\{\\text\{ref\}\}\(v\\mid x\)\}is the implicit\-reward margin between the optimal verdictv∗v^\{\*\}and a rejected verdictvv, withπθ​\(v∣x\)\\pi\_\{\\theta\}\(v\\mid x\)denoting the autoregressive probability of producing a response with parsed verdictvvunder the structured\-verdict format of Eq\.[4](https://arxiv.org/html/2605.08715#S2.E4)\. The class\-conditioned datasets carry𝒟BS\\mathcal\{D\}\_\{\\text\{BS\}\}:x=τ0:k∗−1x=\\tau\_\{0:k^\{\*\}\-1\},v∗=Continuev^\{\*\}=\\textsc\{Continue\},v≠Continuev\\neq\\textsc\{Continue\}; and𝒟BE\\mathcal\{D\}\_\{\\text\{BE\}\}:x=τ0:k∗x=\\tau\_\{0:k^\{\*\}\},v∗=\(Alarm,k∗,a∗\)v^\{\*\}=\(\\textsc\{Alarm\},k^\{\*\},a^\{\*\}\),v≠v∗v\\neq v^\{\*\}\. Since the two subsets differ attk∗t\_\{k^\{\*\}\}, jointly minimizingℒBPPO\\mathcal\{L\}\_\{\\text\{BPPO\}\}forcesπθ\\pi\_\{\\theta\}to flip its verdict at the decisive step, yielding BPPO checkpointπθ1\\pi\_\{\\theta\_\{1\}\}as initialization for Stage 2\.

##### Stage 2: Three\-Axis Verdict Sharpening\.

Stage 2 sharpens this risk\-anticipation prior into precise step\-level localization under a reward operationalizing the structural, temporal, and causal dimensions of an audit verdict\. Each rollout produces a structured verdict<think\>⋯\\cdots</think\> <answer\>y^\\hat\{y\}</answer\>, wherey^=\(k^,a^,r^\)\\hat\{y\}=\(\\hat\{k\},\\,\\hat\{a\},\\,\\hat\{r\}\)carries the predicted decisive step, responsible agent, and a brief reason describing what went wrong; forSafeverdicts,k^\\hat\{k\}holds theSafelabel anda^,r^\\hat\{a\},\\hat\{r\}are null\. We score each rollout against ground truthy∗=\(k∗,a∗\)y^\{\*\}=\(k^\{\*\},a^\{\*\}\)along three orthogonal axes corresponding to the*what*,*where*, and*who*\. The structural axis \(*what*\) is a binary format gateG​\(y^\)∈\{0,1\}G\(\\hat\{y\}\)\\in\\\{0,1\\\}that screens schema validity, JSON well\-formedness, and content grounding\. The temporal axis \(*where*\) scores step\-localization fidelity by a gaussian centered at the ground truth step,

rstep​\(k^,k∗\)=exp⁡\(−\(k^−k∗\)22​σstep2\)\.r\_\{\\text\{step\}\}\(\\hat\{k\},k^\{\*\}\)=\\exp\\\!\\left\(\-\\frac\{\(\\hat\{k\}\-k^\{\*\}\)^\{2\}\}\{2\\sigma\_\{\\text\{step\}\}^\{2\}\}\\right\)\.\(10\)The causal axis \(*who*\) scoresragent​\(a^,a∗\)r\_\{\\text\{agent\}\}\(\\hat\{a\},a^\{\*\}\)at full credit on exact role match and a partial credit on mismatch\. The three axes compose into a class\-symmetric reward through a gated form,

R​\(y^,y∗\)=G​\(y^\)⋅Rcontent​\(y^,y∗\)−ηG⋅\(1−G​\(y^\)\),R\(\\hat\{y\},y^\{\*\}\)=G\(\\hat\{y\}\)\\cdot R\_\{\\text\{content\}\}\(\\hat\{y\},y^\{\*\}\)\-\\eta\_\{G\}\\cdot\\bigl\(1\-G\(\\hat\{y\}\)\\bigr\),\(11\)whereRcontentR\_\{\\text\{content\}\}returns\+1\+1for correctly\-flaggedSafeprefixes,ws​rstep\+wa​ragentw\_\{s\}\\,r\_\{\\text\{step\}\}\+w\_\{a\}\\,r\_\{\\text\{agent\}\}\(withws\+wa=1w\_\{s\}\+w\_\{a\}=1\) for correctly\-flaggedAlarmprefixes, and−1\-1for cross\-class errors\. The class\-symmetric±1\\pm 1design prevents class\-bias drift during training, while the soft penalty−ηG\-\\eta\_\{G\}on format violations preserves gradient signal during the early phase before the policy learns the schema\.

We optimizeRRvia GRPO, applying two adaptations specific to our coarse\-to\-fine setup: \(i\) we anchor the reference policyπref\\pi\_\{\\text\{ref\}\}at the Stage 1 BPPO checkpointπθ1\\pi\_\{\\theta\_\{1\}\}so that the KL regularizer pullsπθ\\pi\_\{\\theta\}back toward the risk\-anticipation prior learned in Stage 1; \(ii\) we estimate the KL divergence with the low\-variance k3 estimatorD^KL​\(πθ∥πref\)\\hat\{D\}\_\{\\text\{KL\}\}\(\\pi\_\{\\theta\}\\\|\\pi\_\{\\text\{ref\}\}\)\[[44](https://arxiv.org/html/2605.08715#bib.bib44)\], which is non\-negative by construction and reduces gradient noise on long\-trajectory rollouts\. With these adaptations, the RL objective is formulated as:

ℒGRPO​\(θ\)=−𝔼​\[min⁡\(ρj,t​\(θ\)​Aj,clip​\(ρj,t​\(θ\),1−ϵ,1\+ϵ\)​Aj\)\]\+βKL​D^KL​\(πθ∥πθ1\),\\mathcal\{L\}\_\{\\text\{GRPO\}\}\(\\theta\)=\-\\,\\mathbb\{E\}\\Bigl\[\\min\\\!\\bigl\(\\rho\_\{j,t\}\(\\theta\)\\,A\_\{j\},\\;\\;\\mathrm\{clip\}\(\\rho\_\{j,t\}\(\\theta\),\\,1\-\\epsilon,\\,1\+\\epsilon\)\\,A\_\{j\}\\bigr\)\\Bigr\]\+\\beta\_\{\\text\{KL\}\}\\,\\hat\{D\}\_\{\\text\{KL\}\}\\\!\\bigl\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{\\theta\_\{1\}\}\\bigr\),\(12\)with token\-level importance ratioρj,t​\(θ\)\\rho\_\{j,t\}\(\\theta\)andπref\\pi\_\{\\text\{ref\}\}anchored atπθ1\\pi\_\{\\theta\_\{1\}\}to prevent drift from the risk\-anticipation prior\. Together, the two stages produce*AgentForesight*\-7B, a compact online auditorfonlinef\_\{\\text\{online\}\}that combines a risk\-anticipation prior with precise step\-level localization, issuing per\-stepContinue/Alarmverdicts on unfolding multi\-agent trajectories\.

## 4Experiments

### 4\.1Experimental setups

##### Datasets\.

We evaluate*AgentForesight*\-7B under the strict online auditing protocol of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2)on two datasets\.\(1\)AFTraj\-22Kheld\-out split\.AFTraj\-22Kis curated from off\-the\-shelf multi\-agent frameworks \(AutoGen\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\], MetaGPT\[[17](https://arxiv.org/html/2605.08715#bib.bib17)\], Smolagents\[[41](https://arxiv.org/html/2605.08715#bib.bib41)\]\) on three representative task corpora, namelyMath\(MATH\-500\[[16](https://arxiv.org/html/2605.08715#bib.bib16)\]\),Coding\(HumanEval\+ and MBPP\+\[[29](https://arxiv.org/html/2605.08715#bib.bib29)\]\), andAgentic\(GAIA\[[33](https://arxiv.org/html/2605.08715#bib.bib33)\], HotpotQA\[[57](https://arxiv.org/html/2605.08715#bib.bib57)\]\)\. We hold out15%15\\%ofAFTraj\-22Kunder a trajectory\-grouped split that places each safe trajectory and its injected unsafe variants in the same partition to prevent train\-test leakage, and report per\-domain plus overall results\.\(2\) Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\], an established external benchmark for multi\-agent failure attribution whose trajectories are disjoint fromAFTraj\-22K, evaluating cross\-construction generalization beyond ourAFTraj\-22Kheld\-out test split\.

##### Baselines\.

We compare*AgentForesight*\-7B against three baseline categories\.\(1\) Open\-source small LLMs: Llama\-3\.2\-3B\[[13](https://arxiv.org/html/2605.08715#bib.bib13)\], Gemma\-3\-4B\[[10](https://arxiv.org/html/2605.08715#bib.bib10)\], Qwen2\.5\-7B\-Instruct, Qwen3\-8B\[[56](https://arxiv.org/html/2605.08715#bib.bib56)\], Qwen3\-32B\.\(2\) Proprietary LLMs: GPT\-4\.1\[[36](https://arxiv.org/html/2605.08715#bib.bib36)\], Gemini\-3\-Flash\[[7](https://arxiv.org/html/2605.08715#bib.bib7)\], Claude\-Haiku\-4\.5\[[1](https://arxiv.org/html/2605.08715#bib.bib1)\], DeepSeek\-V4\-Flash, DeepSeek\-V4\-Pro\[[6](https://arxiv.org/html/2605.08715#bib.bib6)\]\.\(3\) Methodological baselines: four paradigms instantiated on the same Qwen2\.5\-7B\-Instruct to isolate paradigm effects from backbone capability, including uncertainty quantification \(Perplexity\-7B\[[8](https://arxiv.org/html/2605.08715#bib.bib8)\]\), tree\-search prompting \(ToT\-7B\[[58](https://arxiv.org/html/2605.08715#bib.bib58)\]\), self\-reflection \(Reflexion\-7B\[[48](https://arxiv.org/html/2605.08715#bib.bib48)\]\), and post\-hoc failure attribution \(AgentDebug\-7B\[[67](https://arxiv.org/html/2605.08715#bib.bib67)\]\)\. All baselines except AgentDebug\-7B follow our online auditing protocol; AgentDebug\-7B observes the full completed trajectory and serves as a reference for the gap between post\-hoc attribution and online auditing\.

##### Metrics\.

Online auditing requires exact localization of the first decisive error rather than mere binary detection\. We adopt two complementary metrics: Exact\-Step F1 \(Exact\-F1↑\\uparrow\) is the harmonic mean of step\-level recall and precision on decisive\-step predictions, penalizing both missed errors and wrong localizations\. Absolute Step Shift \(ASS↓\\downarrow\) averages\|k^−k∗\|\|\\hat\{k\}\-k^\{\*\}\|over detected unsafe trajectories, remaining informative when alarms miss the exact step\. See Appendix[B\.2](https://arxiv.org/html/2605.08715#A2.SS2)for detailed definitions\.

##### Implementation Details\.

We instantiate*AgentForesight*\-7B from Qwen2\.5\-7B\-Instruct\[[56](https://arxiv.org/html/2605.08715#bib.bib56)\]and train it onAFTraj\-22Kfollowing the coarse\-to\-fine recipe of Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)\. Training uses verl\[[46](https://arxiv.org/html/2605.08715#bib.bib46)\]on2×2\{\\times\}NVIDIA H200 GPUs with vLLM\-accelerated rollouts\.Stage 1uses BPPO withβ=0\.1\\beta=0\.1\(cf\. Eq\.[9](https://arxiv.org/html/2605.08715#S3.E9)\), learning rate5×10−75\{\\times\}10^\{\-7\}, and33epochs;Stage 2uses GRPO with group sizeG=8G=8, KL coefficientβKL=10−3\\beta\_\{\\text\{KL\}\}=10^\{\-3\}, and learning rate10−610^\{\-6\}\(cf\. Eq\.[12](https://arxiv.org/html/2605.08715#S3.E12)\)\. During evaluation, we follow a*strict step\-by\-step incremental walk*: the auditor is queried at every prefixτ0:k\\tau\_\{0:k\}with greedy decoding, and both safe and unsafe trajectories are walked through their full length to surface any false alarm\.

We provide detailed experimental setups in Appendix[B](https://arxiv.org/html/2605.08715#A2)\.

Table 1:Online auditing evaluation on theAFTraj\-22K\. Both safe and unsafe samples are evaluated under the online auditing protocol of Section[2](https://arxiv.org/html/2605.08715#S2)\.Bold==best results,underline==second\-best results\.†AgentDebug\-7B detects zero unsafe trajectories with no step\-shift samples to average\. We use "—" to mark its ASS as undefined\.MathCodingAgenticOverallMethodExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowOpen\-Source LLMsLlama3\.2\-3B8\.144\.8621\.052\.9416\.132\.3014\.413\.38Gemma3\-4B1\.155\.7812\.904\.598\.293\.046\.924\.37Qwen2\.5\-7B\-Instruct10\.393\.8038\.202\.2614\.002\.9621\.052\.75Qwen3\-8B21\.954\.6527\.852\.5733\.641\.5728\.362\.65Qwen3\-32B18\.634\.0720\.002\.8340\.001\.5926\.912\.82Proprietary LLMsGPT\-4\.124\.393\.9510\.813\.5040\.681\.2927\.432\.67Gemini\-3\-Flash40\.522\.4819\.422\.7326\.091\.7029\.742\.19Claude\-Haiku\-4\.519\.754\.1223\.912\.8031\.951\.6925\.532\.85DeepSeek\-V4\-Flash42\.772\.7838\.101\.2732\.531\.4237\.651\.94DeepSeek\-V4\-Pro50\.342\.6049\.320\.9641\.771\.3146\.561\.77Qwen2\.5\-7B\-Instruct basedPerplexity\-7B\[[8](https://arxiv.org/html/2605.08715#bib.bib8)\]2\.314\.0926\.562\.1916\.572\.5014\.113\.02ToT\-7B\[[58](https://arxiv.org/html/2605.08715#bib.bib58)\]20\.384\.407\.024\.0624\.842\.1318\.523\.39Reflexion\-7B\[[48](https://arxiv.org/html/2605.08715#bib.bib48)\]16\.574\.399\.524\.5039\.131\.4423\.383\.17AgentDebug\-7B†\[[67](https://arxiv.org/html/2605.08715#bib.bib67)\]0\.00—28\.574\.052\.821\.009\.633\.76*AgentForesight\-7B*\(ours\)77\.360\.9678\.870\.1848\.700\.5466\.440\.59

### 4\.2Main results

##### Performance comparison onAFTraj\-22K\.

Table[1](https://arxiv.org/html/2605.08715#S4.T1)reports the performance comparison onAFTraj\-22Kacross three domains\. Overall,*AgentForesight*\-7B reaches66\.4466\.44Exact\-F1,19\.8819\.88points above the strongest proprietary baseline DeepSeek\-V4\-Pro, and tightens overall ASS from1\.771\.77to0\.590\.59\(3×3\\times\)\. Per\-domain,*AgentForesight*\-7B performs better on both Exact\-F1 and ASS in every domain, with the largest Exact\-F1 gains on Math \(77\.3677\.36vs\.50\.3450\.34\) and Coding \(78\.8778\.87vs\.49\.3249\.32\)\. The coarse\-to\-fine recipe of Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)lifts the Qwen2\.5\-7B\-Instruct backbone by3\.16×3\.16\\timeson Exact\-F1, while AgentDebug\-7B, the post\-hoc reference with full\-trajectory hindsight, ranks lowest at9\.639\.63overall Exact\-F1\. These results show that AgentForesight’s gains stem from its coarse\-to\-fine recipe tailored to online auditing, not from scaling backbones or re\-purposing existing post\-hoc attributors\.

Table 2:Online auditing evaluation on the Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]benchmark\. All evaluated under the online auditing protocol of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2)ModelStep\-AccAgent\-AccASSLlama3\.2\-3B28\.5747\.622\.57Gemma3\-4B6\.9818\.603\.09Qwen2\.5\-7B\-Instruct36\.5958\.542\.41Qwen3\-8B29\.4155\.882\.79GPT\-4\.138\.1066\.672\.38Gemini\-3\-Flash32\.5653\.492\.47DeepSeek\-V4\-Flash37\.2165\.122\.35*AgentForesight\-7B*\(ours\)57\.6973\.081\.62

##### Generalization to external benchmark\.

*AgentForesight*\-7B further transfers to the external Who&When benchmark \(Table[2](https://arxiv.org/html/2605.08715#S4.T2)\), whose trajectories come from multi\-agent frameworks disjoint fromAFTraj\-22K\. It leads all three metrics, exceeding the strongest baseline GPT\-4\.1 by19\.5919\.59points on Step\-Acc and6\.416\.41on Agent\-Acc, and reducing ASS from2\.352\.35\(DeepSeek\-V4\-Flash\) to1\.621\.62\. Since these trajectories are entirely unseen at training time, the transfer indicates that*AgentForesight*\-7B captures online auditing signal that generalizes beyondAFTraj\-22K’s framework choices rather than overfitting to its curation artifacts\.

### 4\.3Ablation and Further Analysis

##### Stage\-wise contributions to AgentForesight performance\.

Figure[3](https://arxiv.org/html/2605.08715#S4.F3)ablates the two stages of our coarse\-to\-fine recipe on the Qwen2\.5\-7B\-Instruct base\. Each stage individually lifts Exact\-F1 from the base21\.121\.1\(Stage 1 to35\.635\.6, Stage 2 to50\.450\.4\), and combining them yields66\.466\.4for*AgentForesight*\-7B, exceeding either single stage by at least1616points\. The breakdown reveals a clear division of labor: Stage 2 alone already handles Math \(63\.663\.6\) and Coding \(72\.772\.7\) where decisive errors are sharply localizable, but degrades on Agentic \(19\.019\.0, below Stage 1 alone at31\.631\.6\) where the failure boundary is harder to discriminate\. With the risk\-anticipation prior of Stage 1 in the full recipe, Agentic domain recovers to48\.7048\.70\. This validates the predict\-then\-localize coupling, with Stage 1 establishing a learnable failure boundary that Stage 2 sharpens to step\-level precision\.

##### Deployment trade\-off between false alarms and step localization\.

A deployable online auditor must place alarms accurately while rarely interrupting safe trajectories\. Figure[4](https://arxiv.org/html/2605.08715#S4.F4)traces this trade\-off, plotting Step Accuracy \(on𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}\) against False Alarm Rate \(FAR on𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}, fraction of safe trajectories with any raised alarm\)\. We further mark a deployable region at FAR≤20%\\leq 20\\%and Step\-Acc≥50%\\geq 50\\%, the operating point at which downstream triage or recovery routing remains tractable\. Among the ten auditors compared, only*AgentForesight*\-7B \(FAR=2\.4%=2\.4\\%, Step\-Acc=59\.5%=59\.5\\%\) lies inside this region\. The strongest proprietary baseline on both axes, DeepSeek\-V4\-Pro \(FAR=43\.2%=43\.2\\%, Step\-Acc=54\.0%=54\.0\\%\), falls just outside, while other proprietary judges and open\-source 7–8B base models concentrate at high FAR with mid Step\-Acc and the 3–4B LLMs collapse to near\-universal false alarms\. The gap is consistent with our coarse\-to\-fine recipe, where Stage 1’s risk\-anticipation prior suppresses spurious alarms and Stage 2’s three\-axis reward sharpens alarm placement\.

![Refer to caption](https://arxiv.org/html/2605.08715v1/x3.png)
Figure 3:Ablation of the two\-stage*coarse\-to\-fine*recipe onAFTraj\-22K, comparing\+\+Stage 1,\+\+Stage 2, and the full two\-stage*AgentForesight*\-7B\.![Refer to caption](https://arxiv.org/html/2605.08715v1/x4.png)
Figure 4:Deployment trade\-off across all auditors onAFTraj\-22Kwith False Alarm Rate↓\\downarrow\(𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}\) vs\. Step Accuracy↑\\uparrow\(𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}\) and a shaded deployable region\.

### 4\.4Case Study

![Refer to caption](https://arxiv.org/html/2605.08715v1/x5.png)Figure 5:Case study of*online auditing*, comparing predictions from DeepSeek\-V4\-Pro, Gemini\-3\-Flash, and*AgentForesight*\-7B\.##### Where strong baselines miss or mislocate\.

Figure[5](https://arxiv.org/html/2605.08715#S4.F5)shows an agentic trajectory whose decisive error commits at Step 3, where the search\_agent returns the wrong townHorwichinstead of the gold answerBoltonand the Manager propagates it to completion\.*AgentForesight*\-7B alone returns Step 5 with search\_agent as the responsible agent\. The two strong proprietary baselines fail in opposite directions\. Gemini\-3\-Flash flags Step 2 on the Manager’s planning thought, while DeepSeek\-V4\-Pro returnsSafeafter monitoring the whole trajectory\. This shows that effective online auditing demands both refraining from premature alarms on safe prefixes and detecting decisive errors that strong baselines miss entirely\.

## 5Conclusion

In this paper, we presentAgentForesight, an*online auditing*perspective on agentic failure analysis, recasting it from post\-hoc diagnosis of completed trajectories into a per\-step continue\-or\-alarm decision on each unfolding prefix\. Building on this view, we introduceAFTraj\-22K, a curated corpus pairing strictly filtered safe runs with multi\-judge verified*decisive error*annotations across Coding, Math, and Agentic domains\. We develop*AgentForesight*\-7B, a compact online auditor trained via a*coarse\-to\-fine*reinforcement learning recipe that first equips it with a risk\-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs and then sharpens this prior into precise step\-level localization under a three\-axis reward jointly targeting the*what*,*where*, and*who*of an audit verdict\. Extensive experiments on bothAFTraj\-22Kand the external Who&When benchmark validate the effectiveness of*AgentForesight*\-7B\. Beyond advancing online auditing, our framework paves the way for runtime safeguards that intervene before downstream propagation locks in the failure, marking a step toward deployment\-ready oversight of multi\-agent systems\.

## References

- \[1\]Anthropic\.Introducing claude haiku 4\.5\.[https://www\.anthropic\.com/news/claude\-haiku\-4\-5](https://www.anthropic.com/news/claude-haiku-4-5), October 2025\.Accessed: 2026\-05\-02\.
- \[2\]Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, and David Farhi\.Monitoring reasoning models for misbehavior and the risks of promoting obfuscation\.arXiv preprint arXiv:2503\.11926, 2025\.
- \[3\]Mert Cemri, Melissa Z Pan, Shuyi Yang, Lakshya A Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, et al\.Why do multi\-agent llm systems fail?arXiv preprint arXiv:2503\.13657, 2025\.
- \[4\]Chi\-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu\.Chateval: Towards better llm\-based evaluators through multi\-agent debate\.arXiv preprint arXiv:2308\.07201, 2023\.
- \[5\]Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al\.Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168, 2021\.
- \[6\]DeepSeek\-AI\.Deepseek\-v4: Towards highly efficient million\-token context intelligence, 2026\.
- \[7\]Tulsee Doshi\.Gemini 3 flash: Frontier intelligence built for speed\.[https://blog\.google/products\-and\-platforms/products/gemini/gemini\-3\-flash/](https://blog.google/products-and-platforms/products/gemini/gemini-3-flash/), December 2025\.Google Blog\. Accessed: 2026\-05\-01\.
- \[8\]Ekaterina Fadeeva, Roman Vashurin, Akim Tsvigun, Artem Vazhentsev, Sergey Petrakov, Kirill Fedyanin, Daniil Vasilev, Elizaveta Goncharova, Alexander Panchenko, Maxim Panov, et al\.Lm\-polygraph: Uncertainty estimation for language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 446–461, 2023\.
- \[9\]Lang Feng, Zhenghai Xue, Tingcong Liu, and Bo An\.Group\-in\-group policy optimization for llm agent training\.arXiv preprint arXiv:2505\.10978, 2025\.
- \[10\]Gemma Team\.Gemma 3 technical report\.arXiv preprint arXiv:2503\.19786, March 2025\.
- \[11\]Alireza Ghafarollahi and Markus J Buehler\.Sciagents: automating scientific discovery through bioinspired multi\-agent intelligent graph reasoning\.Advanced Materials, 37\(22\):2413523, 2025\.
- \[12\]Ali Essam Ghareeb, Benjamin Chang, Ludovico Mitchener, Angela Yiu, Caralyn J Szostkiewicz, Jon M Laurent, Muhammed T Razzak, Andrew D White, Michaela M Hinks, and Samuel G Rodriques\.Robin: A multi\-agent system for automating scientific discovery\.arXiv preprint arXiv:2505\.13400, 2025\.
- \[13\]Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al\-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783, 2024\.
- \[14\]Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al\.A survey on llm\-as\-a\-judge\.The Innovation, 2024\.
- \[15\]Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al\.Deepseek\-r1: Incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948, 2025\.
- \[16\]Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\.Measuring mathematical problem solving with the math dataset\.arXiv preprint arXiv:2103\.03874, 2021\.
- \[17\]Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, et al\.Metagpt: Meta programming for a multi\-agent collaborative framework\.InThe twelfth international conference on learning representations, 2023\.
- \[18\]Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou\.Large language models cannot self\-correct reasoning yet\.arXiv preprint arXiv:2310\.01798, 2023\.
- \[19\]Zhenlan Ji, Daoyuan Wu, Pingchuan Ma, Zongjie Li, and Shuai Wang\.Testing and understanding erroneous planning in llm agents through synthesized user inputs\.arXiv preprint arXiv:2404\.17833, 2024\.
- \[20\]Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan\.Swe\-bench: Can language models resolve real\-world github issues?arXiv preprint arXiv:2310\.06770, 2023\.
- \[21\]Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Arik, Dong Wang, Hamed Zamani, and Jiawei Han\.Search\-r1: Training llms to reason and leverage search engines with reinforcement learning\.arXiv preprint arXiv:2503\.09516, 2025\.
- \[22\]Neil Kale, Chen Bo Calvin Zhang, Kevin Zhu, Ankit Aich, Paula Rodriguez, Scale Red Team, Christina Q Knight, and Zifan Wang\.Reliable weak\-to\-strong monitoring of llm agents\.arXiv preprint arXiv:2508\.19461, 2025\.
- \[23\]Jonathan Kutasov, Yuqi Sun, Paul Colognese, Teun van der Weij, Linda Petrini, Chen Bo Calvin Zhang, John Hughes, Xiang Deng, Henry Sleight, Tyler Tracy, et al\.Shade\-arena: Evaluating sabotage and monitoring in llm agents\.arXiv preprint arXiv:2506\.15740, 2025\.
- \[24\]Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem\.Camel: Communicative agents for" mind" exploration of large language model society\.Advances in neural information processing systems, 36:51991–52008, 2023\.
- \[25\]Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, et al\.Atbench: A diverse and realistic trajectory benchmark for long\-horizon agent safety\.arXiv preprint arXiv:2604\.02022, 2026\.
- \[26\]Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, and Pan Lu\.In\-the\-flow agentic system optimization for effective planning and tool use\.arXiv preprint arXiv:2510\.05592, 2025\.
- \[27\]Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe\.Let’s verify step by step\.InThe twelfth international conference on learning representations, 2023\.
- \[28\]Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al\.Advances and challenges in foundation agents: From brain\-inspired intelligence to evolutionary, collaborative, and safe systems\.arXiv preprint arXiv:2504\.01990, 2025\.
- \[29\]Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang\.Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation\.Advances in neural information processing systems, 36:21558–21572, 2023\.
- \[30\]Shuo Liu, Zeyu Liang, Xueguang Lyu, and Christopher Amato\.Llm collaboration with multi\-agent reinforcement learning\.InProceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 32150–32158, 2026\.
- \[31\]Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu\.G\-eval: Nlg evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 conference on empirical methods in natural language processing, pages 2511–2522, 2023\.
- \[32\]Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al\.Self\-refine: Iterative refinement with self\-feedback\.Advances in neural information processing systems, 36:46534–46594, 2023\.
- \[33\]Grégoire Mialon, Clémentine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom\.Gaia: a benchmark for general ai assistants\.InThe Twelfth International Conference on Learning Representations, 2023\.
- \[34\]Ning Miao, Yee Whye Teh, and Tom Rainforth\.Selfcheck: Using llms to zero\-shot check their own step\-by\-step reasoning\.arXiv preprint arXiv:2308\.00436, 2023\.
- \[35\]Kaiwen Ning, Jiachi Chen, Jingwen Zhang, Wei Li, Zexu Wang, Yuming Feng, Weizhe Zhang, and Zibin Zheng\.Defining and detecting the defects of large language model\-based autonomous agents\.IEEE Transactions on Software Engineering, 2026\.
- \[36\]OpenAI\.Gpt\-5 system card\.[https://openai\.com/index/gpt\-5\-system\-card/](https://openai.com/index/gpt-5-system-card/), August 2025\.System card\. Accessed: 2026\-05\-01\.
- \[37\]Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al\.Training language models to follow instructions with human feedback\.Advances in neural information processing systems, 35:27730–27744, 2022\.
- \[38\]Charles Packer, Vivian Fang, Shishir\_G Patil, Kevin Lin, Sarah Wooders, and Joseph\_E Gonzalez\.Memgpt: towards llms as operating systems\.2023\.
- \[39\]Chen Qian, Peng Wang, Dongrui Liu, Junyao Yang, Dadi Guo, Ling Tang, Jilin Mei, Qihan Ren, Shuai Shao, Yong Liu, et al\.The why behind the action: Unveiling internal drivers via agentic attribution\.arXiv preprint arXiv:2601\.15075, 2026\.
- \[40\]Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn\.Direct preference optimization: Your language model is secretly a reward model\.Advances in neural information processing systems, 36:53728–53741, 2023\.
- \[41\]Aymeric Roucher, A Villanova del Moral, Thomas Wolf, Leandro von Werra, and Erik Kaunismäki\.smolagents: A smol library to build great agentic systems\.Hugging Face, 2025\.
- \[42\]Timo Schick, Jane Dwivedi\-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom\.Toolformer: Language models can teach themselves to use tools\.Advances in neural information processing systems, 36:68539–68551, 2023\.
- \[43\]Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni, Axel Højmark, Felix Hofstätter, Jérémy Scheurer, Alexander Meinke, Jason Wolfe, Teun van der Weij, Alex Lloyd, et al\.Stress testing deliberative alignment for anti\-scheming training\.arXiv preprint arXiv:2509\.15541, 2025\.
- \[44\]John Schulman\.Approximating KL divergence\.[http://joschu\.net/blog/kl\-approx\.html](http://joschu.net/blog/kl-approx.html), 2020\.Blog post\.
- \[45\]Shuai Shao, Qihan Ren, Chen Qian, Boyi Wei, Dadi Guo, Jingyi Yang, Xinhao Song, Linfeng Zhang, Weinan Zhang, Dongrui Liu, et al\.Your agent may misevolve: Emergent risks in self\-evolving llm agents\.arXiv preprint arXiv:2509\.26354, 2025\.
- \[46\]Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu\.Hybridflow: A flexible and efficient rlhf framework\.InProceedings of the Twentieth European Conference on Computer Systems, pages 1279–1297, 2025\.
- \[47\]Zeru Shi, Kai Mei, Mingyu Jin, Yongye Su, Chaoji Zuo, Wenyue Hua, Wujiang Xu, Yujie Ren, Zirui Liu, Mengnan Du, et al\.From commands to prompts: Llm\-based semantic file system for aios\.InInternational Conference on Learning Representations, volume 2025, pages 33108–33131, 2025\.
- \[48\]Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: Language agents with verbal reinforcement learning\.Advances in neural information processing systems, 36:8634–8652, 2023\.
- \[49\]Yoo Yeon Sung, Hannah Kim, and Dan Zhang\.Verila: A human\-centered evaluation framework for interpretable verification of llm agent failures\.arXiv preprint arXiv:2503\.12651, 2025\.
- \[50\]Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui\.Math\-shepherd: Verify and reinforce llms step\-by\-step without human annotations\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 9426–9439, 2024\.
- \[51\]Xingyao Wang, Boxuan Li, Yufan Song, Frank F Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al\.Openhands: An open platform for ai software developers as generalist agents\.arXiv preprint arXiv:2407\.16741, 2024\.
- \[52\]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al\.Chain\-of\-thought prompting elicits reasoning in large language models\.Advances in neural information processing systems, 35:24824–24837, 2022\.
- \[53\]Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, et al\.Autogen: Enabling next\-gen llm applications via multi\-agent conversations\.InFirst conference on language modeling, 2024\.
- \[54\]Zhiheng Xi, Jixuan Huang, Chenyang Liao, Baodai Huang, Honglin Guo, Jiaqi Liu, Rui Zheng, Junjie Ye, Jiazheng Zhang, Wenxiang Chen, et al\.Agentgym\-rl: Training llm agents for long\-horizon decision making through multi\-turn reinforcement learning\.arXiv preprint arXiv:2509\.08755, 2025\.
- \[55\]Zhiheng Xi, Chenyang Liao, Guanyu Li, Zhihao Zhang, Wenxiang Chen, Binghai Wang, Senjie Jin, Yuhao Zhou, Jian Guan, Wei Wu, et al\.Agentprm: Process reward models for llm agents via step\-wise promise and progress\.InProceedings of the ACM Web Conference 2026, pages 4184–4195, 2026\.
- \[56\]An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\.Qwen3 technical report\.arXiv preprint arXiv:2505\.09388, 2025\.
- \[57\]Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning\.Hotpotqa: A dataset for diverse, explainable multi\-hop question answering\.InProceedings of the 2018 conference on empirical methods in natural language processing, pages 2369–2380, 2018\.
- \[58\]Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan\.Tree of thoughts: Deliberate problem solving with large language models\.Advances in neural information processing systems, 36:11809–11822, 2023\.
- \[59\]Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\.React: Synergizing reasoning and acting in language models\.InThe eleventh international conference on learning representations, 2022\.
- \[60\]Boxuan Zhang, Yi Yu, Jiaxuan Guo, and Jing Shao\.Dive into the agent matrix: A realistic evaluation of self\-replication risk in llm agents\.arXiv preprint arXiv:2509\.25302, 2025\.
- \[61\]Boxuan Zhang and Ruqi Zhang\.Cot\-uq: Improving response\-wise uncertainty quantification in llms with chain\-of\-thought\.InFindings of the Association for Computational Linguistics: ACL 2025, pages 26114–26133, 2025\.
- \[62\]Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou, Kun Wang, and Shuicheng Yan\.Agentracer: Who is inducing failure in the llm agentic systems?arXiv preprint arXiv:2509\.03312, 2025\.
- \[63\]Shaokun Zhang, Ming Yin, Jieyu Zhang, Jiale Liu, Zhiguang Han, Jingyang Zhang, Beibin Li, Chi Wang, Huazheng Wang, Yiran Chen, et al\.Which agent causes task failures and when? on automated failure attribution of llm multi\-agent systems\.arXiv preprint arXiv:2505\.00212, 2025\.
- \[64\]Chujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin, Keming Lu, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin\.Processbench: Identifying process errors in mathematical reasoning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), pages 1009–1024, 2025\.
- \[65\]Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al\.Judging llm\-as\-a\-judge with mt\-bench and chatbot arena\.Advances in neural information processing systems, 36:46595–46623, 2023\.
- \[66\]Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, et al\.Webarena: A realistic web environment for building autonomous agents\.arXiv preprint arXiv:2307\.13854, 2023\.
- \[67\]Kunlun Zhu, Zijia Liu, Bingxuan Li, Muxin Tian, Yingxuan Yang, Jiaxun Zhang, Pengrui Han, Qipeng Xie, Fuyang Cui, Weijia Zhang, et al\.Where llm agents fail and how they can learn from failures\.arXiv preprint arXiv:2509\.25370, 2025\.
- \[68\]Mingchen Zhuge, Wenyi Wang, Louis Kirsch, Francesco Faccio, Dmitrii Khizbullin, and Jürgen Schmidhuber\.Language agents as optimizable graphs\.arXiv preprint arXiv:2402\.16823, 2024\.
- \[69\]Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al\.Agent\-as\-a\-judge: Evaluate agents with agents\.arXiv preprint arXiv:2410\.10934, 2024\.

## Appendices

## Reproducibility Statement

To facilitate reproducibility, we summarize the key experimental details and provide the necessary resources in the submitted supplementary materials\.

- •Datasets\.AFTraj\-22Kis constructed by us from publicly available task corpora and is composed of three domains, with Coding sourced from HumanEval\+ and MBPP\+\[[29](https://arxiv.org/html/2605.08715#bib.bib29)\], Math from MATH\-500\[[16](https://arxiv.org/html/2605.08715#bib.bib16)\], and Agentic from GAIA\[[33](https://arxiv.org/html/2605.08715#bib.bib33)\]and HotpotQA\[[57](https://arxiv.org/html/2605.08715#bib.bib57)\], all gathered through off\-the\-shelf multi\-agent frameworks AutoGen\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\], MetaGPT\[[17](https://arxiv.org/html/2605.08715#bib.bib17)\], and Smolagents\[[41](https://arxiv.org/html/2605.08715#bib.bib41)\]with GPT\-5\.4\-mini as the unified backbone \(details in Appendix[B\.1](https://arxiv.org/html/2605.08715#A2.SS1)and Appendix[B\.3](https://arxiv.org/html/2605.08715#A2.SS3)\)\. The external transfer benchmark Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]is publicly released\.
- •Assumption\.Our method follows the*online auditing*setting introduced in Section[2](https://arxiv.org/html/2605.08715#S2), where a trained auditor is queried at every prefixτ0:k\\tau\_\{0:k\}of an unfolding multi\-agent trajectory and must commit to a continue\-or\-alarm verdict using only the visible window\. We keep the online paradigm consistent across all experiments\.
- •Open source\.We include our source code in the submitted supplementary materials\. The release containsAFTraj\-22Kconstruction pipeline and the coarse\-to\-fine recipe for training*AgentForesight*\-7B\.
- •Environment\.Both training stages are conducted on2×2\\timesNVIDIA H200 GPUs using Python 3\.10 and PyTorch 2\.9\. Key hyperparameters of both stages, including learning rate, batch size, group size are reported in Table[4](https://arxiv.org/html/2605.08715#A2.T4)of Appendix[B\.3](https://arxiv.org/html/2605.08715#A2.SS3)\.

## Appendix AAlgorithmic Pipeline

Input:frameworksℳ\\mathcal\{M\}, tasks𝒯\\mathcal\{T\}

Output:

𝒟AFTraj\\mathcal\{D\}\_\{\\text\{AFTraj\}\}
1ex

𝒟succ,𝒟fail←RollOut​\(ℳ,𝒯\)\\mathcal\{D\}\_\{\\text\{succ\}\},\\mathcal\{D\}\_\{\\text\{fail\}\}\\leftarrow\\textsc\{RollOut\}\(\\mathcal\{M\},\\mathcal\{T\}\)
//Trajectory Collection

𝒟safe←\{τ∈𝒟succ:ϕj​\(τ\)=1,∀j∈ℱ\}\\mathcal\{D\}\_\{\\text\{safe\}\}\\leftarrow\\\{\\tau\\in\\mathcal\{D\}\_\{\\text\{succ\}\}:\\phi\_\{j\}\(\\tau\)\{=\}1,\\forall j\\in\\mathcal\{F\}\\\}
//Verified Safe Curation

1ex

𝒟failinj←∅\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}\\leftarrow\\emptyset
//Constructive Stream

for*τ∈𝒟*safe*\\tau\\in\\mathcal\{D\}\_\{\\text\{safe\}\}*do

sample

\(kinj,c\)\(k\_\{\\text\{inj\}\},c\),

τ~←Inject​\(τ,kinj,c\)\\tilde\{\\tau\}\\leftarrow\\textsc\{Inject\}\(\\tau,k\_\{\\text\{inj\}\},c\)
if*Ω​\(τ~\)=0\\Omega\(\\tilde\{\\tau\}\)\{=\}0*thenadd

\(τ~,kinj,akinj\)\(\\tilde\{\\tau\},k\_\{\\text\{inj\}\},a\_\{k\_\{\\text\{inj\}\}\}\)to

𝒟failinj\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}
1ex

𝒟failnat←∅\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{nat\}\}\\leftarrow\\emptyset
//Diagnostic Stream

for*τ∈𝒟*fail*\\tau\\in\\mathcal\{D\}\_\{\\text\{fail\}\}*do

\(k∗,a∗\)←ProposeVerify​\(τ\)\(k^\{\*\},a^\{\*\}\)\\leftarrow\\textsc\{ProposeVerify\}\(\\tau\)
if*accepted*thenadd

\(τ,k∗,a∗\)\(\\tau,k^\{\*\},a^\{\*\}\)to

𝒟failnat\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{nat\}\}
1exreturn

𝒟safe∪𝒟failinj∪𝒟failnat\\mathcal\{D\}\_\{\\text\{safe\}\}\\cup\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}\\cup\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{nat\}\}

Algorithm 1AFTraj\-2KConstruction PipelineAlgorithm[1](https://arxiv.org/html/2605.08715#algorithm1)executes multi\-agent rollouts to obtain\(𝒟succ,𝒟fail\)\(\\mathcal\{D\}\_\{\\text\{succ\}\},\\mathcal\{D\}\_\{\\text\{fail\}\}\), filters𝒟succ\\mathcal\{D\}\_\{\\text\{succ\}\}through the three predicates ofℱ\\mathcal\{F\}to retain𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}, and then exercises both failure streams in parallel\. The verified\-safe pool𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}plays a dual role, serving as positive supervision for the auditor and the scaffold from which the constructive stream injects decisive errors with by\-construction labels, while the diagnostic stream recovers the unknown decisive step on naturally\-failed trajectories via propose\-and\-verify, with their union producing𝒟AFTraj\\mathcal\{D\}\_\{\\text\{AFTraj\}\}\.

Input:

𝒟AFTraj\\mathcal\{D\}\_\{\\text\{AFTraj\}\}, baseπθ0\\pi\_\{\\theta\_\{0\}\}

Output:

πθ\\pi\_\{\\theta\}
1ex

𝒟pair←BuildBoundaryPairs​\(𝒟unsafe\)\\mathcal\{D\}\_\{\\text\{pair\}\}\\leftarrow\\textsc\{BuildBoundaryPairs\}\(\\mathcal\{D\}\_\{\\text\{unsafe\}\}\)
//Stage 1: Failure\-Boundary Alignment

𝒟pref←SampleClassify​\(πθ0,𝒟pair\)\\mathcal\{D\}\_\{\\text\{pref\}\}\\leftarrow\\textsc\{SampleClassify\}\(\\pi\_\{\\theta\_\{0\}\},\\mathcal\{D\}\_\{\\text\{pair\}\}\)
πθ1←arg⁡minπθ⁡ℒBPPO\\pi\_\{\\theta\_\{1\}\}\\leftarrow\\arg\\min\_\{\\pi\_\{\\theta\}\}\\mathcal\{L\}\_\{\\text\{BPPO\}\}on

𝒟pref\\mathcal\{D\}\_\{\\text\{pref\}\}
1ex

πθ,πref←πθ1\\pi\_\{\\theta\},\\pi\_\{\\text\{ref\}\}\\leftarrow\\pi\_\{\\theta\_\{1\}\}
//Stage 2: Three\-Axis Verdict Sharpening

repeat

sample batch

ℬ⊂𝒟AFTraj\\mathcal\{B\}\\subset\\mathcal\{D\}\_\{\\text\{AFTraj\}\}
\{y^j\}j=1G∼πθ\\\{\\hat\{y\}\_\{j\}\\\}\_\{j=1\}^\{G\}\\sim\\pi\_\{\\theta\}for

x∈ℬx\\in\\mathcal\{B\},

sj←R​\(y^j,y∗\)s\_\{j\}\\leftarrow R\(\\hat\{y\}\_\{j\},y^\{\*\}\)
//Eq\.[11](https://arxiv.org/html/2605.08715#S3.E11)

Aj←\(sj−μ\)/\(σ\+ε\)A\_\{j\}\\leftarrow\(s\_\{j\}\{\-\}\\mu\)/\(\\sigma\{\+\}\\varepsilon\), update

πθ\\pi\_\{\\theta\}on

ℒGRPO\\mathcal\{L\}\_\{\\text\{GRPO\}\}
//Eq\.[12](https://arxiv.org/html/2605.08715#S3.E12)

until*converged*

return

πθ\\pi\_\{\\theta\}

Algorithm 2TrainingAgentForesight\-7BAlgorithm[2](https://arxiv.org/html/2605.08715#algorithm2)trains*AgentForesight*\-7B in two stages\. Stage 1 builds boundary pairs𝒟pair\\mathcal\{D\}\_\{\\text\{pair\}\}from𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}, classifies base\-policy rollouts into preferences𝒟pref\\mathcal\{D\}\_\{\\text\{pref\}\}, and minimizes the dual\-subset BPPO loss of Eq\.[9](https://arxiv.org/html/2605.08715#S3.E9)to yieldπθ1\\pi\_\{\\theta\_\{1\}\}\. Stage 2 then sharpensπθ1\\pi\_\{\\theta\_\{1\}\}under the three\-axis reward of Eq\.[11](https://arxiv.org/html/2605.08715#S3.E11)via the GRPO update of Eq\.[12](https://arxiv.org/html/2605.08715#S3.E12), withπθ1\\pi\_\{\\theta\_\{1\}\}frozen as the reference policyπref\\pi\_\{\\text\{ref\}\}so the KL regularizer pullsπθ\\pi\_\{\\theta\}back toward the boundary alignment learned in Stage 1 rather than toward the generic baseπθ0\\pi\_\{\\theta\_\{0\}\}\.

## Appendix BAdditional Experiment Setups

### B\.1Details of Datasets

##### Our Proposed:AFTraj\-22K\.

AFTraj\-22Kcomprises2,2722\{,\}272family\-level multi\-agent trajectories spanning the three domains targeted by online auditing, with safe trajectories retained under our three\-predicate filter and unsafe trajectories obtained from two complementary streams \(Section[3\.1](https://arxiv.org/html/2605.08715#S3.SS1)\)\. Table[3](https://arxiv.org/html/2605.08715#A2.T3)reports the per\-domain composition, and Figure[6](https://arxiv.org/html/2605.08715#A2.F6)shows the per\-domain distribution of the decisive\-error step\.

Table 3:Per\-domain composition ofAFTraj\-22K\. Counts are reported at the un\-expanded family level \(one row per labelled trajectory\)\. Train and test families are obtained from a stratified split that places each safe trajectory and all of its variants in the same partition\. TheOverallcolumn reports per\-row sums for counts and weighted statistics for averages\.Metric / DomainCodingMathAgenticOverallBenchmarksHumanEval\+, MBPP\+\[[29](https://arxiv.org/html/2605.08715#bib.bib29)\]MATH\-500\[[16](https://arxiv.org/html/2605.08715#bib.bib16)\]GAIA\[[33](https://arxiv.org/html/2605.08715#bib.bib33)\], HotpotQA\[[57](https://arxiv.org/html/2605.08715#bib.bib57)\]—Multi\-Agent SystemsAutoGen\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\], MetaGPT\[[17](https://arxiv.org/html/2605.08715#bib.bib17)\]AutoGen\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\]Smolagents\[[41](https://arxiv.org/html/2605.08715#bib.bib41)\]—Verified Safe3613954021,158Unsafe2473974701,114Train Families5176767471,940Test Families91116125332Total6087928722,272Avg\. \# turns10\.016\.28\.411\.5

![Refer to caption](https://arxiv.org/html/2605.08715v1/x6.png)Figure 6:Distribution of the decisive\-error step \(normalized by trajectory lengthNN\) across the three domains ofAFTraj\-22Kand their aggregate\. Dashed lines mark the per\-panel mean\. The three domains exhibit qualitatively distinct shapes, where Coding errors concentrate in the late half, Math errors spread across the trajectory with a long tail, and Agentic errors are front\-loaded, while theOverallpanel shows that across the full1,1141\{,\}114unsafe trajectories the decisive step is broadly distributed throughout the trajectory rather than clustered at any fixed prefix position, supporting the claim that an online auditor must be calibrated to commit at the step the trajectory actually goes wrong\.The2,2722\{,\}272family\-level entries in Table[3](https://arxiv.org/html/2605.08715#A2.T3)are obtained from a strict additional pass on top of our raw curated trajectory pool\. On Coding,411411verified\-safe trajectories from the AutoGen Swarm pool are reduced to361361after rejecting degenerate runs whose tester agent never independently invoked the test harness\. On Math,396396verified\-safe trajectories are nearly all retained \(395395\)\. On the Agentic side, the GAIA and HotpotQA pools contribute301301verified\-safe trajectories combined, supplemented by two additional Agentic task pools covering expert\-team coordination and tool\-safety scenarios that bring the Agentic safe count to402402\. Failure trajectories on the Agentic side are dominated by the diagnostic stream \(322322\), since GAIA and HotpotQA naturally produce many failed runs whose decisive step can be recovered through the propose\-and\-verify procedure of Section[3\.1](https://arxiv.org/html/2605.08715#S3.SS1)\. Failure trajectories on Math and Coding are dominated by the constructive stream \(333333and247247respectively\), because verified\-safe trajectories in these closed\-form domains are abundant and easier to perturb in a controlled manner\.

##### External Benchmark: Who&When\.

We additionally evaluate on the Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]benchmark as a strictly external test bed\. Who&When provides127127multi\-agent systems with annotated decisive\(agent,step\)\(\\text\{agent\},\\text\{step\}\)pairs, spanning both algorithm\-generated agentic systems built via the CaptainAgent framework and a hand\-crafted Magentic\-One pool\. All trajectories in Who&When are entirely disjoint fromAFTraj\-22Kboth in terms of agentic system construction and in terms of underlying tasks\. Following the original protocol, only failed trajectories with verified decisive errors are released\. We evaluate every model on this benchmark under the same online auditing protocol of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2), walking each trajectory step by step and recording the earliest alarm\.

### B\.2Details on Evaluation Metrics

For each test trajectoryτ\\tauwith ground\-truth labely​\(τ\)∈\{Safe,Unsafe\}y\(\\tau\)\\in\\\{\\textsc\{Safe\},\\textsc\{Unsafe\}\\\}, the auditor performs the strict step\-by\-step incremental walk of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2)and at each prefix emits a structured verdict carrying a categorical label, a predicted decisive stepk^​\(τ\)\\hat\{k\}\(\\tau\), and a responsible agent\. We denote byd​\(τ\)d\(\\tau\)the earliest prefix at which the verdict turns intoAlarm, set to∞\\inftywhen no alarm is raised\. Let𝒰=\{τ:y​\(τ\)=Unsafe\}\\mathcal\{U\}=\\\{\\tau:y\(\\tau\)=\\textsc\{Unsafe\}\\\}and𝒰det=\{τ∈𝒰:d​\(τ\)<∞\}\\mathcal\{U\}\_\{\\text\{det\}\}=\\\{\\tau\\in\\mathcal\{U\}:d\(\\tau\)<\\infty\\\}denote the unsafe set and its alarm\-triggered subset\. Three metrics summarize auditor quality along complementary axes\.

##### Exact\-Step F1 \(Exact\-F1↑\\uparrow\)\.

Step Recall is the fraction of unsafe trajectories whose decisive step is exactly localized, while Step Precision is the same fraction restricted to alarm\-triggered trajectories,

Recallstep=\|\{τ∈𝒰:k^​\(τ\)=k∗​\(τ\)\}\|\|𝒰\|,Precisionstep=\|\{τ∈𝒰det:k^​\(τ\)=k∗​\(τ\)\}\|\|𝒰det\|,\\text\{Recall\}\_\{\\text\{step\}\}=\\frac\{\|\\\{\\tau\\in\\mathcal\{U\}:\\hat\{k\}\(\\tau\)=k^\{\*\}\(\\tau\)\\\}\|\}\{\|\\mathcal\{U\}\|\},\\qquad\\text\{Precision\}\_\{\\text\{step\}\}=\\frac\{\|\\\{\\tau\\in\\mathcal\{U\}\_\{\\text\{det\}\}:\\hat\{k\}\(\\tau\)=k^\{\*\}\(\\tau\)\\\}\|\}\{\|\\mathcal\{U\}\_\{\\text\{det\}\}\|\},\(13\)and Exact\-F1 is their harmonic mean\. Step Recall penalizes missed errors and Step Precision penalizes incorrectly localized alarms within the detected pool, so Exact\-F1 jointly captures both failure modes under exact\-match localization\.

##### Absolute Step Shift \(ASS↓\\downarrow\)\.

For each detected unsafe trajectory, ASS measures the absolute distance between the predicted and ground\-truth decisive steps, averaged over𝒰det\\mathcal\{U\}\_\{\\text\{det\}\},

ASS=1\|𝒰det\|​∑τ∈𝒰det\|k^​\(τ\)−k∗​\(τ\)\|\.\\text\{ASS\}=\\frac\{1\}\{\|\\mathcal\{U\}\_\{\\text\{det\}\}\|\}\\sum\_\{\\tau\\in\\mathcal\{U\}\_\{\\text\{det\}\}\}\\bigl\|\\hat\{k\}\(\\tau\)\-k^\{\*\}\(\\tau\)\\bigr\|\.\(14\)ASS remains informative even when alarms miss the exact step, providing a graded signal of localization quality that the binary correctness in Exact\-F1 cannot capture, and it is undefined on𝒰∖𝒰det\\mathcal\{U\}\\setminus\\mathcal\{U\}\_\{\\text\{det\}\}because there is no reported step to compare against\.

### B\.3Details of Implementations

#### B\.3\.1Implementation details ofAFTraj\-22KConstruction

##### Trajectory Collection\.

We instantiate three multi\-agent system templates corresponding to the three domains ofAFTraj\-22K\. For Coding, we use AutoGen Swarm\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\]with two roles, CodeWriter and CodeTester, that hand off control on demand and terminate on the sentinelFINAL\_VERIFIED\_TESTS\_PASSED, run on HumanEval\+ and MBPP\+\[[29](https://arxiv.org/html/2605.08715#bib.bib29)\]\. For Math, we use the same AutoGen Swarm template with MathSolver and Verifier roles terminating on the sentinelANSWER\_VERIFIED, run on MATH\-500\[[16](https://arxiv.org/html/2605.08715#bib.bib16)\]\. For Agentic, we use Smolagents\[[41](https://arxiv.org/html/2605.08715#bib.bib41)\]with a CodeAgent Manager that delegates web search and Wikipedia retrieval to a ToolCallingAgent search\_agent, run on GAIA\[[33](https://arxiv.org/html/2605.08715#bib.bib33)\]and HotpotQA\[[57](https://arxiv.org/html/2605.08715#bib.bib57)\]; for GAIA we additionally handle file attachments by injecting their content \(parsed for text formats and base64\-encoded for images, with audio transcribed via Whisper\) into the agent prompt\. The backbone LLM is uniformly GPT\-5\.4\-mini across all four sub\-benchmarks, with greedy decoding and a per\-task step budget capped at4040\.

##### Verified Safe Curation\.

Each successful rolloutτ∈𝒟succ\\tau\\in\\mathcal\{D\}\_\{\\text\{succ\}\}is admitted to𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}only if it passes three independent predicates\. The outcome predicateϕoutcome\\phi\_\{\\text\{outcome\}\}enforces strict equivalence against the reference, instantiated as sympy symbolic equivalence withLaTeXnormalization on Math, the GAIA official scorer with number/list normalization on GAIA, an article\-insensitive normalizer with conservative person\-name and location\-suffix variants on HotpotQA, and subprocess test execution with a1515s timeout on Coding\. The integrity predicateϕintegrity\\phi\_\{\\text\{integrity\}\}rejects any trajectory containing tool errors, serialization failures, empty predictions, or environment\-limited terminations\. The coherence predicateϕcoherence\\phi\_\{\\text\{coherence\}\}uses a GPT\-5\.4 judge to verify that each turn remains aligned with the declared sub\-goal\. Curation is realized as a multi\-pass pipeline of post\-generation batch validation followed by a strict cross\-pass audit, retaining only trajectories that survive all three predicates\.

##### Curation of Failure Trajectories with Decisive Error Annotations\.

The two complementary streams introduced in Section[3\.1](https://arxiv.org/html/2605.08715#S3.SS1)are realized as follows\. The*constructive stream*starts fromτ∈𝒟safe\\tau\\in\\mathcal\{D\}\_\{\\text\{safe\}\}, samples an injection stepkinjk\_\{\\text\{inj\}\}uniformly within the agent\-controlled prefix, and draws a fault categoryccfrom a domain\-specific catalog\. For Math the catalog is \{computation\_slip,premature\_finalization,verification\_shortcut,verdict\_misread\} \(\|𝒞Math\|=4\|\\mathcal\{C\}\_\{\\text\{Math\}\}\|=4\), and for Coding it is \{code\_bug,verification\_skip,verdict\_misread\} \(\|𝒞Coding\|=3\|\\mathcal\{C\}\_\{\\text\{Coding\}\}\|=3\)\. The fault distributionπfault\\pi\_\{\\text\{fault\}\}is realized in two complementary modes, a turn\-rewriting mode that statically rewrites the targeted agent turn for short\-horizon math and coding trajectories, and a live\-replay mode that re\-executes the multi\-agent system from the corrupted prefix to obtain authentic downstream propagation for tool\-augmented agentic trajectories, with category\-specific injections includingtool\_injection,prompt\_injection,verification\_shortcut,solver\_premature\_verdict,verifier\_text\_shortcut, andfinal\_verdict\_override\. A post\-injection acceptance check rejects candidates whose outcome flips back to success or whose targeted turn was not in fact modified, after which each accepted candidate is admitted to𝒟failinj\\mathcal\{D\}\_\{\\text\{fail\}\}^\{\\text\{inj\}\}with the by\-construction label\(k∗,a∗\)=\(kinj,akinj\)\(k^\{\*\},a^\{\*\}\)=\(k\_\{\\text\{inj\}\},a\_\{k\_\{\\text\{inj\}\}\}\)\. The*diagnostic stream*operates onτ∈𝒟fail\\tau\\in\\mathcal\{D\}\_\{\\text\{fail\}\}, where the decisive step is unknown and must be recovered\. We useP=5P=5independent proposer calls that return candidate\(kcand,akcand\)\(k\_\{\\text\{cand\}\},a\_\{k\_\{\\text\{cand\}\}\}\)pairs, andV=3V=3independent verifier calls that re\-check each unique candidate along four binary criteria\(sexists,ssubstantive,sdecisive,searliest\)\(s\_\{\\text\{exists\}\},s\_\{\\text\{substantive\}\},s\_\{\\text\{decisive\}\},s\_\{\\text\{earliest\}\}\)\. A candidate is admitted only if its strict\-support count, defined as the number of verifiers under which all four criteria simultaneously hold, exceeds the majority threshold⌊V/2⌋\+1=2\\lfloor V/2\\rfloor\+1=2\. For each trajectory, the highest\-strict\-support candidate is selected, with ties broken by the verifier confidence margin\. Both proposers and verifiers are instantiated by GPT\-5\.4 at temperature0\.20\.2to retain modest diversity while keeping the decisive criteria stable\.

#### B\.3\.2Implementation details of training

##### Stage 1: Failure Boundary Alignment\.

We implement the dual\-subset BPPO objective of Eq\.[9](https://arxiv.org/html/2605.08715#S3.E9)with a custom FSDP trainer launched on2×2\\timesNVIDIA H200 GPUs, optimizing Qwen2\.5\-7B\-Instruct with the reference policy frozen at the same base\. To fit the8,1928\{,\}192\-token boundary\-pair prompts within memory, we adopt88\-bit AdamW from bitsandbytes with weight decay0\.010\.01, bfloat16 mixed precision, and gradient checkpointing, and use a cosine learning\-rate schedule with a5050\-step linear warmup\. The trainer is deliberately framework\-light to keep the boundary\-pair gradient flow auditable across the BS and BE subsets\.

##### Stage 2: Three\-Axis Verdict Sharpening\.

Stage 2 is implemented on top of the verl framework\[[46](https://arxiv.org/html/2605.08715#bib.bib46)\], initializing both the trainable policyπθ\\pi\_\{\\theta\}and the frozen reference policyπref\\pi\_\{\\text\{ref\}\}from the Stage 1 checkpointπθ1\\pi\_\{\\theta\_\{1\}\}\. The KL term is applied directly in the loss rather than folded into the reward \(verl flagsuse\_kl\_loss=Trueanduse\_kl\_in\_reward=False\) and is estimated with verl’slow\_var\_kloption, which is the same k3 estimator referenced in Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)\. Each prompt is rolled outG=8G=8times by a vLLM backend with rollout temperature1\.01\.0and top\-pp1\.01\.0, while validation uses greedy decoding at temperature0\.10\.1\. The custom three\-axis reward of Eq\.[11](https://arxiv.org/html/2605.08715#S3.E11)is wrapped under verl’s DAPO reward manager with a soft overlong\-response buffer that smoothly penalizes responses approaching the4,0964\{,\}096\-token response budget, preventing reward saturation from rollouts that exceed the budget\. Full hyperparameters of both stages are reported in Table[4](https://arxiv.org/html/2605.08715#A2.T4)\.

Table 4:Training hyperparameters for Stage\-1 and Stage\-2 of*AgentForesight*\-7B\.HyperparameterStage 1Stage 2Base policyQwen2\.5\-7B\-Instruct \(πθ0\\pi\_\{\\theta\_\{0\}\}\)Stage 1 checkpointπθ1\\pi\_\{\\theta\_\{1\}\}Frozen reference policyπθ0\\pi\_\{\\theta\_\{0\}\}πθ1\\pi\_\{\\theta\_\{1\}\}Trainer frameworkcustom FSDP BPPOverl\[[46](https://arxiv.org/html/2605.08715#bib.bib46)\]Training samples1,9021\{,\}902boundary pairs1,9401\{,\}940prompts \(×G\\times Grollouts\)Learning rate5×10−75\\times 10^\{\-7\}1×10−61\\times 10^\{\-6\}LR schedulecosine\+\+5050\-step warmupconstantOptimizer8\-bit AdamW \(bitsandbytes\)AdamW \(verl default\)β\\beta/βKL\\beta\_\{\\text\{KL\}\}β=0\.1\\beta=0\.1βKL=10−3\\beta\_\{\\text\{KL\}\}=10^\{\-3\}KL estimator—low\-variance k3Effective batch1616\(1×161\\times 16grad\-accum\)3232Group sizeGG—88rollouts / promptEpochs3388Max prompt length8,1928\{,\}1928,1928\{,\}192Max response length—4,0964\{,\}096Mixed precisionbfloat16bfloat16Gradient checkpointingenabledenabledRollout decoding—vLLM,T=1\.0T=1\.0, top\-p=1\.0p=1\.0Validation decoding—greedy,T=0\.1T=0\.1Hardware2×2\\timesH200 \(FSDP\)2×2\\timesH200 \(FSDP\)

#### B\.3\.3Implementation details of baselines

##### LLM auditors under online auditing\.

The five open\-source small\-size LLMs \(Llama\-3\.2\-3B, Gemma\-3\-4B, Qwen2\.5\-7B\-Instruct, Qwen3\-8B, Qwen3\-32B\) and the five proprietary LLMs \(GPT\-4\.1, Gemini\-3\-Flash, Claude\-Haiku\-4\.5, DeepSeek\-V4\-Flash, DeepSeek\-V4\-Pro\) are all evaluated under the strict step\-by\-step incremental walk of Section[2](https://arxiv.org/html/2605.08715#S2)\. At each stepk=0,1,…,\|τ\|−1k=0,1,\\ldots,\|\\tau\|\-1, the auditor is queried with the prefixτ0:k\\tau\_\{0:k\}wrapped in the same system prompt and incremental\-view user prompt as*AgentForesight*\-7B \(Appendix[E](https://arxiv.org/html/2605.08715#A5)\), and emits a JSON verdicty^∈\{Safe\}∪\{\(k^,a^,f^\)\}\\hat\{y\}\\in\\\{\\textsc\{Safe\}\\\}\\cup\\\{\(\\hat\{k\},\\hat\{a\},\\hat\{f\}\)\\\}\. The walk halts on the first prefix at which the verdict raises an alarm with a parseable step index, and that step is recorded as the predicted decisive step \(*first\-alarm*aggregation\)\. All ten LLM auditors decode greedily \(T=0\.0T=0\.0\) with a per\-call response budget of1,5001\{,\}500tokens; the five HuggingFace models are loaded in bfloat16 on a single H200 GPU, and the five proprietary models are queried through their respective public APIs\.

##### Perplexity\-7B and ToT\-7B\.

Both methodological baselines are instantiated on Qwen2\.5\-7B\-Instruct and produce a scalar per\-step score under prefix\-only context\.Perplexity\-7B\[[8](https://arxiv.org/html/2605.08715#bib.bib8)\]computes the length\-normalized log\-likelihoodLN\-LLk=1Tk​∑t=1Tklog⁡p​\(tokt∣τ0:k−1,prev\-tokens\)\\text\{LN\-LL\}\_\{k\}=\\frac\{1\}\{T\_\{k\}\}\\sum\_\{t=1\}^\{T\_\{k\}\}\\log p\(\\text\{tok\}\_\{t\}\\mid\\tau\_\{0:k\-1\},\\text\{prev\-tokens\}\)for every agent turnkk, with non\-agent turns \(user, environment, tool, system\) skipped\.ToT\-7B\[[58](https://arxiv.org/html/2605.08715#bib.bib58)\]replaces the log\-likelihood with a value rating in\{Sure,Likely,Impossible\}\\\{\\textsc\{Sure\},\\textsc\{Likely\},\\textsc\{Impossible\}\\\}obtained via greedy decoding on a step\-level evaluator prompt that observes the prefixτ0:k−1\\tau\_\{0:k\-1\}and the candidate turn at stepkk, mapped to scores\{2,1,0\}\\\{2,1,0\\\}\. The per\-trajectory verdict in both cases follows a*first\-crossing*decision rule, scanning steps in temporal order and raising an alarm at the firstkkat whichscorek\\text\{score\}\_\{k\}falls below a tuned thresholdθ\\thetaand emitting SAFE otherwise\. The thresholdθ\\thetais tuned on a held\-out training split by maximizing detection F1 and is then frozen for evaluation on the held\-out test split\.

##### Reflexion\-7B\.

Reflexion\-7B\[[48](https://arxiv.org/html/2605.08715#bib.bib48)\]repurposes the verbal self\-reflection module of Reflexion as a per\-step error detector on Qwen2\.5\-7B\-Instruct\. At every prefixτ0:k\\tau\_\{0:k\}the critic LLM is shown the system role, the prior conversation historyτ0:k−1\\tau\_\{0:k\-1\}, and the candidate turn at stepkk, and asked the binary question*“Does this step contain a critical error? Answer with only ‘yes’ or ‘no’\.”*We then read the next\-token distribution at the assistant header and compute a per\-step probability of correctnesspok,k=P​\(‘no’\)/\(P​\(‘yes’\)\+P​\(‘no’\)\)p\_\{\\text\{ok\},k\}=P\(\\text\{\`no'\}\)/\\bigl\(P\(\\text\{\`yes'\}\)\+P\(\\text\{\`no'\}\)\\bigr\), where each ofP​\(‘yes’\)P\(\\text\{\`yes'\}\)andP​\(‘no’\)P\(\\text\{\`no'\}\)is unioned over single\-token capitalization variants\{yes,Yes,YES\}\\\{\\text\{\{yes\}\},\\text\{\{Yes\}\},\\text\{\{YES\}\}\\\}and\{no,No,NO\}\\\{\\text\{\{no\}\},\\text\{\{No\}\},\\text\{\{NO\}\}\\\}, and non\-agent turns \(user, environment, tool, system\) are skipped\. The per\-trajectory verdict is then obtained by a*first\-crossing*decision rule, scanning steps in temporal order and raising an alarm at the firstkkat whichpok,kp\_\{\\text\{ok\},k\}drops below a thresholdθ\\thetaand emitting SAFE otherwise\. The thresholdθ\\thetais tuned on a clean300300\-trajectory held\-out pool by maximizing detection F1 over a1919\-quantile sweep of the score distribution and is then frozen for evaluation on the held\-out test split\. The critic decodes a single forward pass per step in bfloat16 on one H200 GPU\.

##### AgentDebug\-7B\.

AgentDebug\-7B\[[67](https://arxiv.org/html/2605.08715#bib.bib67)\]is the only baseline that consumes the full completed trajectoryτ0:\|τ\|−1\\tau\_\{0:\|\\tau\|\-1\}in a single shot, mirroring the post\-hoc protocol of the original AgentDebug paper\. We adopt its Phase 2 critical\-step identification prompt with one adaptation, namely that we extend the response schema to allow a SAFE outcome alongside the original\{critical\_step,critical\_agent,error\_type,root\_cause,evidence\}\\\{\\text\{critical\\\_step\},\\text\{critical\\\_agent\},\\text\{error\\\_type\},\\text\{root\\\_cause\},\\text\{evidence\}\\\}JSON so the same baseline can be evaluated on both safe and unsafe trajectories\. The judge LLM is Qwen2\.5\-7B\-Instruct served via a vLLM endpoint, decoding greedily with a1,5001\{,\}500\-token response budget\. Its low average Exact\-F1 in Table[1](https://arxiv.org/html/2605.08715#S4.T1)reflects a structural mismatch between the post\-hoc whole\-trajectory training prior, where every input is assumed to carry a known failure outcome, and online auditing’s prefix\-restricted contract that additionally requires reliable separation of safe from unsafe trajectories, with the model often committing to an early non\-decisive step on safe trajectories and over\-attributing to early exploration turns on unsafe ones\.

## Appendix CDetailed Related Work

##### LLM\-based agentic systems\.

Building on single\-agent reasoning paradigms such as chain\-of\-thought\[[52](https://arxiv.org/html/2605.08715#bib.bib52),[61](https://arxiv.org/html/2605.08715#bib.bib61)\]and ReAct\[[59](https://arxiv.org/html/2605.08715#bib.bib59)\], together with tool\-augmented backbones like Toolformer\[[42](https://arxiv.org/html/2605.08715#bib.bib42)\], semantic file systems\[[47](https://arxiv.org/html/2605.08715#bib.bib47)\], and memory architectures such as MemGPT\[[38](https://arxiv.org/html/2605.08715#bib.bib38)\], recent work organizes LLMs into multi\-agent systems that coordinate specialized roles via tool use and inter\-agent communication\. Handcrafted frameworks fix agent roles and protocols, including AutoGen\[[53](https://arxiv.org/html/2605.08715#bib.bib53)\], MetaGPT\[[17](https://arxiv.org/html/2605.08715#bib.bib17)\], and Camel\[[24](https://arxiv.org/html/2605.08715#bib.bib24)\], while partially\-automated approaches such as GPTSwarm\[[68](https://arxiv.org/html/2605.08715#bib.bib68)\]optimize prompts or inter\-agent topology end\-to\-end\. These frameworks are deployed across a growing set of long\-horizon benchmarks spanning scientific assistance and open\-ended web navigation, including GAIA\[[33](https://arxiv.org/html/2605.08715#bib.bib33)\]and WebArena\[[66](https://arxiv.org/html/2605.08715#bib.bib66)\]\. Our work targets this entire spectrum of deployed systems through*online auditing*, treating the underlying agentic system as a black box and auditing its trajectory step by step at deployment time, without modifying the agents, tools, or inter\-agent protocol\.

##### Failure analysis and post\-hoc attribution for LLM agents\.

A growing body of work characterizes how multi\-agent systems fail\. MAST\[[3](https://arxiv.org/html/2605.08715#bib.bib3)\]catalogs fourteen prevalent failure modes spanning task disobedience, role misuse, and reasoning\-action mismatches across popular frameworks, while complementary studies on agentic verification\[[49](https://arxiv.org/html/2605.08715#bib.bib49)\]and on the definition and detection of agent defects\[[35](https://arxiv.org/html/2605.08715#bib.bib35)\]formalize where errors arise within and across modules\. Building on this characterization, the closest line to ours formulates*failure attribution*as identifying the responsible\(agent,step\)\(\\text\{agent\},\\text\{step\}\)pair from a completed trajectory\. Who&When\[[63](https://arxiv.org/html/2605.08715#bib.bib63)\]curates failure logs from 127 multi\-agent systems and benchmarks all\-at\-once, step\-by\-step, and binary\-search prompting baselines for attribution\. AgenTracer\[[62](https://arxiv.org/html/2605.08715#bib.bib62)\]introduces an automated counterfactual\-replay and fault\-injection pipeline for labelling decisive errors and trains AgenTracer\-8B with a multi\-granular reward over the full trajectory\. AgentDebug\[[67](https://arxiv.org/html/2605.08715#bib.bib67)\]derives a five\-module error taxonomy spanning memory, reflection, planning, action, and system\-level failures, and uses LLM\-generated corrective feedback to re\-execute the run from its root cause\. A parallel*agentic attribution*line\[[39](https://arxiv.org/html/2605.08715#bib.bib39)\]attributes the decisive action of a completed trajectory to internal drivers, e\.g\., specific memory entries or tool observations, via temporal likelihood dynamics\. Self\-correction approaches such as reflexion\-style retries\[[48](https://arxiv.org/html/2605.08715#bib.bib48)\]and self\-refine\[[32](https://arxiv.org/html/2605.08715#bib.bib32)\]share the same operational stance, triggering a corrective rollout once an outcome has been observed\. All these formulations consume a completed trajectory and identify the responsible component in hindsight, by construction forfeiting the opportunity to intervene while execution is still unfolding\. Our*online auditing*reframing instead commits the auditor at every step under prefix\-restricted observation, a strictly stronger demand that, as Table[1](https://arxiv.org/html/2605.08715#S4.T1)shows, even a directly re\-purposed AgentDebug\-7B fails to satisfy under naive per\-step re\-application\.

##### Agent monitoring and runtime safety\.

A complementary line of work treats the auditor as a separate process that runs alongside, rather than inside, the agent\. Reliable Weak\-to\-Strong Monitoring\[[22](https://arxiv.org/html/2605.08715#bib.bib22)\]systematizes a red\-team workflow over agent and monitor situational awareness, and shows that a hybrid hierarchical\-sequential scaffold lets a weaker monitor reliably oversee a stronger agent, the closest conceptual analogue to our 7B auditor monitoring stronger underlying agents\. Baker et al\.\[[2](https://arxiv.org/html/2605.08715#bib.bib2)\]demonstrate that chain\-of\-thought monitoring of reasoning agents catches reward hacking far better than action\-only monitoring, but warn that using the monitor’s signal directly as the agent’s reward induces obfuscated reward hacking, a constraint we respect by leaving the underlying agentic system untouched and never feeding the auditor’s verdict back as the agent’s training signal\. SHADE\-Arena\[[23](https://arxiv.org/html/2605.08715#bib.bib23)\]benchmarks sabotage detection over1717task pairs and reports a maximum monitor AUC of0\.870\.87, framing reliable trajectory\-level monitoring as still far from safety\-critical thresholds\. Stress\-testing deliberative alignment for anti\-scheming\[[43](https://arxiv.org/html/2605.08715#bib.bib43)\]reduces covert behavior ino​3o3/o​4o4\-mini through training\-stage interventions but is honest that residual situational awareness contaminates the gain\. Our work differs from this monitoring/safety line in two operational respects, namely we issue a per\-step continue\-or\-alarm verdict rather than a trajectory\-level binary judgement, and we ground the auditor in a curated corpus of decisive\-error annotations rather than red\-team or sabotage trajectories\.

##### Reinforcement learning for agentic LLMs\.

Reinforcement learning has been widely used to shape the policy of agentic LLMs themselves during rollout\. Search\-R1\[[21](https://arxiv.org/html/2605.08715#bib.bib21)\]optimizes a search\-augmented LLM with GRPO under an outcome reward, AgentGym\-RL\[[54](https://arxiv.org/html/2605.08715#bib.bib54)\]extends GRPO\-style updates to long\-horizon agent training\. AgentFlow\[[26](https://arxiv.org/html/2605.08715#bib.bib26)\]co\-trains coordination and reasoning roles via policy\-gradient updates, GiGPO\[[9](https://arxiv.org/html/2605.08715#bib.bib9)\]introduces a two\-level critic\-free advantage that combines episode\-level GRPO with anchor\-state grouping for step\-level credit assignment, and AgentPRM\[[55](https://arxiv.org/html/2605.08715#bib.bib55)\]introduces a process reward model that supplies step\-level supervision during rollout\. Closely related, recent work studies the new failure modes introduced when agents are allowed to self\-evolve\[[45](https://arxiv.org/html/2605.08715#bib.bib45)\]\. AgenTracer\[[62](https://arxiv.org/html/2605.08715#bib.bib62)\]departs from this train\-the\-agent stance by training a tracer network with a composite reward, but still does so over completed trajectories\. Foundationally, credit\-assignment ideas from cooperative LLM\-agent training\[[30](https://arxiv.org/html/2605.08715#bib.bib30)\]inform how scalar outcome rewards can be redistributed across agents and steps\. We adopt GRPO\[[15](https://arxiv.org/html/2605.08715#bib.bib15)\]as our Stage 2 optimizer because the group\-relative advantage eliminates the need for a learned critic at the prompt lengths typical of multi\-agent trajectories, and combine it with Boundary\-Pair Preference Optimization \(BPPO\), a preference\-optimization\[[40](https://arxiv.org/html/2605.08715#bib.bib40)\]variant tailored to adjacent safe/unsafe boundary pairs, for Stage 1 boundary alignment\. Unlike prior agentic RL work that trains the agent itself to act more reliably, our coarse\-to\-fine recipe leaves the underlying agentic system untouched and instead trains an external auditor that runs alongside the deployed system and emits per\-step continue\-or\-alarm verdicts under prefix\-restricted observation \(Section[3\.2](https://arxiv.org/html/2605.08715#S3.SS2)\)\.

##### LLM\-as\-judge, reward models, and step\-level critics\.

Using LLMs themselves to evaluate other LLMs’ outputs has become a standard practice\[[14](https://arxiv.org/html/2605.08715#bib.bib14)\], with judges deployed as zero\-shot critics of completed answers on benchmarks like MT\-Bench\[[65](https://arxiv.org/html/2605.08715#bib.bib65)\], metrics such as G\-Eval\[[31](https://arxiv.org/html/2605.08715#bib.bib31)\], multi\-agent debate panels\[[4](https://arxiv.org/html/2605.08715#bib.bib4)\], and self\-checking modules\[[34](https://arxiv.org/html/2605.08715#bib.bib34)\]\. Trained reward models for RLHF score complete responses against learned human preferences\[[37](https://arxiv.org/html/2605.08715#bib.bib37)\]\. Process reward models score intermediate reasoning steps in math, including PRM800K\[[27](https://arxiv.org/html/2605.08715#bib.bib27)\]and Math\-Shepherd\[[50](https://arxiv.org/html/2605.08715#bib.bib50)\], with ProcessBench\[[64](https://arxiv.org/html/2605.08715#bib.bib64)\]providing a step\-level error detection benchmark\. Recent agent\-specific critics extend these ideas to scoring tool\-use trajectories, including Agent\-as\-a\-Judge\[[69](https://arxiv.org/html/2605.08715#bib.bib69)\]\. Our auditor is closest to these step\-level critics in that it evaluates partial reasoning rather than completed outputs, but it differs in two operationally important ways\. First, it produces a structured verdict consisting of a categorical label, a step index, and a responsible agent, rather than a scalar quality score, which lets it act directly as a deployment\-time monitor\. Second, it is trained for the online auditing protocol rather than zero\-shot prompted, and the GPT\-4\.1 and DeepSeek\-V4\-Pro baselines in Section[4](https://arxiv.org/html/2605.08715#S4)show this gap is not closed by sheer model scale\.

## Appendix DAdditional Experimental Results

This section reports the full numerical results behind the two analysis figures of Section[4](https://arxiv.org/html/2605.08715#S4)and adds a per\-call efficiency analysis that complements the deployment\-level discussion of Figure[4](https://arxiv.org/html/2605.08715#S4.F4)\.

### D\.1Full Two\-Stage Ablation Results

Table[5](https://arxiv.org/html/2605.08715#A4.T5)reports the per\-domain Exact\-F1 and ASS values that underlie Figure[3](https://arxiv.org/html/2605.08715#S4.F3), separating the contribution of each stage of our coarse\-to\-fine recipe\. Stage 1 alone \(BPPO on adjacent safe/unsafe boundary pairs\) lifts overall Exact\-F1 from21\.0521\.05to35\.6335\.63by establishing a learnable failure boundary\. Stage 2 alone \(GRPO under the three\-axis reward\) reaches50\.4250\.42but exhibits a clear domain split, sharpening Math \(63\.6463\.64\) and Coding \(72\.7372\.73\) where decisive errors are sharply localizable, yet underperforming Stage 1 on Agentic \(19\.0519\.05vs\.31\.5831\.58\) where the failure boundary is harder to discriminate\. The full recipe combines both stages and reaches66\.4466\.44overall, with Agentic recovering to48\.7048\.70\. The very tight ASS values that Stage 2 alone attains on Math and Coding \(0\.030\.03and0\.170\.17\) reflect that it raises very few alarms but places them precisely; layering Stage 1’s risk\-anticipation prior trades a small ASS overhead for the substantial Exact\-F1 gains visible across all four column groups\.

Table 5:Full per\-domain results of the two\-stage*coarse\-to\-fine*ablation, expanding Figure[3](https://arxiv.org/html/2605.08715#S4.F3)with both Exact\-F1 and ASS\.Bold==best per column\.MathCodingAgenticOverallConfigurationExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowExact\-F1↑\\uparrowASS↓\\downarrowBase \(Qwen2\.5\-7B\-Instruct\)10\.393\.8038\.202\.2614\.002\.9621\.052\.75\+\+Stage 1 \(BPPO only\)38\.243\.1037\.971\.6531\.581\.7735\.632\.30\+\+Stage 2 \(GRPO only\)63\.640\.0372\.730\.1719\.052\.1950\.420\.55Stage 1\+\+Stage 2 \(*AgentForesight*\-7B, ours\)77\.360\.9678\.870\.1848\.700\.5466\.440\.59

### D\.2Full Deployment Trade\-Off Results

Table[6](https://arxiv.org/html/2605.08715#A4.T6)reports the False Alarm Rate \(FAR\) and Step Accuracy values behind the scatter plot in Figure[4](https://arxiv.org/html/2605.08715#S4.F4)\. The two columns measure the auditor’s behavior on the two complementary halves ofAFTraj\-22K, with FAR computed on𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}and Step Accuracy computed on𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}\. Only*AgentForesight*\-7B operates inside the deployable region of FAR≤20%\\leq 20\\%and Step\-Acc≥50%\\geq 50\\%\(FAR=2\.37%=2\.37\\%, Step\-Acc=59\.51%=59\.51\\%\); the strongest proprietary baseline DeepSeek\-V4\-Pro lies just outside \(FAR=43\.20%=43\.20\\%, Step\-Acc=53\.99%=53\.99\\%\), while the smaller open\-source backbones collapse to near\-universal false alarms\. The gap is consistent with our coarse\-to\-fine recipe, where Stage 1’s risk\-anticipation prior suppresses spurious alarms on safe prefixes and Stage 2’s three\-axis reward sharpens alarm placement on unsafe runs\.

Table 6:Full results behind Figure[4](https://arxiv.org/html/2605.08715#S4.F4): False Alarm Rate on𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\}and Step Accuracy on𝒟unsafe\\mathcal\{D\}\_\{\\text\{unsafe\}\}for every auditor evaluated onAFTraj\-22K\.Bold==best per column,underline==second\-best\.MethodFAR \(%\)↓\\downarrowStep\-Acc \(%\)↑\\uparrowOpen\-Source LLMsLlama3\.2\-3B90\.5320\.86Gemma3\-4B97\.6310\.43Qwen2\.5\-7B\-Instruct46\.1536\.20Qwen3\-8B56\.8038\.04Proprietary LLMsGPT\-4\.185\.8038\.04Gemini\-3\-Flash67\.8638\.04Claude\-Haiku\-4\.568\.6433\.13DeepSeek\-V4\-Flash59\.7647\.24DeepSeek\-V4\-Pro43\.2053\.99*AgentForesight*\-7B \(ours\)2\.3759\.51
### D\.3Computational and Cost Analysis

Beyond detection quality, an online auditor must be cheap enough to be queried at every prefix without rate\-limiting the host system\. Table[7](https://arxiv.org/html/2605.08715#A4.T7)reports per\-call deployment cost along two axes that matter at scale\. Wall\-clock latency is hard\-measured from the eval logs as total elapsed seconds divided by total audit calls, and API cost is computed from per\-call input and output token counts at the official 2026\-05 pricing of each provider\. Open\-source auditors incur only compute time and no per\-call charge\.*AgentForesight*\-7B serves audits locally at1\.031\.03s/call on a single H200 and4\.734\.73s/call on the smaller RTX 4500 Ada, undercutting every API\-served baseline at the same backbone scale and outperforming the strongest proprietary baseline DeepSeek\-V4\-Pro \(25\.7725\.77s/call, $2\.972 per1​k1kcalls\) by roughly25×25\\timeson latency at zero per\-call charge\. This deployment profile, combined with the prefix\-restricted online auditing protocol of Section[2](https://arxiv.org/html/2605.08715#S2), lets a single H200 colocate the auditor with a host agent without rate\-limiting the underlying multi\-agent pipeline\.

Table 7:Per\-call deployment efficiency of the auditor\.Latency\(s/call, wall\-clock from eval logs\) andAPI cost\($/1k calls, computed from per\-call input/output tokens at official 2026\-05 pricing\)\. Open\-source models incur only compute time,Bold==best,underline==second\-best\.ModelParamsHostingLatency \(s/call\)↓\\downarrow$/1k calls↓\\downarrowOpen\-Source LLMs \(local\)Llama\-3\.2\-3B3BLocal \(H200\)2\.17—Gemma\-3\-4B4BLocal \(H200\)4\.81—Qwen2\.5\-7B\-Instruct7BLocal \(H200\)2\.32—Qwen3\-8B8BLocal \(H200\)13\.54—Proprietary LLMs \(API\)GPT\-4\.1—API3\.61$3\.690Gemini\-3\-Flash—API12\.06$0\.918Claude\-Haiku\-4\.5—API1\.28$1\.831DeepSeek\-V4\-Flash∼\\sim671B\-MoEAPI10\.62$0\.239DeepSeek\-V4\-Pro∼\\sim671B\-MoEAPI25\.77$2\.972*AgentForesight*\-7B \(ours\)7BLocal \(H200\)1\.03—
### D\.4Failure Mode Analysis

We inspect the two failure modes of*AgentForesight*\-7B onAFTraj\-22Kto characterize the boundary of the method\.Type A, false alarms on safe trajectories, occurs in only 4/169 safe runs \(FAR=2\.37%\\text\{FAR\}=2\.37\\%, matching Table[6](https://arxiv.org/html/2605.08715#A4.T6)\)\.Type B, mis\-localized alarms on unsafe trajectories, is dominated by off\-by\-one shifts \(2121/2828cases,75%75\\%\)\. Both modes concentrate in<10%<10\\%of the evaluation set and do not reverse the\+19\.88\+19\.88Exact\-F1 lead and3×3\\timestighter ASS reported in Table[1](https://arxiv.org/html/2605.08715#S4.T1)\. We walk through one representative trajectory per mode below\.

##### Type A: false alarm during in\-turn verifier self\-correction\.

The trajectory in the box below answers\(3​x−2\)​\(4​x\+1\)−\(3​x−2\)​4​x\+1\(3x\-2\)\(4x\+1\)\-\(3x\-2\)4x\+1atx=4x=4with gold1111and ends safely\. The Verifier’s tool calls surface a transientVERIFIED\_WRONG:13≠1113\\neq 11string in its scratchpad at step1212that the same Verifier itself contradicts at step1313by re\-deriving3​x−13x\-1and substitutingx=4x=4to recover1111, before closing the trajectory at step1414\.*AgentForesight*\-7B’s risk\-anticipation prior fires on this surface signal one turn earlier, conflating an intra\-turn self\-correction loop with a decisive error\. The pattern accounts for all three Math instances among the four Type A cases\.

A safe Math trajectory \(1515turns total; verification block at steps1010–1414shown\):Step 10 \(Verifier\)\.*\[empty turn; opens the verification block\.\]* Step 11 \(environment\)\.compute→\\to"3\*x \- 1";compute→\\to"11";verify\_answer→\\to"VERIFIED\_WRONG: 13≠\\neq11"\. Step 12 \(Verifier\)\.Reads back:3\*x \- 1 / 11 / VERIFIED\_WRONG: 13≠\\neq11\. Step 13 \(Verifier\)\.*“The final answer is incorrect\. Expanding…\\dotsgives3​x−13x\-1\. Now substitutex=4x=4:3​\(4\)−1=113\(4\)\-1=11\. So the correct answer is𝟏𝟏\\mathbf\{11\}\.”* Step 14 \(Verifier\)\.ANSWER\_VERIFIED: 11\.*AgentForesight*\-7B verdict\.answer: step1010/Verifier; reason:*“Verifier verified13≠1113\\neq 11\.”*

##### Type B: off\-by\-one upstream localization\.

The trajectory in the box below answers*“Are both Cypress and Ajuga genera?”*\(gold*no*\)\. The Manager emits the unverified “Yes” at step11, repeats it at step22, and commits it through the python interpreter at steps33–44to produce the final wrong output\.AFTraj\-22Kannotates the decisive step at the python wrap \(step22\), while*AgentForesight*\-7B localizes one turn earlier at the same Manager’s first emission of the same belief\. The alarm correctly classifies the trajectory as unsafe and identifies the responsible agent; the gap is between the upstream root of the wrong assertion and its downstream commit point, not between two distinct errors\.

An unsafe HotpotQA trajectory \(55turns total; full trajectory shown\):Step 0 \(user\)\.*“Are both Cypress and Ajuga genera?”* Step 1 \(Manager\)\."Yes\."\(no retrieval, no evidence\.\) Step 2 \(Manager\)\."Yes\."\(repeats the same assertion\.\) Step 3 \(Manager\)\.*\[python\_interpreter call wrapping the assertion\.\]* Step 4 \(environment\)\.Execution logs: Last output from code snippet: Yes\.*AgentForesight*\-7B verdict\.answer: step11/Manager; reason:*“Manager incorrectly responded ‘Yes’\.”*

### D\.5Additional Case Study

![Refer to caption](https://arxiv.org/html/2605.08715v1/x7.png)Figure 7:Math case study comparing decisive\-error verdicts from Gemini\-3\-Flash, GPT\-4\.1, and*AgentForesight*\-7B on a MATH\-500 trajectory whose decisive error commits late at Step 6\.##### Late\-committing decisive errors\.

Figure[7](https://arxiv.org/html/2605.08715#A4.F7)shows a Math trajectory where the decisive error commits at Step 6, yet both proprietary baselines commit to their verdict too early\. Gemini\-3\-Flash flags Step 4 on the symbolic tool result and GPT\-4\.1 flags Step 3 on the tool call, both still\-recoverable steps\.*AgentForesight*\-7B alone returns Step 6 with MathSolver as the responsible agent\. This contrast highlights an intrinsic difficulty of online auditing, where decisive errors in agentic systems are typically late\-committing, locally indistinguishable from recoverable steps, and propagation\-revealed\. Generic LLM judges collapse onto the first locally\-suspicious step, while*AgentForesight*\-7B identifies the committing step before propagation reveals it\.

## Appendix EPrompt Templates

This section displays the four most load\-bearing prompts of our pipeline\. The first two govern the online auditor’s task and per\-prefix observation, and the latter two govern the LLM\-as\-judge supervision used by the diagnostic stream of Appendix[B\.3\.1](https://arxiv.org/html/2605.08715#A2.SS3.SSS1)\. All other prompts, including the constructive\-stream injection prompts and baseline\-specific templates, are released in the accompanying code repository\.

##### Online auditor prompts\.

The system prompt below defines the auditor’s role and the strict two\-block response format used by both training stages and evaluation\. The incremental\-view user prompt wraps a partial trajectoryτ0:k\\tau\_\{0:k\}at every prefixkkduring the online auditing protocol of Definition[2\.2](https://arxiv.org/html/2605.08715#S2.Thmtheorem2)\.

`System Prompt for the Online Auditor Incremental\-View User Prompt`

`Diagnostic stream judge prompts\. The propose\-and\-verify procedure of Appendix B\.3\.1 relies on two prompts\. The proposer prompt draws up to three candidate decisive\-error steps from each failed trajectory, while the verifier prompt re\-checks each candidate along the four binary criteria \(sexists,ssubstantive,sdecisive,searliest\)\(s\_\{\\text\{exists\}\},s\_\{\\text\{substantive\}\},s\_\{\\text\{decisive\}\},s\_\{\\text\{earliest\}\}\) that back the strict\-support voting threshold ⌊V/2⌋\+1\\lfloor V/2\\rfloor\+1\. Diagnostic Stream Proposer Prompt Diagnostic Stream Verifier Prompt Appendix F Qualitative Examples from AFTraj\-22K This section displays one trajectory per domain to make the structure of AFTraj\-22K records concrete\. The three examples cover the three orthogonal sources of supervision used in our pipeline \(Section 3\.1\): a verified\-safe trajectory from the Math domain, a constructive\-stream injected unsafe trajectory from the Coding domain, and a diagnostic\-stream natural\-failure unsafe trajectory from the Agentic domain\. Frame color encodes the safe/unsafe label, and the box title carries the source benchmark, the responsible agent on unsafe trajectories, and the decisive step index k∗k^\{\*\} where applicable\. We display only the agent\-controlled turns relevant to the trajectory’s outcome and elide routine user, environment, and intermediate handoff turns; the original step indices are preserved verbatim so that any omitted index can be identified at a glance\. Verified\-safe Math trajectory\. Math \(SAFE\) \| Source: MATH\-500 Framework: AutoGen Swarm \(MathSolver ↔\\leftrightarrow Verifier\)\. Stop sentinel: ANSWER\_VERIFIED\. Length: 12 turns\. Label: \(k∗,a∗\)=\(SAFE,∅\)\(k^\{\*\},a^\{\*\}\)=\(\\text\{SAFE\},\\emptyset\)\. The trajectory exhibits the verified\-safe pattern that all 𝒟safe\\mathcal\{D\}\_\{\\text\{safe\}\} records satisfy: an outcome\-correct prediction \(3\*sqrt\(13\) matches the gold answer\), a sentinel\-clean termination \(ANSWER\_VERIFIED\), and an independent re\-derivation by a second agent\. The Verifier at Step 8 does not merely echo the MathSolver’s answer; it issues its own compute call before invoking verify\_answer, providing the independent\-verification evidence required by the curation predicates of Appendix B\.3\.1\. No tool error or coherence\-predicate violation is recorded across the twelve turns\. Constructive\-stream unsafe Coding trajectory\. Coding \(UNSAFE, injected\) \| Source: HumanEval\+ task 10 Framework: AutoGen Swarm \(CodeWriter ↔\\leftrightarrow CodeTester\)\. Length: 11 turns\. Label: \(k∗,a∗\)=\(9,CodeTester\)\(k^\{\*\},a^\{\*\}\)=\(9,\\text\{CodeTester\}\), fault category == verdict\_misread\. The constructive stream rewrites Step 9 to produce a verdict that contradicts the tool result returned at Step 8\. The run\_tests environment response is unambiguous \(ALL\_TESTS\_PASSED\), yet the CodeTester fabricates an off\-by\-one concern and emits TESTS\_FAILED at Step 10, flipping the trajectory outcome\. This is the canonical verdict\_misread fault category from the Coding catalog 𝒞Coding\\mathcal\{C\}\_\{\\text\{Coding\}\} \(Appendix B\.3\.1\), and the by\-construction label assigns the decisive step to Step 9 with the CodeTester as the responsible agent\. Diagnostic\-stream unsafe Agentic trajectory\. Agentic \(UNSAFE, diagnosed\) \| Source: HotpotQA Framework: Smolagents \(Manager →\\to search\_agent\)\. Length: 8 turns\. Label: \(k∗,a∗\)=\(4,search\_agent\)\(k^\{\*\},a^\{\*\}\)=\(4,\\texttt\{search\\\_agent\}\), fault category == wrong\-granularity retrieval\. On this naturally failed trajectory the Manager’s plan in Step 1 is well\-posed and the Smolagents delegation in Step 3 reaches the correct primary source\. The decisive failure is committed by the search\_agent at Step 4, which conflates the two halves of the question and returns the hotel’s name in place of its geographic location, after which the Manager in Step 5 propagates this answer verbatim\. The diagnostic\-stream propose\-and\-verify pipeline \(Appendix B\.3\.1\) localises this trajectory to \(k∗,a∗\)=\(4,search\_agent\)\(k^\{\*\},a^\{\*\}\)=\(4,\\texttt\{search\\\_agent\}\) with strict\-support voting on the four binary criteria \(sexists,ssubstantive,sdecisive,searliest\)\(s\_\{\\text\{exists\}\},s\_\{\\text\{substantive\}\},s\_\{\\text\{decisive\}\},s\_\{\\text\{earliest\}\}\), where Step 4 is the earliest step at which an answer\-determining error commits\. Appendix G Discussions G\.1 External Auditing vs\. Agent Self\-Reflection A natural alternative to the external\-auditor design of AgentForesight is to delegate the audit to the agent itself, asking the underlying policy to reflect on each prefix and decide whether to continue\. We deliberately reject this design for four complementary reasons, supported by an empirical anchor that is already visible in Table 1\. Generator\-verifier asymmetry plus auditor specialization\. Auditing a multi\-agent prefix is strictly easier than producing one\. The auditor only has to judge whether the trajectory remains on track, while the underlying agents must plan, retrieve, compute, and coordinate\. This generator\-verifier gap is well documented in process supervision for reasoning, where a small dedicated verifier matches or outperforms a much larger generator’s self\-check \[5, 27\]\. The asymmetry is sharpened in our setting because the external auditor can be specialized, with the prefix\-restricted observation contract of Section 2, the \(k∗,a∗\)\(k^\{\*\},a^\{\*\}\) supervision of AFTraj\-22K, and the three\-axis reward of Eq\. 11 all shaped around the audit objective\. None of these affordances are available to a base agent that must remain general\-purpose for task execution, so audit specialization comes for free in the external design and is mutually exclusive with the agent’s primary policy in a self\-reflection design\. Self\-reflection inherits the generator’s prior\. The agent emitted its CoT precisely because, under its current parameters, that reasoning was the most plausible continuation\. Asking the same parameters to re\-evaluate the same CoT reuses the same posterior, which preserves the same blind spots that produced the error in the first place\. The empirical record on intrinsic self\-correction is consistent with this analysis: iterative self\-feedback loops often fail to recover from upstream errors and can degrade correct intermediate steps \[18, 32, 34\]\. An external auditor breaks this loop by construction, since its parameters and supervision are independent of the agent that generated the prefix\. Self\-reflection competes with the agent’s primary objective\. Forcing the agent to reflect at every step layers a meta\-cognitive sub\-task on top of its task\-execution policy and introduces two failure modes\. First, over\-correction, in which reflection rewrites a correct step into an incorrect one, is recurrent across the self\-refinement literature \[48, 32\]\. Second, oscillation across multi\-step coordination, in which reflection\-induced revisions break state continuity between agent handoffs, is particularly costly in our multi\-agent setting where downstream agents condition on upstream outputs\. An external auditor decouples these objectives entirely, leaving the agent’s policy and execution dynamics untouched while still committing per\-step verdicts\. Modular operational benefits\. Externalizing the auditor yields three deployment properties that an in\-agent self\-reflection cannot replicate\. First, the auditor can be retrained or hot\-swapped as the underlying agent population evolves, without re\-training the agents themselves\. Second, the deployment threshold that trades false\-alarm rate against step accuracy can be tuned post\-hoc on the auditor alone, supporting the calibrated continue\-or\-alarm contract surfaced in Figure 4\. Third, the auditor’s verdicts are inspectable and never entangled with the agent’s hidden self\-talk, which mitigates the obfuscated\-reward\-hacking risk identified for monitor signals fed back into agent training \[2\] and aligns with the weak\-monitor\-over\-strong\-agent design pattern of \[22\], where a smaller specialized monitor reliably oversees a stronger underlying system\. Empirical anchor\. Our main results provide direct evidence for this design choice\. Reflexion\-7B in Table 1 instantiates self\-reflection on the same Qwen2\.5\-7B\-Instruct backbone that we use for AgentForesight\-7B; despite identical capacity, it reaches only 23\.3823\.38 overall Exact\-F1 with 3\.173\.17 ASS, whereas AgentForesight\-7B reaches 66\.4466\.44 Exact\-F1 and 0\.590\.59 ASS\. Holding the backbone fixed and varying only the audit paradigm, the external\-auditor design recovers a 2\.84×2\.84\\times Exact\-F1 improvement and a 5\.4×5\.4\\times tighter ASS, confirming that the gains documented in Section 4 are not artefacts of model scale but of the design choice to externalize and specialize the audit\. G\.2 Limitations Limitations\. We acknowledge two practical considerations of AgentForesight\-7B\. First, the online auditing protocol of Section 2 requires the auditor to be queried at every prefix of an unfolding trajectory, which introduces a lightweight operational dependency relative to a one\-shot post\-hoc attributor; we deliberately keep the auditor at 77B so that it remains practical to colocate with a host agent, and a similar per\-step monitoring contract is shared with concurrent runtime safeguards in agentic systems \[49, 19\], but deployments that forgo any per\-step instrumentation are out of our current scope\. Second, although AFTraj\-22K already spans three deployment\-faithful domains \(Coding, Math, and Agentic web/QA in Table 3\) and AgentForesight\-7B transfers cleanly to the external Who&When benchmark \(Table 2\), benchmarked coverage of failure structure in long\-horizon embodied agents and open\-ended scientific discovery remains an open frontier shared across the failure\-attribution literature rather than specific to our setting, and extending AFTraj\-22K and the three\-axis reward of Eq\. 11 to these domains is a natural follow\-up direction\. Despite these considerations, AgentForesight\-7B reaches 66\.4466\.44 average Exact\-F1 on AFTraj\-22K, 19\.8819\.88 points above the strongest proprietary baseline and a 3×3\\times tighter ASS, and retains its advantage on the external Who&When benchmark, indicating that the prefix\-restricted online auditing perspective coupled with our coarse\-to\-fine recipe offers a principled foundation for runtime audit of multi\-agent systems\. G\.3 Broader Impact The rapid deployment of multi\-agent LLM systems into real\-world workflows has substantially amplified the risk of cascaded failures spilling into tool calls, file writes, and external API requests with potentially irreversible side effects\. By committing a continue\-or\-alarm verdict at every step of an unfolding trajectory, AgentForesight\-7B directly contributes to mitigating these harms: a calibrated online auditor enables agent operators, platform owners, and downstream users to halt a trajectory at the moment a decisive error first becomes visible, supporting the integrity of agentic deployment in code\-execution sandboxes, mathematical reasoning, and tool\-augmented web and QA agents\. The three\-axis verdict on what, where, and who that we surface also yields interpretable evidence \(Figure 5\), which can be inspected and audited, in line with calls for transparent decision\-making in AI\-driven runtime safeguards\. We are aware of two potential negative effects worth noting\. First, any deployment\-time auditor can become a target of adaptive evasion: agents whose training signals overlap with auditing signals could attempt to mask decisive errors behind auditor\-friendly prefixes, so we recommend that production deployments combine AgentForesight with complementary signals such as post\-hoc attribution audits and provenance logs, and refresh the auditor as the underlying agent population evolves\. Second, false alarms, i\.e\., useful trajectories prematurely halted, can adversely affect agentic\-system end users and operators; deployers should expose calibrated alarm confidence rather than treat AgentForesight verdicts as hard kill switches, and pair the auditor with a tiered intervention policy or a human\-in\-the\-loop in safety\-critical settings\.`

Similar Articles

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv cs.AI

This paper introduces SearchAuditBench, a benchmark of 1,243 failed long-horizon search-agent trajectories with expert annotations, and SearchAuditor, a multi-perspective auditing framework that localizes, attributes, and repairs agent failures. Experiments show SearchAuditor outperforms baselines, achieving a 32.3% end-to-end pass rate with frontier models like GPT-5.5.

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

arXiv cs.AI

StepFinder is a lightweight framework that uses LLMs only in the feature construction phase to encode execution logs into temporal semantic sequences, then applies parameter-efficient temporal and attention modules for failure attribution in multi-agent systems. It reduces inference time by 79% compared to the fastest LLM-based method on the Who&When benchmark.