Pistis Technical Report
Summary
The Pistis technical report introduces a family of multimodal large language models (27B and 9B parameters) developed through a novel post-training framework combining distillation and reinforcement learning, with specialized variants for deep reasoning and agentic tasks.
View Cached Full Text
Cached at: 09/25/26, 09:27 AM
# Pistis Technical Report
Source: [https://arxiv.org/html/2609.28554](https://arxiv.org/html/2609.28554)
###### Abstract
We introduce thePistismodel family, comprising 27B\- and 9B\-parameter multimodal large language models built on Qwen3\.6 and Qwen3\.5, respectively, and developed through a general and scalable post\-training framework\. The framework first establishes a strong foundation through large\-scale multimodal supervised fine\-tuning \(SFT\)\. Building on this SFT foundation, we proposeInterleaved Distillation and Reinforcement Learning \(IDRL\), a novel post\-training paradigm that tightly integrates on\-policy distillation and reinforcement learning within a single training loop\. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit assignment for long\-horizon agentic trajectories, leading to stronger performance while mitigating common capability trade\-offs\. At both model scales, the framework produces two specialized variants:Pistis\-Thinking, designed to strengthen deep multimodal reasoning, andPistis\-Agentic, which additionally incorporates agentic trajectory data to support long\-horizon planning, iterative reasoning, and tool use\. Pistis\-Agentic is particularly strong in multimodal search\. Both scales outperform their corresponding base models\. Beyond model\-parameter optimization, we further introducePistis\-Auto\-Harnessing \(PAH\), a system\-level method that automatically improves the agent’s inference harness through iterative optimization\. Experiments demonstrate that PAH enhances the model performance without updating the model parameters or increasing the interaction budget\.
## Introduction
Figure 1:Overview of Pistis capabilities and performance\. The understanding and reasoning panel compares Pistis\-Thinking with the corresponding base models on MathVista\-mini, MM\-Vet, HallusionBench, OCRBench, AI2D\-test, RefCOCO\+ testA, Charades\-STA \(64 frames\), and MVBench \(8 frames\)\([Lu et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib2);[Yu et al\., 2023](https://arxiv.org/html/2609.28554#bib.bib20);[Guan et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib21);[Liu et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib3);[Kembhavi et al\., 2016](https://arxiv.org/html/2609.28554#bib.bib22);[Yu et al\., 2016](https://arxiv.org/html/2609.28554#bib.bib23);[Gao et al\., 2017](https://arxiv.org/html/2609.28554#bib.bib24);[Li et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib25)\)\. The agentic and search panel compares Pistis\-Agentic with the corresponding base models on MMSearch, BrowseComp\-VL, LiveVQA, PinchBench, TreeBench, HRBench4K, LogicVista, and MathVerse\-mini\([Jiang et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib27);[Geng et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib28);[Fu et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib29);[Kilo Code, 2026](https://arxiv.org/html/2609.28554#bib.bib31);[Wang et al\., 2025a](https://arxiv.org/html/2609.28554#bib.bib30);[Wang et al\., 2025c](https://arxiv.org/html/2609.28554#bib.bib17);[Xiao et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib18);[Zhang et al\., 2024b](https://arxiv.org/html/2609.28554#bib.bib26)\)\. All values match Tables[4](https://arxiv.org/html/2609.28554#S2.T4)and[5](https://arxiv.org/html/2609.28554#S3.T5)\.Multimodal large language models \(MLLMs\) have made rapid progress in understanding images, videos, and documents, driven by strong foundation architectures such as Qwen3\.5 and Qwen3\.6\. Despite these advances, post\-training remains a key bottleneck for unlocking higher\-level reasoning, robust generalization, and real\-world usability\. Existing pipelines usually apply supervised fine\-tuning and reinforcement learning in separate stages, leaving their complementary strengths underused: distillation provides dense supervision but can keep the student close to the teacher’s distribution; reinforcement learning optimizes task reward but often collapses policy entropy and destabilizes training; and long\-horizon agentic tasks require process supervision that a final\-answer reward alone cannot provide\.
In this work, we present thePistismodel family, comprising 27B\- and 9B\-parameter multimodal large language models built on Qwen3\.6 and Qwen3\.5, respectively, together with a general post\-training framework\. The framework first establishes a shared reasoning foundation through large\-scale multimodal SFT on structured reasoning data\. For the agentic variants, we augment this stage with agentic trajectories covering long\-horizon planning, iterative reasoning, and tool use across textual and multimodal domains\.
Building on this foundation, we proposeInterleaved Distillation and Reinforcement Learning \(IDRL\), a training paradigm that integrates on\-policy distillation and reinforcement learning within a single training loop\. Standard multimodal post\-training treats distillation and RL as disjoint stages, preventing their signals from interacting during optimization\. On\-policy distillation provides dense token\-level supervision by training the student to match the teacher’s distribution on student\-generated sequences\([Agarwal et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib14);[Lu and Thinking Machines Lab, 2025](https://arxiv.org/html/2609.28554#bib.bib15)\), but it is sensitive to the initial student–teacher overlap and can fail without a suitable cold start\([Li et al\., 2026](https://arxiv.org/html/2609.28554#bib.bib16)\)\. IDRL retains both objectives but alternates between their updates rather than applying them sequentially or merging them into a single weighted loss\. The two objectives can therefore shape each other across training phases without competing within the same update\. Empirically, this interleaving improves exploration, stabilizes optimization, and outperforms pure on\-policy distillation \(OPD\), pure RL, and their static combination in our ablations\. For long\-horizon agentic trajectories, IDRL also incorporates step\-level positive\-advantage suppression \(PAS\) to prevent rejected or ineffective intermediate actions from receiving positive credit solely because the final trajectory succeeds, thereby improving credit assignment\. IDRL thus transfers knowledge from strong teachers while optimizing task\-level objectives across diverse verifiable tasks, ranging from reasoning\-centric problems to interactive agentic settings such as tool\-integrated reasoning and deep research\.
Complementing this model\-level optimization, we introducePistis\-Auto\-Harnessing \(PAH\), a system\-level method that improves inference\-time orchestration while keeping the agentic policy and tool interface fixed\. An Optimization Agent uses development trajectories to propose and validate bounded harness revisions, retaining a candidate only when it improves a pre\-specified development metric; the selected harness is then frozen for evaluation\. We instantiate PAH on multimodal search, but its trace\-guided propose–validate–update procedure applies more broadly whenever an executable harness and measurable development feedback are available\.
Using this framework, we instantiate Thinking and Agentic variants at both the 27B and 9B scales\. The Thinking models build on reasoning\-data SFT to strengthen deep multimodal reasoning, whereas the Agentic models additionally incorporate agentic trajectories and are specialized for long\-horizon, tool\-integrated interaction\. On the 24 non\-grounding benchmarks, Pistis\-27B\-Thinking is comparable to Qwen3\.8\-27B \(82\.3 vs\. 82\.4\), while it achieves the highest grounding average of 80\.5, exceeding Qwen3\.8\-27B by 8\.6 points\. Pistis\-9B\-Thinking achieves the highest averages among the compared models at a similar scale on both benchmark groups\. The Agentic variants also achieve higher overall averages than their corresponding base models, with particularly strong gains in multimodal search\. Relative to Qwen3\.6\-27B, Pistis\-27B\-Agentic improves BrowseComp\-VL, MMSearch, VDR\-testmini, and LiveVQA by 8\.8, 3\.3, 3\.2, and 9\.7 points, respectively\. The corresponding gains for Pistis\-9B\-Agentic over Qwen3\.5\-9B are 8\.4, 8\.0, 3\.2, and 11\.4 points \(Figure[1](https://arxiv.org/html/2609.28554#S1.F1)and Table[5](https://arxiv.org/html/2609.28554#S3.T5)\)\. Together, these results demonstrate strong generalization across model scales and task settings\.
Our main contributions are:
- —Pistis model family\.We developPistis\-Thinkingmodels for deep multimodal reasoning andPistis\-Agenticmodels for long\-horizon, tool\-integrated interaction at both the 27B and 9B scales\.
- —Interleaved Distillation and Reinforcement Learning\.We propose IDRL, which alternates on\-policy distillation and RL updates within one training loop, rather than running them as separate stages or combining them in a static joint loss\. This design preserves policy entropy, reduces objective interference, and stabilizes optimization\. For long\-horizon agentic training, it incorporates PAS to improve credit assignment over intermediate actions\.
- —Pistis\-Auto\-Harnessing\.We propose an automated closed outer loop that iteratively improves the inference workflow, prompts, skills, evidence representation, and routing logic from development trajectories while keeping the Pistis\-Agentic policy and tool interface fixed\.
- —Empirical validation\.Across public multimodal and agentic benchmarks, the 27B and 9B Pistis models achieve higher overall averages than their corresponding base models, with particular strengths in multimodal search and Claw\-Style interaction\. Training\-dynamics and downstream ablations validate IDRL, including gains on the Claw\-Style task group\. PAH is further evaluated on a multimodal search benchmark, VDR\-testmini, under the same environment\-interaction limit\.
## Approach
Figure 2:Overview of Interleaved Distillation and Reinforcement Learning \(IDRL\)\.The Pistis family is built on Qwen3\.5\-9B and Qwen3\.6\-27B base models and follows a shared post\-training framework: large\-scale multimodal supervised fine\-tuning \(SFT\) followed by reinforcement\-based optimization\. The two scales share the same framework while using task\-specific data compositions and optimization configurations\. We first build a strong multimodal foundation from structured reasoning and agentic trajectories and then introduce Interleaved Distillation and Reinforcement Learning \(IDRL\) as our core algorithmic contribution\. At inference time, we further introduce our automated harness optimization method—Pistis\-Auto\-Harnessing \(PAH\)\. Pistis\-Thinking is trained with SFT on reasoning data and subsequently optimized on reasoning\-centric tasks, whereas Pistis\-Agentic additionally incorporates agentic trajectories during SFT and is specialized on interactive agentic tasks\. Finally, we describe the infrastructure supporting efficient training, evaluation, and inference\.
### 2\.1Supervised Fine\-Tuning
We construct a large\-scale, diverse multimodal SFT corpus of approximately 3\.2M QA pairs, comprisingreasoning dataandagentic trajectory data, together with a quality\-control pipeline that filters noisy samples and strengthens the resulting SFT models\. The reasoning data forms the shared SFT foundation for both models, while the agentic trajectory data is additionally included when fine\-tuning Pistis\-Agentic, allowing it to build agentic skills on top of the same reasoning foundation\.
Reasoning Data\.The reasoning data spans a broad range of domains and tasks, including mathematical reasoning; chart, figure, table, and document understanding; scientific and diagram reasoning; medical visual question answering; code generation; general logical reasoning; and general real\-world visual question answering \(including knowledge\-based and scene\-text reading\)\. The data are aggregated from public repositories and carefully designed synthetic prompts, with high\-quality responses generated by a strong teacher model under controlled prompting that elicits detailed step\-by\-step reasoning, grounded visual analysis, and consistent answer formulation\. Each instance adopts a structured output format in which the intermediate reasoning trace is enclosed by<think\>…</think\>, followed by two newline characters and the final answer without<answer\>tags; this separation enables reliable parsing and fine\-grained supervision over both reasoning quality and answer correctness\. All samples are normalized into a unified schema and pass a quality\-control pipeline: we validate tag usage, logical completeness, and a minimum reasoning length, remove malformed or underspecified samples, and, for tasks with reference answers, perform answer\-level consistency checks to discard mismatched or unverifiable outputs\.
Agentic Trajectory Data\.For the agentic SFT mixture, tool\-integrated reasoning \(TIR\), search, and general agent trajectories account for approximately 40%, 20%, and 40%, respectively\. TIR covers mathematical reasoning, visual perception, and visual\-logic tasks\. Search data includes both text\-based and multimodal retrieval\. The remaining trajectories are high\-quality open\-source examples spanning software engineering, code generation, tool use, and long\-horizon environment interaction\.
For self\-generated TIR and search data, we sample up to four response trajectories for each query and evaluate them with task\-specific rewards\. We discard queries for which all sampled trajectories fail or all succeed\. For a query with an intermediate success rate, we randomly retain one successful trajectory for SFT\. This selection focuses supervision on moderately difficult examples, for which a successful solution is informative while substantial room for improvement remains\. We further filter trajectories by final\-answer correctness and replace correct but low\-quality open\-source reasoning with cleaner regenerated trajectories when appropriate\.
Training Details\.We train with the AdamW optimizer using a base learning rate of1×10−51\\times 10^\{\-5\}and a cosine decay schedule\. To improve efficiency and reduce memory fragmentation, we adopt the Liger kernel, apply sequence packing up to a maximum length of 32,768 tokens, and preserve native\-resolution inputs for fine\-grained visual perception\. Training runs for three epochs with a global batch size of 1,536\. Supervised fine\-tuning takes approximately 9,216 GPU\-hours on the combined reasoning and agentic trajectory data\.
### 2\.2Interleaved Distillation and Reinforcement Learning
Modern large language models increasingly rely on post\-training to improve general capabilities and task performance\. However, existing paradigms often create limited synergy between supervised distillation and reinforcement learning, leaving their complementary strengths under\-exploited\.
To address this limitation, we propose Interleaved Distillation and Reinforcement Learning \(IDRL\), a training paradigm that integrates distillation and RL through an interleaved optimization process\. Specifically, IDRL alternates between On\-Policy Distillation \(OPD\) and RL during training, with flexible step budgets for each phase\. This design enables dynamic interaction between the two learning signals, rather than treating them as isolated or sequential stages\.
Intuitively, RL and OPD oppose each other: RL sharpens the policy toward high\-reward outputs, reducing entropy, whereas OPD draws the student toward the teacher’s broader distribution with a dense token\-level signal that preserves entropy\. Summing the objectives forces these opposing updates into a single gradient, where they compete directly, whereas alternating them lets each step follow a single objective while the two still influence each other across phases\. We formalize this schedule next, and analyze its effect on entropy and stability in the following subsections\.
Formally, letssdenote the current global training step\. We define the durations, or step budgets, for the OPD and RL phases asSOPDS\_\{\\text\{OPD\}\}andSRLS\_\{\\text\{RL\}\}, respectively\. The optimization strategyΠs\\Pi\_\{s\}at stepssis governed by a periodic scheduling function:
Πs=\{OPD,if\(smodS\)<SOPDands<S∗RL,otherwise\\Pi\_\{s\}=\\begin\{cases\}\\text\{OPD\},&\\text\{if \}\(s\\bmod S\)<S\_\{\\text\{OPD\}\}\\text\{ and \}s<S^\{\*\}\\\\ \\text\{RL\},&\\text\{otherwise\}\\end\{cases\}\(1\)
whereS=SOPD\+SRLS=S\_\{\\text\{OPD\}\}\+S\_\{\\text\{RL\}\}represents the total length of one full interleaved cycle\. We switch to pure RL afterS∗S^\{\*\}steps, whereS∗S^\{\*\}is determined by monitoring the entropy plateau\. The transition to pure RL is motivated by two observations: \(i\) once the entropy plateaus, continued OPD contributes little additional output diversity, and \(ii\) by continually regularizing the student toward the teacher’s distribution, it upper\-bounds the student at the teacher’s capability, foreclosing any super\-teacher gains\. We apply IDRL to the 9B variants, using reasoning\-centric mixtures for Pistis\-9B\-Thinking and interactive agentic mixtures for Pistis\-9B\-Agentic\. The corresponding 27B variants are optimized with pure RL under the same RL hyperparameters and serve as the frozen OPD teachers for 9B models\. For reproducibility, Table[1](https://arxiv.org/html/2609.28554#S2.T1)summarizes the training paradigm, teacher assignment, and optimization hyperparameters for each Pistis variant\.
Table 1:Post\-training configurations and compute\. The 9B variants use IDRL, whereas the 27B variants use pure RL\.SOPDS\_\{\\mathrm\{OPD\}\}andSRLS\_\{\\mathrm\{RL\}\}denote the numbers of OPD and RL steps in each interleaved cycle, respectively, andS∗S^\{\*\}denotes the step at which IDRL switches to pure RL\.ConfigurationPistis\-9BThinkingPistis\-27BThinkingPistis\-9BAgenticPistis\-27BAgenticTraining paradigmIDRLPure RLIDRLPure RLOPD teacherPistis\-27B\-Thinking–Pistis\-27B\-Agentic–SOPDS\_\{\\mathrm\{OPD\}\}\(steps/cycle\)5–5–SRLS\_\{\\mathrm\{RL\}\}\(steps/cycle\)5–5–Switch stepS∗S^\{\*\}300–100–Distillation weightα\\alpha1–1–JSD coefficientβ\\beta0\.5–0\.5–Teacher top\-kk50–50–Maximum responses length32k32k100k100kSAPO temperatureτpos\\tau\_\{\\mathrm\{pos\}\}1111SAPO temperatureτneg\\tau\_\{\\mathrm\{neg\}\}1\.051\.051\.051\.05Learning rate1e\-61e\-61e\-61e\-6
In the following, we detail the two components of IDRL, on\-policy distillation and reinforcement learning, and then analyze why interleaving them preserves entropy and stabilizes optimization\.
#### 2\.2\.1On\-Policy Distillation
We use on\-policy distillation to transfer teacher behavior to the student under the student’s own sampled contexts\. We formulate this distillation process following the Generalized Knowledge Distillation \(GKD\) framework of[Agarwal et al\. \(2024\)](https://arxiv.org/html/2609.28554#bib.bib14), instantiated with the Jensen–Shannon divergence \(JSD\) as the training objective\.
##### Setup\.
Letπθ\\pi\_\{\\theta\}denote the student policy being optimized, andπT\\pi\_\{T\}the teacher policy corresponding to a promptxxdrawn from distribution𝒟\\mathcal\{D\}\. Given an output sequenceyysampled on\-policy from the student, the token\-level discrepancy between teacher and student is measured by the generalized JSD at each positiontt:
ℓOPD\(t\)=DJSD\(β\)\(πT\(⋅∣x,y<t\)∥πθ\(⋅∣x,y<t\)\)=βDKL\(πT∥M\)\+\(1−β\)DKL\(πθ∥M\),\\ell^\{\(t\)\}\_\{\\text\{OPD\}\}=D\_\{\\mathrm\{JSD\}\(\\beta\)\}\\\!\\left\(\\pi\_\{T\}\(\\cdot\\mid x,y\_\{<t\}\)\\,\\\|\\,\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}\)\\right\)=\\beta\\,D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{T\}\\,\\big\\\|\\,M\\right\)\+\(1\-\\beta\)\\,D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{\\theta\}\\,\\big\\\|\\,M\\right\),\(2\)whereM=βπT\(⋅∣x,y<t\)\+\(1−β\)πθ\(⋅∣x,y<t\)M=\\beta\\,\\pi\_\{T\}\(\\cdot\\mid x,y\_\{<t\}\)\+\(1\-\\beta\)\\,\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}\)is the mixture distribution andβ∈\(0,1\)\\beta\\in\(0,1\)controls the asymmetry between the teacher\-to\-mixture and student\-to\-mixture KL terms\. Compared to the plain reverse KL objective, the JSD formulation is bounded and provides gradient signal from both directions, which we find leads to more stable training in practice\.
##### Top\-kkVocabulary Restriction\.
Computing the full\-vocabulary divergence at every token is expensive and dominated by near\-zero probability mass\. We therefore restrict the computation to the top\-kktokens selected by the*teacher*distribution at each decoding step\. Let𝒱t⊆𝒱\\mathcal\{V\}\_\{t\}\\subseteq\\mathcal\{V\}denote this top\-kkindex set with\|𝒱t\|=k\|\\mathcal\{V\}\_\{t\}\|=k\. We approximateπT\(v∣x,y<t\)≈0\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\\approx 0for allv∉𝒱tv\\notin\\mathcal\{V\}\_\{t\}\. In practice, this approximation is acceptable for current LLMs, whose output distributions are typically sharp and concentrated on a small fraction of the vocabulary\.
Under this approximation, we expand each KL term in the JSD separately over𝒱t\\mathcal\{V\}\_\{t\}and its complement𝒱∖𝒱t\\mathcal\{V\}\\setminus\\mathcal\{V\}\_\{t\}\. For the first term in Eq\.[2](https://arxiv.org/html/2609.28554#S2.E2), sinceπT\(v∣x,y<t\)≈0\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\\approx 0forv∉𝒱tv\\notin\\mathcal\{V\}\_\{t\}, the sum reduces directly to the top\-kktokens:
DKL\(πT∥M\)=∑v∈𝒱tπT\(v∣x,y<t\)logπT\(v∣x,y<t\)M\(v\)\+∑v∉𝒱tπT\(v∣x,y<t\)logπT\(v∣x,y<t\)M\(v\)⏟≈0\.D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{T\}\\,\\big\\\|\\,M\\right\)=\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}\+\\underbrace\{\\sum\_\{v\\notin\\mathcal\{V\}\_\{t\}\}\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}\}\_\{\\approx\\;0\}\.\(3\)For the second term, we first note that outside𝒱t\\mathcal\{V\}\_\{t\}, the mixture simplifies toM\(v\)=βπT\(v∣x,y<t\)\+\(1−β\)πθ\(v∣x,y<t\)≈\(1−β\)πθ\(v∣x,y<t\)M\(v\)=\\beta\\,\\pi\_\{T\}\(v\\mid x,y\_\{<t\}\)\+\(1\-\\beta\)\\,\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\approx\(1\-\\beta\)\\,\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\. Substituting this approximation intoDKL\(πθ∥M\)D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{\\theta\}\\,\\big\\\|\\,M\\right\):
DKL\(πθ∥M\)\\displaystyle D\_\{\\mathrm\{KL\}\}\\\!\\left\(\\pi\_\{\\theta\}\\,\\big\\\|\\,M\\right\)=∑v∈𝒱tπθ\(v∣x,y<t\)logπθ\(v∣x,y<t\)M\(v\)\+∑v∉𝒱tπθ\(v∣x,y<t\)logπθ\(v∣x,y<t\)M\(v\)\\displaystyle=\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}\+\\sum\_\{v\\notin\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}≈∑v∈𝒱tπθ\(v∣x,y<t\)logπθ\(v∣x,y<t\)M\(v\)\+log11−β∑v∉𝒱tπθ\(v∣x,y<t\)\\displaystyle\\approx\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}\+\\log\\frac\{1\}\{1\-\\beta\}\\sum\_\{v\\notin\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)=∑v∈𝒱tπθ\(v∣x,y<t\)logπθ\(v∣x,y<t\)M\(v\)\+log11−β\(1−∑v∈𝒱tπθ\(v∣x,y<t\)\)\\displaystyle=\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\log\\frac\{\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\}\{M\(v\)\}\+\\log\\frac\{1\}\{1\-\\beta\}\\left\(1\-\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\\right\)\(4\)Combining both terms, and omitting the shared conditioning on\(x,y<t\)\(x,y\_\{<t\}\)for readability, the per\-token OPD objectiveℓOPD\(t\)\\ell^\{\(t\)\}\_\{\\text\{OPD\}\}is:
ℓOPD\(t\)=β∑v∈𝒱tπT\(v\)logπT\(v\)M\(v\)\+\(1−β\)\[∑v∈𝒱tπθ\(v\)logπθ\(v\)M\(v\)\+log11−β\(1−∑v∈𝒱tπθ\(v\)\)\]\\ell^\{\(t\)\}\_\{\\text\{OPD\}\}=\\beta\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{T\}\(v\)\\log\\frac\{\\pi\_\{T\}\(v\)\}\{M\(v\)\}\+\(1\-\\beta\)\\left\[\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\)\\log\\frac\{\\pi\_\{\\theta\}\(v\)\}\{M\(v\)\}\+\\log\\frac\{1\}\{1\-\\beta\}\\left\(1\-\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\)\\right\)\\right\]\(5\)Note thatπθ\\pi\_\{\\theta\}is*not*renormalized over𝒱t\\mathcal\{V\}\_\{t\}\. The residual term1−∑v∈𝒱tπθ\(v\)1\-\\sum\_\{v\\in\\mathcal\{V\}\_\{t\}\}\\pi\_\{\\theta\}\(v\)instead analytically accounts for the student mass outside the teacher’s top\-kksupport, avoiding the distortion introduced by renormalization while remaining computationally efficient\.
##### Choice of divergence and top\-kk\.
The choice of divergence is what makes OPD an effective entropy regularizer inside IDRL\. Minimizing a reverse KLDKL\(πθ∥πT\)D\_\{\\mathrm\{KL\}\}\(\\pi\_\{\\theta\}\\,\\\|\\,\\pi\_\{T\}\)is*mode\-seeking*: it is zero\-forcing and drives the student to concentrate probability on a few teacher modes, which, like reward\-driven RL, reduces policy entropy\. A forward KLDKL\(πT∥πθ\)D\_\{\\mathrm\{KL\}\}\(\\pi\_\{T\}\\,\\\|\\,\\pi\_\{\\theta\}\)is instead*mass\-covering*, forcing the student to spread probability across the teacher’s support and thereby preserving entropy\. The generalized JSD interpolates between these two regimes, and the top\-kkset controls how much of the teacher’s support the student must cover\. With a sufficiently broad top\-kk, OPD pulls the student toward the teacher’s comparatively high\-entropy distribution and counteracts the entropy collapse induced by RL\. This mechanism matches the OPD\-variant ablation in Section[3](https://arxiv.org/html/2609.28554#S3)\(Figure[6](https://arxiv.org/html/2609.28554#S3.F6)\): the mode\-seeking reverse KL collapses entropy, the overly narrow top\-55JSD still collapses because coverage is insufficient, whereas the broader top\-5050JSD sustains high and stable entropy\. We therefore setk=50k=50\. Crucially, this top\-kkobjective remains informative only when the student already places appreciable mass on the teacher’s top\-kktokens, i\.e\., when the two distributions overlap; we find this condition breaks down at initialization in multimodal settings, which we address next\.
##### Applying OPD to Multimodal Settings\.
Directly applying OPD to MLLMs yields only limited gains\. As shown in Table[2](https://arxiv.org/html/2609.28554#S2.T2), with a Qwen3\.5\-9B student and a Qwen3\.6\-27B teacher on Geo3K, off\-policy SFT and on\-policy OPD improve the base student only marginally, from 65\.0 to 65\.4 and 65\.7, respectively\. The modest gain from OPD suggests that on\-policy sampling alone does not resolve the difficulty\. Instead, the objective may be limited by a mismatch between the initial student and teacher distributions\.
To investigate, we prepend a cold\-start SFT phase before OPD, fine\-tuning the student on teacher\-generated responses before switching to on\-policy OPD\. As shown in Figure[3](https://arxiv.org/html/2609.28554#S2.F3), direct OPD drops sharply at the start, whereas cold\-start OPD remains stable, improves consistently, and converges above the base student\. This result suggests thatthe initial distribution gap between student and teacher is the key obstacle, and that bridging it via cold\-start SFT is sufficient to restore effective distillation\. While[Li et al\. \(2026\)](https://arxiv.org/html/2609.28554#bib.bib16)identify low initial token\-distribution overlap as the governing failure condition for OPD in text\-only settings, our results show that the same mechanism extends to multimodal models, and that cold\-start SFT remains an effective remedy in this more challenging regime\.
ModelGeo3KQwen3\.6\-27B \(Teacher\)70\.2Qwen3\.5\-9B \(Student\)65\.0\+ off\-policy SFT65\.4\+ on\-policy OPD65\.7Table 2:Geo3K accuracy under different training configurations\.
Figure 3:Geo3K accuracy over training steps with different training strategies\.
#### 2\.2\.2Reinforcement Learning with Verifiable Reward
We conduct RL over a broad spectrum of textual and multimodal tasks whose outputs can be scored reliably, in most cases by predefined rules or executable programs\. Pistis\-Thinking is trained on reasoning and grounding tasks, spanning STEM and visual reasoning, image grounding, visual counting, and temporal video grounding\. For Pistis\-Agentic, IDRL focuses on two core capabilities: TIR and search, which account for approximately 60% and 40% of the mixture, respectively\. The TIR portion covers mathematical reasoning, visual perception, and visual\-logic reasoning; the search portion covers general web search, visual fact retrieval, and visual document retrieval\. The training corpus is assembled from open\-source and proprietary resources with strict preprocessing and human annotation\.
For the largest task family, visual reasoning, we initially curate roughly 80K candidate STEM problems from open\-source platforms and proprietary K\-12 data, using an LLM to remove proof\-style questions, convert multiple\-choice items into open\-ended form \(reducing reward hacking\), filter by model\-estimated difficulty, and exclude problems solvable without the image\. For multimodal queries, we sample 16 candidate responses per query from strong MLLMs and discard queries whose responses are all incorrect\. After source\-specific filtering and deduplication, this pool is combined with the other task sources; pilot RL experiments then prune sources with low improvement potential, yielding a final RL corpus of roughly 30K high\-quality queries\. During training, we again sample 16 responses per query, remove overly easy queries \(pass rate above 90%\), and merge task\-specific data into mixed\-task batches with a fixed, empirically tuned sampling ratio; grounding data combines general\-purpose and GUI tasks\.
##### Optimization with SAPO\.
For policy optimization we adopt SAPO\([Gao et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib19)\), a smooth, adaptive policy\-gradient algorithm\. For each queryxxwe sample a group ofGGresponses\{yi\}i=1G\\\{y\_\{i\}\\\}\_\{i=1\}^\{G\}fromπθold\\pi\_\{\\theta\_\{\\text\{old\}\}\}, score each with a verifiable rewardRiR\_\{i\}, and assign a group\-relative advantage shared across the tokens of a response,
Ai=Ri−mean\(\{Rj\}j=1G\)std\(\{Rj\}j=1G\)\.A\_\{i\}=\\frac\{R\_\{i\}\-\\mathrm\{mean\}\(\\\{R\_\{j\}\\\}\_\{j=1\}^\{G\}\)\}\{\\mathrm\{std\}\(\\\{R\_\{j\}\\\}\_\{j=1\}^\{G\}\)\}\.\(6\)SAPO then maximizes
𝒥RL\(θ\)=𝔼x∼𝒟,\{yi\}∼πθold\[1G∑i=1G1\|yi\|∑t=1\|yi\|fi,t\(ri,t\(θ\)\)Ai\],ri,t\(θ\)=πθ\(yi,t∣x,yi,<t\)πθold\(yi,t∣x,yi,<t\),\\mathcal\{J\}\_\{\\text\{RL\}\}\(\\theta\)=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\,\\\{y\_\{i\}\\\}\\sim\\pi\_\{\\theta\_\{\\text\{old\}\}\}\}\\\!\\left\[\\frac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\frac\{1\}\{\|y\_\{i\}\|\}\\sum\_\{t=1\}^\{\|y\_\{i\}\|\}f\_\{i,t\}\\\!\\big\(r\_\{i,t\}\(\\theta\)\\big\)\\,A\_\{i\}\\right\],\\qquad r\_\{i,t\}\(\\theta\)=\\frac\{\\pi\_\{\\theta\}\(y\_\{i,t\}\\mid x,y\_\{i,<t\}\)\}\{\\pi\_\{\\theta\_\{\\text\{old\}\}\}\(y\_\{i,t\}\\mid x,y\_\{i,<t\}\)\},\(7\)where, in place of the hard PPO\-style clipping used by GRPO, the importance ratio is reweighted by a smooth, temperature\-controlled gate
fi,t\(r\)=4τi,tσ\(τi,t\(r−1\)\),τi,t=\{τpos,Ai\>0,τneg,Ai≤0,f\_\{i,t\}\(r\)=\\frac\{4\}\{\\tau\_\{i,t\}\}\\,\\sigma\\\!\\big\(\\tau\_\{i,t\}\(r\-1\)\\big\),\\qquad\\tau\_\{i,t\}=\\begin\{cases\}\\tau\_\{\\text\{pos\}\},&A\_\{i\}\>0,\\\\\[2\.0pt\] \\tau\_\{\\text\{neg\}\},&A\_\{i\}\\leq 0,\\end\{cases\}\(8\)withσ\\sigmathe sigmoid and asymmetric temperaturesτpos,τneg\\tau\_\{\\text\{pos\}\},\\tau\_\{\\text\{neg\}\}\. The induced update weightfi,t′\(r\)=4σ\(τi,t\(r−1\)\)\(1−σ\(τi,t\(r−1\)\)\)f\_\{i,t\}^\{\\prime\}\(r\)=4\\,\\sigma\(\\tau\_\{i,t\}\(r\-1\)\)\\,\(1\-\\sigma\(\\tau\_\{i,t\}\(r\-1\)\)\)peaks atri,t=1r\_\{i,t\}=1\(on\-policy\) and decays smoothly as the ratio departs from11, so off\-policy updates are attenuated continuously rather than clipped discontinuously, forming a smooth trust region that stabilizes training across task types and model scales\.
##### Step\-Level Positive\-Advantage Suppression\.
Trajectory\-level rewards alone can incorrectly reinforce intermediate tool calls in a successful trajectory\. To improve credit assignment, we identify assistant steps that are rejected or do not contribute to later trajectory states, including malformed tool calls, repeated or substantially similar calls, calls issued after a final\-answer instruction, and calls whose observations cannot be incorporated into the context\. After advantage estimation, we set the positive advantages of all tokens in these steps to zero, while retaining zero or negative advantages\. Thus, an ineffective intermediate action is not positively reinforced even if the trajectory eventually receives a positive outcome, but it remains penalized in unsuccessful trajectories\. This asymmetric treatment separates the quality of intermediate actions from the final trajectory\-level outcome and yields more precise credit assignment for multi\-turn agent interaction\.
##### Reward System
We supervise RL with a hybrid, task\-aware reward formulation\. Our guiding principle is to prefer rewards that can be*automatically verified*by a rule or an executable check whenever the task admits one; such rewards are precise and reproducible, and they avoid the dominant failure mode of model\-based judging, where the policy learns to exploit the judge rather than solve the task\. Concretely, a global format reward first checks that the reasoning tags are present and correctly paired\. Tasks with a single correct answer, including STEM problems, chart numerical questions, and OCR, receive a binary correctness reward through rule\-based exact match, while spatial and temporal grounding receive a continuous IoU reward against the ground truth\. For open\-ended tasks where no deterministic check exists, namely long\-document QA, general VQA, and agentic search, we fall back to model\-based evaluation with a strong judge model\. Table[3](https://arxiv.org/html/2609.28554#S2.T3)summarizes the design for each task type\.
Table 3:Task\-aware reward design in the IDRL stage of Pistis, including a global format constraint and domain\-specific correctness rewards\.Reward scopeTask domainRuleModelBinaryReward design detailsFormatAll Domains✓✓Score11if<think\>and</think\>tags are present and strictly paired; otherwise00\.STEMMath✓✓Numerical calculation & multiple\-choice: Rule\-based exact match; score11for correct,00for incorrect\.Physics✓✓Chemistry✓✓Long DocumentChart & OCRLong Document✓Open\-ended QA: Model\-based evaluation using a strong judge model\.Chart✓✓Numerical calculation & multiple\-choice: Rule\-based exact match; score11for correct,00for incorrect\.OCR✓✓Rule\-based exact match; score11for correct,00for incorrect\.General VQAVQA✓Open\-ended QA: Model\-based evaluation using a strong judge model\.GroundingSpatial✓Score equals the IoU between the prediction and ground truth\.Temporal✓AgentSearch✓Model\-based evaluation using a strong judge model\.
#### 2\.2\.3Benefits of Interleaved Distillation and Reinforcement Learning
To understand why*interleaving*the two objectives is more effective than simply combining them, we examine the objective optimized at each step\. At stepss, the scheduleΠs\\Pi\_\{s\}defined above activates exactly one objective, so the per\-step objective can be written as the single indicator\-gated sum
ℒs\(θ\)=−1\[Πs=RL\]𝒥RL\(θ\)\+α1\[Πs=OPD\]ℒOPD\(θ\),\\displaystyle\\mathcal\{L\}\_\{s\}\(\\theta\)=\-\\,\\mathbbm\{1\}\[\\Pi\_\{s\}=\\text\{RL\}\]\\,\\mathcal\{J\}\_\{\\text\{RL\}\}\(\\theta\)\+\\alpha\\,\\mathbbm\{1\}\[\\Pi\_\{s\}=\\text\{OPD\}\]\\,\\mathcal\{L\}\_\{\\text\{OPD\}\}\(\\theta\),\(9\)where𝟙\[⋅\]\\mathbbm\{1\}\[\\cdot\]is the indicator function andα\>0\\alpha\>0weights distillation\. Here𝒥RL\(θ\)\\mathcal\{J\}\_\{\\text\{RL\}\}\(\\theta\)is the reward\-maximizing SAPO objective defined above, andℒOPD\(θ\)=𝔼x∼𝒟,y∼πθ\[1\|y\|∑tℓOPD\(t\)\]\\mathcal\{L\}\_\{\\text\{OPD\}\}\(\\theta\)=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\,y\\sim\\pi\_\{\\theta\}\}\\big\[\\tfrac\{1\}\{\|y\|\}\\sum\_\{t\}\\ell^\{\(t\)\}\_\{\\text\{OPD\}\}\\big\]is the OPD loss averaging the per\-token JSDℓOPD\(t\)\\ell^\{\(t\)\}\_\{\\text\{OPD\}\}of Eq\.[2](https://arxiv.org/html/2609.28554#S2.E2)\(mixtureM=βπT\+\(1−β\)πθM=\\beta\\pi\_\{T\}\+\(1\-\\beta\)\\pi\_\{\\theta\}, all distributions conditioned on\(x,y<t\)\(x,y\_\{<t\}\)\)\. Since exactly one indicator is nonzero at any step, IDRL never optimizes the two terms at once, in contrast to the joint objectiveℒjoint\(θ\)=−𝒥RL\(θ\)\+αℒOPD\(θ\)\\mathcal\{L\}\_\{\\text\{joint\}\}\(\\theta\)=\-\\mathcal\{J\}\_\{\\text\{RL\}\}\(\\theta\)\+\\alpha\\mathcal\{L\}\_\{\\text\{OPD\}\}\(\\theta\), which always sums them\. The distillation phases act as a teacher anchor loosely analogous to the KL\-to\-reference term in KL\-regularized RL, but stronger: the reference is a more capable teacher rather than a frozen copy of the policy, so distillation both transfers new capability and re\-broadens the output distribution\. Two properties explain why this alternation is preferable\.
\(i\) Distillation preserves entropy and provides a dense signal\.The OPD phases minimizeℒOPD\\mathcal\{L\}\_\{\\text\{OPD\}\}, which vanishes only whenπθ=πT\\pi\_\{\\theta\}=\\pi\_\{T\}and thus pulls the student toward the teacher distribution\. Through the identityDJSD\(β\)\(πT∥πθ\)=H\(M\)−βH\(πT\)−\(1−β\)H\(πθ\)D\_\{\\mathrm\{JSD\}\(\\beta\)\}\(\\pi\_\{T\}\\\|\\pi\_\{\\theta\}\)=H\(M\)\-\\beta H\(\\pi\_\{T\}\)\-\(1\-\\beta\)H\(\\pi\_\{\\theta\}\), whereH\(⋅\)H\(\\cdot\)denotes entropy and the teacher is fixed, the JSD objective includes an entropy\-related term while also depending on the mixture entropyH\(M\)H\(M\)\. Rather than relying on this term alone, the key effect is that the mass\-covering component of JSD encourages the student to cover a broader teacher support than mode\-seeking reverse KL, which empirically sustains higher policy entropy in our ablations\. This teacher\-anchored signal counteracts the entropy collapse driven by the reward term, keeping the policy’s outputs diverse and complementing the mode\-seeking\-versus\-mass\-covering view from the on\-policy distillation discussion above\. The distillation signal is also*dense*: it supplies a token\-level target at every position, unlike the sparse, sequence\-level scalar reward of RL, which improves credit assignment and stabilizes learning\.
\(ii\) Interleaving avoids gradient conflict\.Writing the two updates explicitly shows when summing them is harmful\. Consider a single decoding position with policy distributionπθ\(⋅∣x,y<t\)\\pi\_\{\\theta\}\(\\cdot\\mid x,y\_\{<t\}\)and teacherπT\(⋅∣x,y<t\)\\pi\_\{T\}\(\\cdot\\mid x,y\_\{<t\}\); both gradients are combinations of the token score functions∇θlogπθ\(v∣x,y<t\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(v\\mid x,y\_\{<t\}\)\. At the on\-policy point the SAPO gate satisfiesfi,t′\(1\)=1f\_\{i,t\}^\{\\prime\}\(1\)=1, so the RL gradient reduces to the policy gradient
gR=∇θ𝒥RL\(θ\)\|θ=θold=𝔼x∼𝒟,\{yi\}∼πθold\[1G∑i=1G1\|yi\|∑tAi∇θlogπθ\(yi,t∣x,yi,<t\)\],g\_\{R\}=\\nabla\_\{\\theta\}\\mathcal\{J\}\_\{\\text\{RL\}\}\(\\theta\)\\big\|\_\{\\theta=\\theta\_\{\\text\{old\}\}\}=\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\,\\\{y\_\{i\}\\\}\\sim\\pi\_\{\\theta\_\{\\text\{old\}\}\}\}\\\!\\Big\[\\tfrac\{1\}\{G\}\\sum\_\{i=1\}^\{G\}\\tfrac\{1\}\{\|y\_\{i\}\|\}\\sum\_\{t\}A\_\{i\}\\,\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(y\_\{i,t\}\\mid x,y\_\{i,<t\}\)\\Big\],\(10\)which forAi\>0A\_\{i\}\>0raises the probability of the realized tokena=yi,ta=y\_\{i,t\}and thus*concentrates*mass \(mode\-seeking\)\. Using∂DJSD\(β\)\(πT∥πθ\)/∂πθ\(v\)=\(1−β\)logπθ\(v\)M\(v\)\\partial D\_\{\\mathrm\{JSD\}\(\\beta\)\}\(\\pi\_\{T\}\\\|\\pi\_\{\\theta\}\)/\\partial\\pi\_\{\\theta\}\(v\)=\(1\-\\beta\)\\log\\frac\{\\pi\_\{\\theta\}\(v\)\}\{M\(v\)\}, the distillation ascent direction, holding the on\-policy samples fixed as in on\-policy distillation, is
gD=−∇θℒOPD\(θ\)\|θ=θold=−\(1−β\)𝔼x∼𝒟,y∼πθold\[1\|y\|∑t∑vπθ\(v\)logπθ\(v\)M\(v\)∇θlogπθ\(v\)\],g\_\{D\}=\-\\nabla\_\{\\theta\}\\mathcal\{L\}\_\{\\text\{OPD\}\}\(\\theta\)\\big\|\_\{\\theta=\\theta\_\{\\text\{old\}\}\}=\-\(1\-\\beta\)\\,\\mathbb\{E\}\_\{x\\sim\\mathcal\{D\},\\,y\\sim\\pi\_\{\\theta\_\{\\text\{old\}\}\}\}\\Big\[\\textstyle\\tfrac\{1\}\{\|y\|\}\\sum\_\{t\}\\sum\_\{v\}\\pi\_\{\\theta\}\(v\)\\log\\tfrac\{\\pi\_\{\\theta\}\(v\)\}\{M\(v\)\}\\,\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(v\)\\Big\],\(11\)which raises probability on tokens withπθ\(v\)<M\(v\)\\pi\_\{\\theta\}\(v\)<M\(v\)and thus*spreads*mass toward the teacher mixture \(mass\-covering\)\. The two act oppositely on the realized tokenaa: up to positive normalization, the RL contribution to the coefficient of∇θlogπθ\(a\)\\nabla\_\{\\theta\}\\log\\pi\_\{\\theta\}\(a\)is proportional toAi\>0A\_\{i\}\>0, whereas the OPD contribution is proportional to−\(1−β\)πθ\(a\)logπθ\(a\)M\(a\)\-\(1\-\\beta\)\\,\\pi\_\{\\theta\}\(a\)\\log\\frac\{\\pi\_\{\\theta\}\(a\)\}\{M\(a\)\}\. Whenever the policy is already more confident onaathan the teacher,πθ\(a\)\>πT\(a\)\\pi\_\{\\theta\}\(a\)\>\\pi\_\{T\}\(a\)\(equivalentlyπθ\(a\)\>M\(a\)\\pi\_\{\\theta\}\(a\)\>M\(a\)\), this OPD contribution is negative, so the two contributions have opposite signs; when such over\-confident tokens dominate the batch, the aggregated gradients conflict,gR⊤gD<0g\_\{R\}^\{\\top\}g\_\{D\}<0, precisely in the over\-sharpened regime that OPD is meant to correct\. The cost of this conflict is clearest in the first\-order improvement each update produces\. Summing the two signals,θ\+=θ\+η\(gR\+αgD\)\\theta^\{\+\}=\\theta\+\\eta\(g\_\{R\}\+\\alpha g\_\{D\}\)with step sizeη\\eta, moves the two objectives by
Δ𝒥RL≈η\(‖gR‖2\+αgR⊤gD\),Δ𝒥D≈η\(α‖gD‖2\+gR⊤gD\),\\Delta\\mathcal\{J\}\_\{\\text\{RL\}\}\\approx\\eta\\big\(\\\|g\_\{R\}\\\|^\{2\}\+\\alpha\\,g\_\{R\}^\{\\top\}g\_\{D\}\\big\),\\qquad\\Delta\\mathcal\{J\}\_\{\\text\{D\}\}\\approx\\eta\\big\(\\alpha\\\|g\_\{D\}\\\|^\{2\}\+g\_\{R\}^\{\\top\}g\_\{D\}\\big\),\(12\)where𝒥D:=−ℒOPD\\mathcal\{J\}\_\{\\text\{D\}\}:=\-\\mathcal\{L\}\_\{\\text\{OPD\}\}is the distillation objective\. The shared cross\-termgR⊤gDg\_\{R\}^\{\\top\}g\_\{D\}penalizes both updates: under conflict each objective improves less than it would in isolation, and oncegR⊤gD<−∥gR∥2/αg\_\{R\}^\{\\top\}g\_\{D\}<\-\\\|g\_\{R\}\\\|^\{2\}/\\alphathe summed step*lowers*the reward objective even though it descends the joint loss\. Interleaving never forms this cross\-term, because each step carries a single gradient: an RL step givesΔ𝒥RL≈η‖gR‖2≥0\\Delta\\mathcal\{J\}\_\{\\text\{RL\}\}\\approx\\eta\\\|g\_\{R\}\\\|^\{2\}\\geq 0and an OPD step givesΔ𝒥D≈ηα‖gD‖2≥0\\Delta\\mathcal\{J\}\_\{\\text\{D\}\}\\approx\\eta\\alpha\\\|g\_\{D\}\\\|^\{2\}\\geq 0, each non\-negative to first order regardless of the angle betweengRg\_\{R\}andgDg\_\{D\}, and vanishing only at a stationary point of the active objective\. Every interleaved update is therefore an uncorrupted, single\-objective step, whereas the summed update can stall or reverse one of the two\. This per\-step guarantee is what keeps optimization well\-behaved; the complementary, cross\-phase effect, in which OPD periodically restores the entropy and exploration that the multi\-step RL phases then exploit, is an empirical property, which we examine next\.
These properties are borne out empirically \(Section[3](https://arxiv.org/html/2609.28554#S3)\): OPD raises and sustains policy entropy where pure RL collapses it, and the interleaved variant shows smoother gradient norms and a clean alternating entropy pattern while the summed \(RL\+OPD\) variant is noisier \(Figure[7](https://arxiv.org/html/2609.28554#S3.F7)\)\. Once the student’s entropy plateaus, switching entirely to RL afterS∗S^\{\*\}steps leverages the enlarged exploration space for stronger and more stable improvement\.
### 2\.3Closed\-Loop Auto\-Harnessing for Multimodal Search
Auto\-Harnessing is a system\-level complement to IDRL for multimodal search\. We instantiate it for Pistis\-Agentic asPistis\-Auto\-Harnessing \(PAH\)\. Whereas IDRL improves the Pistis policy through parameter updates, PAH keeps the model, tool protocol, and evaluation entry point fixed and treats the surrounding inference harness as the optimization target\. PAH denotes the optimization procedure rather than a particular runtime system\. We call the initial system entering this procedure the Baseline Harness and the concrete system selected and frozen at the end of this application the Optimized Harness\. Figure[4](https://arxiv.org/html/2609.28554#S2.F4)separates the development\-time optimization procedure from the runtime behavior of its resulting harness\.
Outer Loop: Closed Harness Optimization by the Optimization Agentdevelopment set onlyCurrent bestharnessAttributionmechanism\-level failure analysisProposalone falsifiablechangeImplementationcode and prompt revisionCanary Gatesmalltargeted checkFull Evaluationcomplete dev set run12345Accept:new best harnessRollback:keep the bestnext round starts from the current best harness; the dev metric alone decides acceptancefreeze code, prompts, and configuration; one\-way evaluation on a disjoint test setFrozen Optimized Harness at Runtimeno Optimization Agent at runtimeWeb environmentsearch, visual retrieval, page readingPistis\-AgenticQuestionand imageFinal answeractions and observations,fixed interaction budgetbudget\-aware convergenceCandidate Ledgercandidates, evidence IDs, unmet constraintsSearch Skillsloaded only when the search is blockedCheckpoints and budget controlbounded prompts, reserved answer budget
Figure 4:Overview of Pistis\-Auto\-Harnessing \(PAH\)\. During development, an Optimization Agent runs a five\-stage closed loop on a fixed development set: it attributes failures, proposes one falsifiable change, implements it, checks it with a small canary run, and evaluates it on the complete development set\. A candidate replaces the current best harness only when the development metric improves; otherwise it is rolled back\. The Optimized Harness is then frozen and evaluated once on a disjoint test set\. At runtime, the Optimized Harness supports the frozen Pistis\-Agentic model with a Candidate Ledger, conditionally loaded Search Skills, and bounded checkpoints with budget control, while the model itself makes every decision and writes the final answer\.#### 2\.3\.1Auto\-Harnessing Procedure
PAH optimizes the harness’s code constraints, prompts, state representation, capability modules, routing rules, and workflow\. Before the loop begins, we register the editable interface, the development metric, and a fixed environment\-interaction budget\. The Pistis\-Agentic model remains frozen throughout, and no additional model is introduced into a task trajectory for routing, judging, voting, or answer selection\. The Optimization Agent operates only during development and is absent from runtime once the Optimized Harness is ready\.
One harness revision follows a five\-stage outer\-loop cycle: attribution, proposal, implementation, canary gate, and full evaluation\. The Optimization Agent reads the complete development\-set run of the current best harness, attributes recurring failures at the mechanism level, and writes a falsifiable proposal specifying its trigger condition, expected state change, and rollback condition\. Each round introduces one independently switchable change so that gains remain attributable\. After implementation, a small canary matched to the target failure type must reach a preset target before the candidate enters a complete development\-set evaluation under a fixed configuration\. A candidate replaces the current best harness only when the pre\-registered development metric improves; otherwise it is automatically rolled back, while mechanism activation rates and other process signals remain diagnostics only\. The development and final test sets contain disjoint samples; the test set is used once after harness optimization and never feeds back into proposal generation or version selection\.
#### 2\.3\.2Baseline and Resulting Optimized Harness
The components described below characterize the particular Optimized Harness obtained in this study, rather than constituting a fixed definition of PAH\. Applying the same optimization procedure under a different task distribution, tool protocol, or budget may produce a different harness structure\.
##### Baseline Harness\.
The Baseline Harness is a model\-driven recurrent state machine\. Given the question, image, prior messages, and previously returned images, the model either emits a<tool\_call\>or a final<answer\>; absent a valid tool call, the trajectory also terminates\. At most one parsed tool is executed per turn, including web and image search, reverse image search, page summarization, and Python\. Returned text is appended as a tool message and returned images are added to the multimodal context before the next model call\. Malformed calls, repeated actions, and highly similar queries trigger retry or warning rules, while tool errors are returned as observations\. A round or token\-budget limit disables further tool calls and forces an answer from the accumulated context\. This flexible loop leaves search decomposition and evidence retention to free\-form next\-action generation\.
##### Optimized Harness\.
The frozen Optimized Harness produced in this study augments the baseline loop with a structured Candidate Ledger, conditionally loaded Search Skills, evidence\-driven checkpoints, and budget\-aware convergence\. These mechanisms were proposed, tested, revised, or rejected by the PAH outer loop; they are properties of the resulting harness rather than manually prescribed components of the optimization method\.
##### Candidate Ledger\.
The Candidate Ledger is a bounded, model\-visible collection of structured candidate records injected into the model context\. Each record contains a candidate entity or answer hypothesis, supporting and refuting evidence IDs, source relations, and unmet constraints\. Every external result receives a stable evidence identifier, and a record can only cite valid identifiers, which keeps provenance traceable\. The harness orders candidates with a deterministic scorer and surfaces evidence gaps, while the same reasoning model still makes every selection and writes the final answer; the ledger assists the model and never answers in its place\. Evidence\-driven checkpoints ask the model to record its current best candidate once enough external results exist, and later updates happen only when new evidence substantially changes, refutes, or adds a candidate\. Checkpoints are bounded in number and budget, so they cannot form reflection loops\. This design turns the choice between continuing to search and answering into a decision constrained jointly by evidence sufficiency, candidate gaps, and the remaining budget\.
##### Search Skills and adaptive workflow\.
The resulting Search Skills are reusable operating procedures for recurring situations such as grounding visual candidates, disambiguating close candidates, tracing a relation chain, verifying the requested terminal field, repairing a failed query, or reconciling conflicting sources\. Each skill records a trigger condition, action steps, success criteria, and a stop condition\. It encodes a reusable procedure distilled from development trajectories, never an answer or a sample\-specific rule\.
At runtime the skills are loaded conditionally\. When the next action and its success condition are already clear, the model acts directly; a skill is loaded only when it would change the next evidence action, and each trajectory loads at most a small number of skills\. An adaptive workflow coordinates three controls: advisory checkpoints after early evidence actions, a hard recovery step that loads the query\-repair skill after a deterministic search failure, and the ledger checkpoint described above\. The workflow also separates exploration from final answering\. Near the interaction limit the harness stops opening new search branches, allows only a verification that can finish within the budget, then freezes further environment interaction and forces a final answer\. All checkpoints have count and budget caps, so the model can proceed when it ignores a prompt or when the remaining budget is low\.
The whole runtime follows a minimal intervention principle: by default the harness preserves the original reasoning path, and it adds a candidate prompt, a routing step, or a recovery action only when a replayable structural state triggers it\. The harness never constructs queries, candidates, or answers on behalf of the model, which keeps every state transition auditable from the trajectory\.
Runtime audits, transfer results, a Candidate Ledger case study, and limitations of the frozen Optimized Harness are provided in Appendix[A](https://arxiv.org/html/2609.28554#A1)\.
#### 2\.3\.3Relation to Automated Agent Optimization
PAH shares with ADAS the use of an LLM\-based optimizer to improve agent\-system implementations rather than model parameters\([Hu et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib32)\)\. Whereas ADAS emphasizes open\-ended invention of code\-defined agent architectures, PAH fixes the policy, task, environment interface, resource accounting, and evaluation protocol, then searches for auditable and reversible changes to the harness surrounding a single policy\.
Like AFlow, PAH treats code\-level control flow rather than only a single instruction as an optimization target\([Zhang et al\., 2024a](https://arxiv.org/html/2609.28554#bib.bib33)\)\. AFlow searches graphs of LLM\-calling nodes with Monte Carlo Tree Search; PAH requires neither a predefined multi\-node graph nor a specific search algorithm, and may revise state representation, tool routing, bounded checkpoints, recovery, budget control, termination logic, and their associated prompts\.
PAH also resembles GEPA in using complete trajectories and natural\-language reflection to diagnose failures and test revisions\([Agrawal et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib34)\)\. GEPA evolves prompts through reflective mutation and a Pareto frontier, whereas PAH may modify both deterministic code and prompts, evolves versions sequentially from the current best harness, and accepts a revision only when the fixed development metric improves\. PAH therefore focuses on state\-conditioned minimal intervention and auditable experimental governance for a frozen\-policy runtime, rather than proposing a general open\-ended, tree\-search, or Pareto\-evolution algorithm\.
### 2\.4Infrastructure for Training, Evaluation and Inference
The scalability of our framework rests on infrastructure for large\-scale training, checkpoint\-level evaluation, and efficient deployment, complementing the IDRL algorithm\. We describe the three components in turn\.
#### 2\.4\.1Training
We build the training infrastructure around reproducible, sandboxed environments with one\-click replication viauv, reducing setup and migration costs across the SFT and RL stages\. For continuous SFT, we adopt VeOmni\([Ma et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib12)\)with a padding\-free dynamic token\-budget batching scheme, meta\-device initialization, and rank\-0\-only checkpoint loading; combined with standard FSDP2 sharding, gradient checkpointing, mixed\-precision training, FlashAttention\-2, and optional Liger\-Kernel, this gives over 80% end\-to\-end speedup over a standard Hugging Face pipeline\.
For on\-policy distillation and RL, we adopt VeRL\([Sheng et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib13)\), whose asynchronous rollout pipeline decouples rollout generation from model optimization and executes them on separate resources, enabling sample generation and parameter updates to proceed in parallel\. This removes the long\-tail bottleneck of synchronous training, where updates are stalled by the slowest rollout\. Using vLLM as the rollout engine, whose high\-throughput serving and KV\-cache reuse amortize decoding overhead across batched rollouts, this design yields an overall∼\\sim2×\\timesend\-to\-end speedup over conventional synchronous RL pipelines\.
Figure 5:Architecture of PistisEvalKit\.
#### 2\.4\.2Evaluation
We evaluate Pistis withPistisEvalKit, built on the open\-source VLMEvalKit\([Duan et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib4)\)\. As shown in Figure[5](https://arxiv.org/html/2609.28554#S2.F5), it adopts a modular design with two decoupled components: aModelmodule that separates standardized pre/post\-processing \(the Wrapper, including agentic workflows such as tool calling\) from inference backends \(the Modeling layer, covering vLLM\([Kwon et al\., 2023](https://arxiv.org/html/2609.28554#bib.bib5)\), Hugging Face Transformers\([Wolf et al\., 2020](https://arxiv.org/html/2609.28554#bib.bib6)\), and API services\), and aBenchmarkmodule that remains compatible with VLMEvalKit benchmarks while extending to internal suites \(e\.g\., Pistis Benchmark\) and agentic tasks\.
For all publicly available models evaluated in our environment, we follow the inference configurations officially recommended by their respective model developers, including the reasoning mode and decoding parameters\. Qwen3\.8\-27B is evaluated with the officially recommendedxhighreasoning effort, while the other reasoning\-enabled baselines use their recommendedthinkingconfigurations\. For Pistis models, we use temperature0\.00\.0and top\-pp1\.01\.0to minimize sampling variance\. Results for Step3\-VL\-10B are taken directly from its technical report rather than reproduced in our evaluation environment\.
Two capabilities are tailored to our development workflow\. First, anautomated training\-evaluation loop: an SDK\-based client orchestrates large\-scale evaluation and automatically evaluates checkpoints during training, enabling fine\-grained monitoring and faster iteration\. Second,native agentic benchmarking: PistisEvalKit supports multi\-turn reasoning and tool calling via DeepEyesV2\([Hong et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib8)\)with an extensible tool ecosystem\. Configuration uses a Hydra\-based YAML system with Pydantic\([Pydantic Developers, 2024](https://arxiv.org/html/2609.28554#bib.bib7)\)typing for one\-command execution of large benchmark suites\.
#### 2\.4\.3Inference
We optimize Pistis online inference on top of vLLM at the system, scheduling, and kernel levels, improving inference speed by over 100% in some prefill scenarios\. Two optimizations follow vLLM’s design and are tuned for Qwen3\.5/Qwen3\.6 visual\-language workloads: anasynchronous pipelinethat overlaps CPU pre/post\-processing with GPU compute by converting synchronous operators, including H2D copies and boolean\-mask operations, to asynchronous forms, reaching up to 99% GPU utilization with multi\-stream; and amulti\-processfrontend/backend split that relieves the Python GIL for CPU\-bound request handling, tokenization, and image processing, using shared memory for large multimodal payloads\.
At the kernel level, we further applyparameter alignment\. NVIDIA’s Tensor Memory Accelerator \(TMA\) on the Hopper architecture requires parameter dimensions to be multiples of 128 to schedule its high\-performance operators\. We identify unaligned ViT parameters in the Qwen3\.5/Qwen3\.6\-based models and pad the affected dimensions to 128\-aligned sizes, enabling high\-performance operators on H20 GPUs\.
Table 4:Performance of Pistis\-27B, Pistis\-9B, and representative multimodal large language models, including Qwen3\.6\-27B, Qwen3\.8\-27B, Qwen3\.5\-9B, Step3\-VL\-10B, Qwen3\-VL\-8B\([Bai et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib1)\), Keye\-VL\-1\.5\([Kwai Keye Team, 2025](https://arxiv.org/html/2609.28554#bib.bib10)\), and InternVL3\.5\([Wang et al\., 2025b](https://arxiv.org/html/2609.28554#bib.bib11)\), on visual benchmarks\. Results for Step3\-VL\-10B are taken directly from its technical report, while all other results are obtained using our unified evaluation environment\. The final two rows report averages separately over 24 non\-grounding benchmarks and six grounding benchmarks\.CategoryBenchmark\\Block1\-1Pistis27B\\Block1\-1Qwen3\.627B\\Block1\-1Qwen3\.827B\\Block1\-1Pistis9B\\Block1\-1Qwen3\.59B\\Block1\-1Step3\-VL10B\\Block1\-1Qwen3\-VL8B\\Block1\-1Keye\-VL\-1\.58B\\Block1\-1InternVL3\.58Bthinkingthinkingxhighthinkingthinkingthinkingthinkingthinkingthinking\\Block4\-1STEMPuzzleMMMUval81\.082\.082\.376\.678\.078\.1171\.671\.473\.4ScienceQAval99\.598\.598\.799\.298\.1\-95\.897\.696\.7MathVistamini\{\}\_\{\\text\{mini\}\}87\.787\.587\.086\.885\.383\.9779\.281\.280\.8MathVersemini\{\}\_\{\\text\{mini\}\}87\.587\.087\.285\.984\.775\.7373\.170\.961\.9\\Block4\-1GeneralVQARealWorldQA84\.684\.285\.881\.481\.074\.4473\.273\.268\.6MMStar80\.581\.380\.179\.778\.977\.4875\.380\.366\.0MM\-Vet77\.976\.977\.178\.175\.2\-71\.173\.270\.0MME91\.689\.187\.389\.889\.3\-84\.686\.084\.6\\Block2\-1AlignmentHallusionBench69\.467\.769\.669\.366\.864\.9161\.864\.259\.1MMVP81\.785\.084\.079\.783\.068\.1678\.778\.772\.7\\Block8\-1DocumentUnderstandingTextVQAval89\.089\.389\.388\.689\.0\-85\.986\.283\.2AI2Dtest\{\}\_\{\\text\{test\}\}93\.292\.692\.692\.491\.289\.3585\.190\.083\.6ChartQAtest\{\}\_\{\\text\{test\}\}84\.886\.186\.484\.486\.0\-84\.282\.778\.1InfoVQAval93\.693\.893\.991\.391\.4\-84\.977\.679\.1DocVQAval96\.095\.795\.995\.894\.6\-93\.292\.592\.3OCRBench91\.288\.085\.391\.188\.986\.7583\.186\.684\.0CharXiv\(DQ\)val95\.095\.394\.893\.593\.3\-89\.378\.880\.5CharXiv\(RQ\)val80\.777\.780\.772\.574\.459\.5254\.346\.249\.3\\Block4\-1SpatialGroundingRefCOCOtestA95\.894\.494\.695\.592\.8\-93\.385\.394\.7RefCOCOtestB91\.689\.889\.591\.186\.8\-87\.474\.588\.7RefCOCO\+testA94\.092\.392\.593\.989\.7\-90\.282\.392\.4RefCOCO\+testB86\.784\.785\.186\.780\.5\-80\.768\.782\.4\\Block2\-1TemporalGroundingCharades\-STA64frame63\.256\.934\.563\.654\.1\-57\.522\.223\.8TACoS128frame51\.540\.335\.049\.834\.1\-34\.43\.04\.4\\Block1\-1Multi\-ImageBLINK73\.674\.080\.771\.670\.066\.7963\.356\.157\.7\\Block4\-1VideoUnderstandingMVBench8frame70\.670\.471\.968\.967\.8\-66\.256\.967\.5TempCompass8frame83\.181\.885\.881\.079\.8\-76\.672\.872\.1MLVU64frame73\.974\.971\.872\.771\.9\-68\.075\.071\.0Video\-MME64frame74\.075\.276\.170\.168\.2\-65\.273\.065\.4\\Block1\-1MultilingualMTVQAtest34\.634\.633\.333\.832\.1\-26\.725\.035\.2\\Block2\-1SummaryNon\-grounding Avg\. \(24\)82\.382\.082\.480\.680\.0\-74\.674\.072\.2Grounding Avg\. \(6\)80\.576\.471\.980\.173\.0\-73\.956\.064\.4
## Experiments
### 3\.1Comparison with Public Benchmarks
#### 3\.1\.1Perception and Reasoning Tasks
As shown in Table[4](https://arxiv.org/html/2609.28554#S2.T4), we report aggregate results separately over 24 non\-grounding benchmarks and six spatial or temporal grounding benchmarks\. On the non\-grounding subset, Pistis\-27B obtains an average score of 82\.3, slightly outperforming Qwen3\.6\-27B \(82\.0\) while remaining essentially on par with Qwen3\.8\-27B \(82\.4\)\. Across all 30 benchmarks, Pistis\-27B records 19 wins, 10 losses, and one tie against Qwen3\.6\-27B, and 18 wins, 11 losses, and one tie against Qwen3\.8\-27B\. The improvements are therefore broad but not uniform\. Pistis\-27B remains weaker on several benchmarks, including MMMU, MMVP, and ChartQA, and trails Qwen3\.8\-27B by 7\.1 points on BLINK\. Its largest margins are concentrated in spatial and temporal grounding, where its six\-benchmark average reaches 80\.5, compared with 76\.4 for Qwen3\.6\-27B and 71\.9 for Qwen3\.8\-27B\.
At the 9B scale, Pistis\-9B achieves a non\-grounding average of 80\.6, exceeding Qwen3\.5\-9B by 0\.6 points, and records 24 wins and six losses over the full set of 30 benchmarks\. It improves on Qwen3\.5\-9B across STEM reasoning, document understanding, and most video\-understanding tasks, while remaining weaker on MMVP \(79\.7 vs\. 83\.0\)\. As with the 27B model, the largest gains occur on grounding benchmarks: Pistis\-9B averages 80\.1 over the six grounding tasks, compared with 73\.0 for Qwen3\.5\-9B\. These results indicate that Pistis provides consistent but generally moderate gains outside grounding, together with substantially stronger spatial and temporal grounding capability\.
#### 3\.1\.2Agentic Tasks
Table 5:Agentic benchmark performance comparison among Pistis\-27B Agentic, Pistis\-9B Agentic, Qwen3\.6\-27B, Qwen3\.8\-27B, Qwen3\.5\-9B, DeepEyesV2\-7B\([Hong et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib8)\), and Thyme\-7B\([Zhang et al\., 2025](https://arxiv.org/html/2609.28554#bib.bib9)\)\. All results in this table are evaluated with prompts that expose the tool schema\. Consequently, scores on benchmarks overlapping with Table[4](https://arxiv.org/html/2609.28554#S2.T4)are not directly comparable because the prompting configurations differ\. MMSearch combines the multimodal and text\-only subsets; BrowseComp\-VL similarly combines Level 1 and Level 2\. We use a general\-purpose ReAct\-based harness in which the agent dynamically selects and sequences atomic tool calls\. This differs from the official MMSearch’s fixed search workflow and VDR\-testmini’s composite tools, which bundle operations such as image search and cropping\. We report the main score on PinchBench averaged over five runs\.CategoryBenchmarkPistis 27BAgenticQwen3\.627BQwen3\.827B xhighPistis 9BAgenticQwen3\.59BDeepEyesV27BThyme7BChartUnderstandingChartQAtest\{\}\_\{\\text\{test\}\}85\.086\.278\.283\.583\.788\.486\.1CharXiv\(DQ\)95\.095\.095\.092\.392\.378\.6\-CharXiv\(RQ\)77\.775\.580\.470\.264\.648\.9\-Real\-WorldPerceptionV\*94\.294\.891\.193\.791\.181\.882\.2TreeBench59\.851\.165\.755\.352\.342\.5\-OCRBench87\.186\.881\.587\.088\.1\-\-SeedBench\-2 Plus76\.976\.376\.675\.174\.478\.6\-HRBench4K91\.391\.088\.589\.087\.977\.977\.0HRBench8K89\.687\.689\.584\.985\.173\.872\.0MME\-RealWorld\-Lite63\.361\.566\.163\.659\.264\.964\.8MultimodalReasoningMathVistamini\{\}\_\{\\text\{mini\}\}87\.786\.187\.383\.582\.071\.970\.0MathVersemini\{\}\_\{\\text\{mini\}\}86\.985\.888\.184\.179\.852\.7\-LogicVista81\.078\.383\.770\.769\.448\.749\.0Search\-OrientedBrowseComp\-VL57\.248\.454\.653\.244\.8\-\-MMSearch78\.074\.774\.772\.764\.763\.7\-VDR\-testmini26\.823\.626\.624\.821\.6\-\-LiveVQA84\.775\.088\.382\.771\.3\-\-Claw\-StylePinchBench86\.585\.787\.877\.674\.8\-\-OverallAverage78\.375\.778\.074\.771\.5\-\-
As shown in Table[5](https://arxiv.org/html/2609.28554#S3.T5), Pistis\-27B\-Agentic and Pistis\-9B\-Agentic obtain higher overall averages than their corresponding Qwen base models across the 18 reported agentic benchmarks\. Pistis\-27B\-Agentic reaches 78\.3, compared with 75\.7 for Qwen3\.6\-27B, while Pistis\-9B\-Agentic reaches 74\.7, compared with 71\.5 for Qwen3\.5\-9B\. Multimodal search is a major source of improvement over these base models\. Pistis\-27B\-Agentic exceeds Qwen3\.6\-27B by 8\.8 points on BrowseComp\-VL, 3\.3 points on MMSearch, 3\.2 points on VDR\-testmini, and 9\.7 points on LiveVQA\. The corresponding gains for Pistis\-9B\-Agentic over Qwen3\.5\-9B are 8\.4, 8\.0, 3\.2, and 11\.4 points, respectively\. On PinchBench, the reported mean scores increase from 85\.7 to 86\.5 at the 27B scale and from 74\.8 to 77\.6 at the 9B scale\. Beyond these interaction benchmarks, Pistis\-27B\-Agentic also exceeds Qwen3\.6\-27B by 8\.7 points on TreeBench and 2\.0 points on HRBench8K\.
Pistis\-27B\-Agentic and Qwen3\.8\-27B obtain similar overall scores \(78\.3 vs\. 78\.0\), with a small numerical advantage for Pistis\-27B\-Agentic and task\-dependent differences\. Pistis\-27B\-Agentic scores higher on ChartQA \(85\.0 vs\. 78\.2\), V\* \(94\.2 vs\. 91\.1\), OCRBench \(87\.1 vs\. 81\.5\), HRBench4K \(91\.3 vs\. 88\.5\), BrowseComp\-VL \(57\.2 vs\. 54\.6\), and MMSearch \(78\.0 vs\. 74\.7\), while Qwen3\.8\-27B scores higher on tasks including TreeBench \(65\.7 vs\. 59\.8\), LiveVQA \(88\.3 vs\. 84\.7\), and PinchBench \(87\.8 vs\. 86\.5\)\. Taken together, these results demonstrate the effectiveness of our post\-training recipe: starting from Qwen3\.6\-27B, Pistis\-27B\-Agentic raises the overall benchmark average from 75\.7 to 78\.3, achieving competitive aggregate performance against Qwen3\.8\-27B while retaining different strengths across individual tasks\.
### 3\.2Held\-Out Evaluation of the PAH\-Optimized Harness
We evaluate the frozenOptimized Harnessproduced by Pistis\-Auto\-Harnessing \(PAH\), using Pistis\-27B\-Agentic as the frozen policy\. During development, the outer loop described in Section[2\.3](https://arxiv.org/html/2609.28554#S2.SS3)runs on 100 development instances that share the source and distribution of the test set but contain no overlapping samples\. All attribution, proposals, canary runs, and version selection use only this development set\. The selected harness is then frozen, including its code, prompts, and configuration, and evaluated once on the VDR\-testmini\. In the primary matched\-budget comparison, the Optimized Harness and the Baseline Harness share the same limit of 15 environment interactions per trajectory, counting search, visual retrieval, and page reading; the 10\- and 30\-interaction Baseline runs are budget references only\.
As shown in Table[6](https://arxiv.org/html/2609.28554#S3.T6), the frozen Optimized Harness reaches 28\.6% accuracy, compared with 26\.8% for the Baseline Harness, a gain of 1\.8 percentage points\. The average number of actual environment interactions is nearly unchanged \(8\.60 versus 8\.68\), so the gain comes from how the fixed budget is spent\. We attribute the improvement to the frozen harness as a whole, covering the Candidate Ledger, the Search Skills, the adaptive workflow, and the budget design; the test set took no part in any version selection\. The result provides descriptive evidence for a complementary system\-level contribution: IDRL supplies the trained agentic policy, whereas PAH organizes how that fixed policy retrieves, preserves, and adjudicates evidence\. The same frozen harness also transfers to Seed\-2\.1\-turbo and GPT\-5\.5, yielding gains of 1\.4 and 4\.0 percentage points, respectively \(Appendix[A\.3](https://arxiv.org/html/2609.28554#A1.SS3)\)\.
A budget sweep provides additional descriptive context\. With maximum interaction limits of 10, 15, and 30, the Baseline Harness reaches 25\.2%, 26\.8%, and 28\.4% accuracy while using 6\.55, 8\.68, and 11\.66 average interactions, respectively\. The Optimized Harness reaches 28\.6% under the matched 15\-interaction limit while using 8\.60 interactions\. It is numerically 0\.2 percentage points higher than the 30\-interaction Baseline while using 26\.2% fewer interactions, suggesting that the observed difference is not attributable solely to a larger search budget\.
Table 6:Frozen Optimized Harness versus Baseline Harness on the VDR\-testmini, with Pistis\-27B\-Agentic as the frozen policy\. The primary comparison uses the same limit of 15 environment interactions per trajectory, and the test set took no part in version selection\. The 10\- and 30\-interaction Baselines are included only as budget references\.HarnessMax\. env\. interactionsAccuracy \(%\)Correct/TotalAvg\. env\. interactionsBaseline Harness1025\.2126/5006\.55Baseline Harness1526\.8134/5008\.68Baseline Harness3028\.4142/50011\.66Optimized Harness1528\.6143/5008\.60
##### Frozen cross\-benchmark transfer\.
We directly transfer the same frozen Optimized Harness to three additional search benchmarks without using their examples to optimize the harness or select a version\. As shown in Table[7](https://arxiv.org/html/2609.28554#S3.T7), performance improves on MMSearch, BrowseComp\-VL, and LiveVQA by 0\.4, 1\.7, and 1\.0 points, respectively\. The consistently positive direction provides evidence that the resulting candidate\-maintenance, evidence\-organization, and convergence procedures transfer beyond VDR\-testmini, although all evaluated tasks remain within multimodal search\.
Table 7:Zero\-target\-tuning transfer of the frozen Optimized Harness\. The harness is not revised or selected on any of the three target benchmarks\.BenchmarkBaseline HarnessOptimized HarnessGainMMSearch78\.078\.4\+0\.4BrowseComp\-VL57\.258\.9\+1\.7LiveVQA84\.785\.7\+1\.0
Together, these evaluations isolate harness\-level gains for fixed policies; we next return to the model\-level contribution and ablate the components of IDRL\.
### 3\.3Ablation Study
Impact of OPD variants\.Figure[6](https://arxiv.org/html/2609.28554#S3.F6)compares the training dynamics of entropy and gradient norm across three OPD variants: Reverse KL \(RKL\), JSD\-5, and JSD\-50\. RKL, which computes the KL divergence solely from the current predicted token\([Lu and Thinking Machines Lab, 2025](https://arxiv.org/html/2609.28554#bib.bib15)\), exhibits severe training instability with large, irregular gradient\-norm spikes\. This mode\-seeking behavior also concentrates probability mass on a narrow set of tokens, ultimately leading to entropy collapse as shown by the low entropy values after convergence\. Replacing RKL with our JSD formulation substantially alleviates gradient instability: both JSD\-5 and JSD\-50 maintain much smoother gradient norms\. However, JSD\-5, which restricts the divergence computation to the top\-5 tokens, still suffers from entropy collapse, suggesting that five candidate tokens are insufficient to prevent over\-concentration\. In contrast, JSD\-50 expands the divergence computation to the top\-50 tokens, providing a broader supervisory signal that prevents over\-concentration of probability mass\. Crucially, this wider token coverage acts as a threshold:once enough candidate tokens are covered, the model entropy rises and stabilizes at a higher level\. This elevated entropy is not merely a training artifact; it reflects more diverse output distributions that are essential for downstream reinforcement learning, where greater sampling diversity translates to richer exploration and more effective policy optimization\. Based on this analysis, we adopt top\-50 tokens for JSD computation in our framework\.
\(a\)\(b\)
Figure 6:Effect of OPD objective design on training stability\. We compare entropy \(a\) and gradient norm \(b\) across RKL, JSD\-5, and JSD\-50\. Broader top\-kkJSD preserves higher policy entropy and yields smoother gradients, while RKL and narrow\-support JSD exhibit entropy collapse or instability\.Effect of IDRL\.Starting from the same SFT checkpoint, we compare pure RL \(SAPO\), pure OPD, a two\-stage sequential pipeline that applies OPD followed by RL \(OPD→\\rightarrowRL\), their joint optimization \(RL\+OPD withℒjoint=−𝒥RL\+αℒOPD\\mathcal\{L\}\_\{\\text\{joint\}\}=\-\\mathcal\{J\}\_\{\\text\{RL\}\}\+\\alpha\\mathcal\{L\}\_\{\\text\{OPD\}\}\), and our interleaved variant \(IDRL\)\. The training dynamics in Figure[7](https://arxiv.org/html/2609.28554#S3.F7)show that pure RL undergoes steady entropy collapse, whereas OPD, RL\+OPD, and IDRL maintain substantially higher entropy\. IDRL exhibits phase\-wise entropy variation consistent with its alternating schedule: entropy tends to decrease during RL phases and recover during OPD phases\. In terms of optimization stability, pure RL develops a rising gradient norm and severe late\-stage spikes, while the methods incorporating OPD maintain lower and more bounded gradient norms\. These results suggest that interleaving preserves the complementary effects of RL and OPD without forcing their potentially conflicting gradients into every update\. Table[8](https://arxiv.org/html/2609.28554#S3.T8)provides the corresponding downstream results over 18 benchmarks\. IDRL achieves the highest reported overall average of 74\.7, compared with 74\.1 for sequential OPD→\\rightarrowRL and approximate averages of 74\.0 for pure RL, 74\.1 for joint RL\+OPD, and 73\.4 for pure OPD\. The comparisons with the joint and sequential variants examine two distinct alternatives to interleaving\. Relative to joint RL\+OPD, IDRL obtains higher scores in all five categories, with differences of 0\.7 on Chart Understanding, 0\.7 on Real\-World Perception, 0\.5 on Multimodal Reasoning, 0\.5 on Search\-Oriented tasks, and 0\.1 on PinchBench\. This pattern favors alternating the objectives over combining them within every update in the evaluated setting\. Relative to sequential OPD→\\rightarrowRL, IDRL obtains higher scores in four categories, with the largest numerical gain on Search\-Oriented tasks \(\+1\.9\) and a gain of 0\.4 on PinchBench, while scoring slightly lower on Real\-World Perception \(78\.4 vs\. 78\.5\)\. This comparison favors repeated alternation over a single transition from distillation to RL in terms of aggregate performance, though not in every category\. Across all evaluated variants, IDRL ranks highest in Chart Understanding, Multimodal Reasoning, Search\-Oriented tasks, and PinchBench\. Together with the observed training dynamics, these results support interleaving as a promising way to combine RL and OPD, without establishing that reduced gradient interference or improved exploration alone explains the downstream differences\.
\(a\)\(b\)
Figure 7:Effect of interleaving OPD and RL during post\-training\. Compared with vanilla RL, vanilla OPD, and the joint RL\+OPD objective, IDRL shows the expected phase\-wise entropy pattern \(a\) and more stable gradient norms \(b\), indicating that alternating the two objectives reduces optimization interference\.Table 8:Comparison of IDRL against vanilla RL, OPD, the sequential OPD\-then\-RL pipeline \(OPD→\\rightarrowRL\), and joint RL\+OPD optimization on Pistis\-9B\-Agentic across agentic benchmark categories\. All models start from the SFT checkpoint\. AVG gives equal weight to each of the 18 individual benchmarks\.CategorySFTRLOPDOPD→\\rightarrowRLRL\+OPDIDRLChart Understanding81\.081\.579\.681\.281\.382\.0Real\-World Perception77\.577\.878\.578\.577\.778\.4Multimodal Reasoning78\.579\.077\.878\.978\.979\.4Search\-Oriented57\.157\.356\.256\.557\.958\.4Claw\-Style77\.177\.574\.477\.277\.577\.6AVG73\.774\.073\.474\.174\.174\.7Effect of Positive\-Advantage Suppression\.We further study step\-level positive\-advantage suppression \(PAS\) within IDRL\. PAS prevents rejected or ineffective assistant steps from receiving positive reinforcement by setting their positive token advantages to zero, while retaining zero or negative advantages\. As shown in Table[9](https://arxiv.org/html/2609.28554#S3.T9), removing PAS causes the largest drops on tasks that place greater demands on multi\-step reasoning and long\-horizon interaction: Search\-Oriented performance decreases by 1\.2 points, Multimodal Reasoning by 0\.9 points, and PinchBench by 0\.8 points \(77\.6 vs\. 76\.8\)\. By contrast, the variant without PAS improves only marginally on Chart Understanding \(\+0\.2\) and Real\-World Perception \(\+0\.1\), where trajectories are typically shorter and intermediate credit assignment is less critical\. Overall, PAS improves the average score over 18 benchmarks by approximately 0\.4 points \(74\.7 vs\. 74\.3\)\. These results suggest that trajectory\-level rewards alone can incorrectly reinforce ineffective intermediate actions, with the resulting noise becoming more consequential as trajectories grow longer\. PAS mitigates this issue by assigning positive credit more selectively, thereby improving learning on reasoning\- and interaction\-intensive tasks\.
Table 9:Ablation of step\-level positive\-advantage suppression \(PAS\) within IDRL\. AVG gives equal weight to each of the 18 individual benchmarks\.Δ\\Deltadenotes the difference between IDRL without PAS and IDRL\.CategoryIDRLIDRLw/o PASΔ\\DeltaChart Understanding82\.082\.2\+0\.2Real\-World Perception78\.478\.5\+0\.1Multimodal Reasoning79\.478\.5\-0\.9Search\-Oriented58\.457\.2\-1\.2Claw\-Style77\.676\.8\-0\.8AVG74\.774\.3\-0\.4
### 3\.4Training Observations and Design Lessons
Generator–target model compatibility\.For SFT data construction, Qwen3\.6\-27B produced trajectories with a lower raw pass rate than Qwen3\.5\-397B\-A17B, yet training on its filtered data yielded stronger downstream performance\. This observation suggests that raw pass rate alone does not determine data quality: a generator that is closer to the target model family may yield a more compatible data distribution, which can be more important than maximizing the generator’s standalone success rate\.
Task\-specific interaction budgets\.We assign different maximum interaction rounds to different task families, using 5 rounds for TIR and 15 rounds for search\. A single shared budget can distort behavior: when TIR tasks are given an unnecessarily long horizon, the model may issue uninformative actions, such as generating blank images, merely to consume the available interaction rounds\.
High\-resolution perception for TIR\.Agentic training on high\-resolution images is particularly effective for TIR tasks\. On 8K desktop screenshots in HRBench and MME\-RealWorld\-Lite, allowing the agent to crop and inspect local regions supplies details that may be missed by a single global view, improving fine\-grained visual grounding and reasoning\.
Web\-page acquisition for search\.Search capability depends not only on query quality and retrieval accuracy, but also on extracting detailed evidence from retrieved snippets and web pages\. In practice, anti\-bot mechanisms can lower the success rate of page\-fetching tools\. Under RL, repeated failures to retrieve useful information discourage the model from invoking the tool, which can in turn reduce overall search capability\. Reliable page acquisition and fine\-grained information extraction should therefore be treated as first\-class components of the search environment\.
### 3\.5Comparison with Business Benchmarks
Motivation for building a business\-specific benchmark\.Recent multimodal foundation models have achieved strong performance on a wide range of general\-purpose benchmarks, demonstrating advances in perception, reasoning, and cross\-modal understanding\. However, these benchmarks primarily evaluate generic capabilities and do not capture the requirements of real\-world content\-safety and business\-integrity scenarios\. Such scenarios have several distinctive properties: \(1\) policy grounding, where decisions must align with explicit policy provisions; \(2\) fine\-grained semantic discrimination, often involving subtle or borderline cases; \(3\) multimodal and temporal reasoning or grounding, requiring joint interpretation of visual and textual signals over time; and \(4\) high\-stakes decision making, where errors have direct practical consequences\. As a result, performance on existing benchmarks does not reliably translate to effectiveness in these domains\.
Pistis Benchmark\.To address this gap, we construct a large\-scale, multimodal, business\-specific benchmark, namedPistis Benchmark, that systematically evaluates multimodal models in content\-safety and business\-integrity scenarios through provision\-driven tasks\. The benchmark is designed to \(i\) cover diverse business scenarios and provision definitions, \(ii\) assess a range of business\-specific atomic abilities across modalities and tasks, and \(iii\) support scalable and robust evaluation through approximately 10,000 high\-quality samples\. This framework enables comprehensive comparison between Pistis and competing models, while providing insights for model development and deployment\. Pistis Benchmark is compatible with VLMEvalKit\([Duan et al\., 2024](https://arxiv.org/html/2609.28554#bib.bib4)\)for standardized evaluation\.
Data engineering\.To ensure consistency and usability, all collected data are standardized at the video level and aligned with a unified schema\. To improve data quality and diversity, we perform multi\-stage multimodal deduplication\. At the frame level \(intra\-video\), key frames are selected based on hybrid similarity \(low\-level features plus deep embeddings\), reducing redundancy while preserving representative visual content\. At the video level \(inter\-video\), multimodal embeddings built from key frames, OCR, and ASR are used to remove semantically similar videos\. This process removes approximately 25% of redundant samples, increasing diversity and reducing evaluation bias\.
Question\-answer generation and filtering\.We formulate Pistis Benchmark primarily as multimodal VQA tasks, enabling flexible and structured evaluation\. The generation pipeline has two stages\. First, an LLM generator produces candidate QA pairs conditioned on the multimodal input and the relevant policy context\. Second, an LLM judge filters the candidates by relevance, accuracy, and quality\. The generated questions span multiple formats, including multiple\-choice, attribution, action recognition, captioning, reasoning, and grounding tasks\. Importantly, questions are designed to require cross\-modal reasoning and policy grounding, rather than surface\-level recognition\. Finally, we apply additional post\-processing to remove low\-quality samples and enforce consistency between the QA pair and the original video data\.
Benchmark composition and tasks\.Pistis Benchmark comprises approximately 10K high\-quality samples spanning both video \(71\.5%\) and image \(28\.5%\) inputs, reflecting the predominantly temporal nature of real\-world content\-safety scenarios\. The benchmark is organized into three question formats—multiple\-choice, open\-ended , and yes/no—and covers a broad spectrum of business\-specific atomic abilities\. Specifically, reasoning requires policy\-grounded inference over multimodal evidence, judging which provision a video may violate and resolving subtle, borderline cases from the interplay between narrative captions and visual composition; VQA covers general question answering over video and image content; and multilingual understanding poses identical policy\-relevant questions across languages to ensure consistent judgments\. OCR extracts on\-screen text verbatim in its original script, while captioning produces neutral, literal descriptions of observable visual content\. Action recognition identifies fine\-grained physical interactions and attributes, and attribution synthesizes structured product details into an accurate description\. This composition allows Pistis Benchmark to evaluate not only surface\-level perception but also fine\-grained, policy\-grounded, cross\-modal decision making\.
Evaluation results\.Table[10](https://arxiv.org/html/2609.28554#S3.T10)reports results for Pistis and other competitive models on Pistis Benchmark\. The results show that†Pistis\-9B \(fine\-tuned on business\-specific data\) obtains the highest aggregate score among the compared models \(89\.7 vs\. 86\.2 for Qwen3\.5\-9B and 84\.3 for Qwen3\-VL\-8B\-Thinking\)\. Pistis improves on OCR, grounding, action recognition, captioning, and video VQA, which are directly relevant to the content\-safety and business\-integrity scenarios evaluated by the benchmark\.
Table 10:Evaluation of business\-specific visual understanding capabilities for content safety and business integrity on Pistis Benchmark\.†Pistis\-9B is further fine\-tuned on business\-specific data, whereas all other models are evaluated in their original released form\.ModalityCapabilityInternVL3\.58BthinkingKeye\-VL\-1\.58BthinkingQwen3\-VL8BthinkingQwen3\.59BthinkingPistis9BthinkingPistis†9BthinkingImageOCR70\.872\.276\.681\.482\.182\.8Grounding52\.656\.871\.269\.773\.880\.3Attribute84\.296\.799\.399\.599\.699\.5Multilingual73\.985\.991\.590\.786\.592\.3VideoReasoning83\.184\.387\.386\.092\.589\.5Action recog\.91\.586\.082\.994\.592\.397\.0Caption72\.273\.365\.565\.569\.274\.9Multilingual90\.488\.190\.992\.290\.593\.3VQA94\.792\.893\.196\.297\.797\.8Avg\. Score79\.381\.884\.386\.287\.189\.7
## Failure Analysis and Future Directions
Beyond the aggregate results in Section[3](https://arxiv.org/html/2609.28554#S3), manual inspection of erroneous Pistis\-Agentic rollouts reveals three recurring process\-level weaknesses\. First, the model may recognize salient visual cues but bind them to an incorrect interpretation of the question\. Second, it may use a tool to confirm a preselected hypothesis rather than obtain a decision\-relevant measurement\. Third, retrieval and reasoning may remain loosely coupled: useful evidence can fail to constrain the final answer, while expressed uncertainty may not trigger retrieval\. These failures often arise even when low\-level perception is adequate, indicating that stronger benchmark performance does not by itself ensure reliable evidence use\.
These patterns help explain the design choices of PAH\. The Candidate Ledger preserves explicit candidate–evidence bindings, evidence\-driven checkpoints require the model to update this state before answering, and Search Skills provide conditional procedures when retrieval is blocked\. However, these mechanisms cannot recover an answer whose correct entity never enters the upstream candidate set\. Future work should therefore combine inference\-time orchestration with process\-level training signals that reward intent verification, decision\-relevant tool use, candidate recall, and consistency between retrieved evidence and final answers\. Detailed trajectories, visual examples, and a fuller discussion appear in Appendix[B](https://arxiv.org/html/2609.28554#A2)\.
## Conclusion
We presented the Pistis model family and the general post\-training framework behind it\. A general recipe, large\-scale supervised fine\-tuning followed by IDRL, produces two strong specialists that share a reasoning\-data SFT foundation: Pistis\-Thinking for deep multimodal reasoning, and Pistis\-Agentic, whose SFT stage additionally includes agentic trajectory data before IDRL specialization\. Pistis\-Agentic shows its strongest gains in multimodal search and also consistently improves Claw\-Style interaction at both scales\. In particular, Pistis\-9B\-Agentic exceeds Qwen3\.5\-9B by 2\.8 points on PinchBench, while Pistis\-27B\-Agentic exceeds Qwen3\.6\-27B by 0\.8 points on the same benchmark\. Our ablations further show that IDRL obtains the strongest aggregate Claw\-Style score and that removing PAS reduces it, connecting these interaction gains to the proposed training design\. More broadly, both Pistis variants perform strongly across multimodal and agentic benchmarks, and interleaving on\-policy distillation with reinforcement learning is more stable and effective than performing either alone or jointly optimizing their losses with static weights\.
Complementing these model\-level contributions, Pistis\-Auto\-Harnessing \(PAH\) provides a system\-level method for multimodal search: a closed outer loop in which an Optimization Agent proposes, validates, and accepts harness revisions on a development set while the model stays frozen\. The resulting Optimized Harness combines a Candidate Ledger, conditionally loaded Search Skills, and an adaptive workflow with budget\-aware termination, and improves VDR\-testmini accuracy from 26\.8% to 28\.6% under a matched environment interaction budget\. The same frozen Optimized Harness also improves MMSearch, BrowseComp\-VL, and LiveVQA without target\-set tuning and transfers positively to Seed\-2\.1\-turbo and GPT\-5\.5\. These results broaden the evidence within multimodal search, while transfer to other agent domains such as coding and Claw\-Style interaction remains unvalidated\. Together, the model\- and system\-level results suggest that specialized multimodal capabilities can be cultivated efficiently from a shared foundation, and we hope Pistis serves as a strong basis for future research\.
At the same time, the failure analysis in Section[4](https://arxiv.org/html/2609.28554#S4)shows that stronger benchmark performance does not eliminate process\-level weaknesses: Pistis\-Agentic can still mis\-bind question intent, use tools to confirm rather than measure, and leave retrieval and reasoning loosely coupled\. Closing this gap will require process\-level supervision that rewards intent verification, decision\-relevant tool use, and consistency between intermediate and final answers, all of which can be integrated into the RL phases of IDRL\. We view narrowing the distance between aggregate accuracy and reliable reasoning as a central direction for future work\.
## Contributors
Names within each group are listed alphabetically by surname\.
Core Contributors\.Heyun Chen, Xiaohan Lan, Jiaxi Li, Zhilin Lu, Qi She, Weiwen Xu, Fei Yu, Yujie Zhong\.
Contributors\.Jinghuan Chen, Zijian Feng, Siyu Jiao, Yiheng Lin, Xinhao Wang, Sihan Yang, Jieyu You, Changbin Zhang, Hengyu Zhang, Xudong Zhang, Yunqing Zhao, Shuai Zheng\.
## References
- Agarwalet al\.\(2024\)R\. Agarwal, N\. Vieillard, Y\. Zhou, P\. Stanczyk, S\. R\. Garea, M\. Geist, and O\. BachemOn\-policy distillation of language models: learning from self\-generated mistakes\.InICLR,Cited by:[§1](https://arxiv.org/html/2609.28554#S1.p3.1),[§2\.2\.1](https://arxiv.org/html/2609.28554#S2.SS2.SSS1.p1.1)\.
- Agrawalet al\.\(2025\)L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang, C\. Potts, K\. Sen, A\. G\. Dimakis, I\. Stoica, D\. Klein, M\. Zaharia, and O\. KhattabGEPA: reflective prompt evolution can outperform reinforcement learning\.arXiv preprint arXiv:2507\.19457\.Cited by:[§2\.3\.3](https://arxiv.org/html/2609.28554#S2.SS3.SSS3.p3.1)\.
- Baiet al\.\(2025\)S\. Bai, Y\. Cai, R\. Chen, K\. Chen, X\. Chen, Z\. Cheng, L\. Deng, W\. Ding, C\. Gao, C\. Ge,et al\.Qwen3\-VL technical report\.arXiv preprint arXiv:2511\.21631\.Cited by:[Table 4](https://arxiv.org/html/2609.28554#S2.T4)\.
- Duanet al\.\(2024\)H\. Duan, J\. Yang, Y\. Qiao, X\. Fang, L\. Chen, Y\. Liu, X\. Dong, Y\. Zang, P\. Zhang, J\. Wang,et al\.VLMEvalKit: an open\-source toolkit for evaluating large multi\-modality models\.InACM MM,Cited by:[§2\.4\.2](https://arxiv.org/html/2609.28554#S2.SS4.SSS2.p1.1),[§3\.5](https://arxiv.org/html/2609.28554#S3.SS5.p2.1)\.
- Fuet al\.\(2025\)M\. Fu, Y\. Peng, B\. Liu, Y\. Wan, and D\. ChenLiveVQA: live visual knowledge seeking\.arXiv preprint arXiv:2504\.05288\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Gaoet al\.\(2025\)C\. Gao, C\. Zheng, X\. Chen, K\. Dang, S\. Liu, B\. Yu, A\. Yang, S\. Bai, J\. Zhou, and J\. LinSoft adaptive policy optimization\.arXiv preprint arXiv:2511\.20347\.Cited by:[§2\.2\.2](https://arxiv.org/html/2609.28554#S2.SS2.SSS2.Px1.p1.1)\.
- Gaoet al\.\(2017\)J\. Gao, C\. Sun, Z\. Yang, and R\. NevatiaTALL: temporal activity localization via language query\.InICCV,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Genget al\.\(2025\)X\. Geng, P\. Xia, Z\. Zhang, X\. Wang, Q\. Wang, R\. Ding, C\. Wang, J\. Wu, Y\. Zhao, K\. Li, Y\. Jiang, P\. Xie, F\. Huang, and J\. ZhouWebWatcher: breaking new frontier of vision\-language deep research agent\.arXiv preprint arXiv:2508\.05748\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Guanet al\.\(2024\)T\. Guan, F\. Liu, X\. Wu, R\. Xian, Z\. Li, X\. Liu, X\. Wang, L\. Chen, F\. Huang, Y\. Yacoob, D\. Manocha, and T\. ZhouHallusionBench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision\-language models\.arXiv preprint arXiv:2310\.14566\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Honget al\.\(2025\)J\. Hong, C\. Zhao, C\. Zhu, W\. Lu, G\. Xu, and X\. YuDeepEyesV2: toward agentic multimodal model\.arXiv preprint arXiv:2511\.05271\.Cited by:[§2\.4\.2](https://arxiv.org/html/2609.28554#S2.SS4.SSS2.p3.1),[Table 5](https://arxiv.org/html/2609.28554#S3.T5)\.
- Huet al\.\(2024\)S\. Hu, C\. Lu, and J\. CluneAutomated design of agentic systems\.arXiv preprint arXiv:2408\.08435\.Cited by:[§2\.3\.3](https://arxiv.org/html/2609.28554#S2.SS3.SSS3.p1.1)\.
- Jianget al\.\(2024\)D\. Jiang, R\. Zhang, Z\. Guo, Y\. Wu, J\. Lei, P\. Qiu, P\. Lu, Z\. Chen, C\. Fu, G\. Song, P\. Gao, Y\. Liu, C\. Li, and H\. LiMMSearch: benchmarking the potential of large models as multi\-modal search engines\.arXiv preprint arXiv:2409\.12959\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Kembhaviet al\.\(2016\)A\. Kembhavi, M\. Salvato, E\. Kolve, M\. Seo, H\. Hajishirzi, and A\. FarhadiA diagram is worth a dozen images\.InECCV,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Kilo Code \(2026\)Kilo CodePinchBench: real\-world benchmarks for openclaw agents\.Note:GitHub repositoryCited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Kwai Keye Team \(2025\)Kwai Keye TeamKwai Keye\-VL\-1\.5 technical report\.arXiv preprint arXiv:2509\.01563\.Cited by:[Table 4](https://arxiv.org/html/2609.28554#S2.T4)\.
- Kwonet al\.\(2023\)W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. Gonzalez, H\. Zhang, and I\. StoicaEfficient memory management for large language model serving with PagedAttention\.InSOSP,Cited by:[§2\.4\.2](https://arxiv.org/html/2609.28554#S2.SS4.SSS2.p1.1)\.
- Liet al\.\(2024\)K\. Li, Y\. Wang, Y\. He, Y\. Li, Y\. Wang, Y\. Liu, Z\. Wang, J\. Xu, G\. Chen, P\. Luo, L\. Wang, and Y\. QiaoMVBench: a comprehensive multi\-modal video understanding benchmark\.InCVPR,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Liet al\.\(2026\)Y\. Li, Y\. Zuo, B\. He, J\. Zhang, C\. Xiao, C\. Qian, T\. Yu, H\. Gao, W\. Yang, Z\. Liu,et al\.Rethinking on\-policy distillation of large language models: phenomenology, mechanism, and recipe\.arXiv preprint arXiv:2604\.13016\.Cited by:[§1](https://arxiv.org/html/2609.28554#S1.p3.1),[§2\.2\.1](https://arxiv.org/html/2609.28554#S2.SS2.SSS1.Px4.p2.1)\.
- Liuet al\.\(2024\)Y\. Liu, Z\. Li, M\. Huang, B\. Yang, W\. Yu, C\. Li, X\. Yin, C\. Liu, L\. Jin, and X\. BaiOCRBench: on the hidden mystery of OCR in large multimodal models\.Science China Information Sciences\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Lu and Thinking Machines Lab \(2025\)K\. Lu and Thinking Machines LabOn\-Policy Distillation\.Note:Thinking Machines Lab: ConnectionismAvailable:[https://thinkingmachines\.ai/blog/on\-policy\-distillation/](https://thinkingmachines.ai/blog/on-policy-distillation/)External Links:[Document](https://dx.doi.org/10.64434/tml.20251026)Cited by:[§1](https://arxiv.org/html/2609.28554#S1.p3.1),[§3\.3](https://arxiv.org/html/2609.28554#S3.SS3.p1.1)\.
- Luet al\.\(2024\)P\. Lu, H\. Bansal, T\. Xia, J\. Liu, C\. Li, H\. Hajishirzi, H\. Cheng, K\. Chang, M\. Galley, and J\. GaoMathVista: evaluating mathematical reasoning of foundation models in visual contexts\.InICLR,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Maet al\.\(2025\)Q\. Ma, Y\. Zheng, Z\. Shi, Z\. Zhao, B\. Jia, Z\. Huang, Z\. Lin, Y\. Li, J\. Yang, Y\. Peng,et al\.VeOmni: scaling any modality model training with model\-centric distributed recipe zoo\.arXiv preprint arXiv:2508\.02317\.Cited by:[§2\.4\.1](https://arxiv.org/html/2609.28554#S2.SS4.SSS1.p1.1)\.
- Pydantic Developers \(2024\)Pydantic DevelopersPydantic: data validation using python type hints\.External Links:[Link](https://docs.pydantic.dev/latest/)Cited by:[§2\.4\.2](https://arxiv.org/html/2609.28554#S2.SS4.SSS2.p3.1)\.
- Shenget al\.\(2025\)G\. Sheng, C\. Zhang, Z\. Ye, X\. Wu, W\. Zhang, R\. Zhang, Y\. Peng, H\. Lin, and C\. WuHybridFlow: a flexible and efficient RLHF framework\.InEuroSys,Cited by:[§2\.4\.1](https://arxiv.org/html/2609.28554#S2.SS4.SSS1.p2.1)\.
- Wanget al\.\(2025a\)H\. Wang, X\. Li, Z\. Huang, A\. Wang, J\. Wang, T\. Zhang, J\. Zheng, S\. Bai, Z\. Kang, J\. Feng, Z\. Wang, and Z\. ZhangTraceable evidence enhanced visual grounded reasoning: evaluation and methodology\.arXiv preprint arXiv:2507\.07999\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Wanget al\.\(2025b\)W\. Wang, Z\. Gao, L\. Gu, H\. Pu, L\. Cui, X\. Wei, Z\. Liu, L\. Jing, S\. Ye, J\. Shao,et al\.InternVL3\.5: advancing open\-source multimodal models in versatility, reasoning, and efficiency\.arXiv preprint arXiv:2508\.18265\.Cited by:[Table 4](https://arxiv.org/html/2609.28554#S2.T4)\.
- Wanget al\.\(2025c\)W\. Wang, L\. Ding, M\. Zeng, X\. Zhou, L\. Shen, Y\. Luo, W\. Yu, and D\. TaoDivide, conquer and combine: a training\-free framework for high\-resolution image perception in multimodal large language models\.InAAAI,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Wolfet al\.\(2020\)T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz,et al\.Transformers: state\-of\-the\-art natural language processing\.InEMNLP: System Demonstrations,Cited by:[§2\.4\.2](https://arxiv.org/html/2609.28554#S2.SS4.SSS2.p1.1)\.
- Xiaoet al\.\(2024\)Y\. Xiao, E\. Sun, T\. Liu, and W\. WangLogicVista: multimodal LLM logical reasoning benchmark in visual contexts\.arXiv preprint arXiv:2407\.04973\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Yuet al\.\(2016\)L\. Yu, P\. Poirson, S\. Yang, A\. C\. Berg, and T\. L\. BergModeling context in referring expressions\.InECCV,Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Yuet al\.\(2023\)W\. Yu, Z\. Yang, L\. Li, J\. Wang, K\. Lin, Z\. Liu, X\. Wang, and L\. WangMM\-Vet: evaluating large multimodal models for integrated capabilities\.arXiv preprint arXiv:2308\.02490\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Zhanget al\.\(2024a\)J\. Zhang, J\. Xiang, Z\. Yu, F\. Teng, X\. Chen, J\. Chen, M\. Zhuge, X\. Cheng, S\. Hong, J\. Wang, B\. Zheng, B\. Liu, Y\. Luo, and C\. WuAFlow: automating agentic workflow generation\.arXiv preprint arXiv:2410\.10762\.Cited by:[§2\.3\.3](https://arxiv.org/html/2609.28554#S2.SS3.SSS3.p2.1)\.
- Zhanget al\.\(2024b\)R\. Zhang, D\. Jiang, Y\. Zhang, H\. Lin, Z\. Guo, P\. Qiu, A\. Zhou, P\. Lu, K\. Chang, P\. Gao, and H\. LiMathVerse: does your multi\-modal LLM truly see the diagrams in visual math problems?\.arXiv preprint arXiv:2403\.14624\.Cited by:[Figure 1](https://arxiv.org/html/2609.28554#S1.F1)\.
- Zhanget al\.\(2025\)Y\. Zhang, X\. Lu, S\. Yin, C\. Fu, W\. Chen, X\. Hu, B\. Wen, K\. Jiang, C\. Liu, T\. Zhang,et al\.Thyme: think beyond images\.arXiv preprint arXiv:2508\.11630\.Cited by:[Table 5](https://arxiv.org/html/2609.28554#S3.T5)\.
## Appendix AAdditional Analysis of the PAH\-Optimized Harness
This appendix analyzes the frozen Optimized Harness produced by PAH\. All statistics below are computed after the harness is frozen and do not feed back into proposal generation or version selection\. They characterize the resulting system rather than isolate the causal effect of any single component\.
### A\.1Runtime Audit of the Candidate Ledger and Search Skills
Table[11](https://arxiv.org/html/2609.28554#A1.T11)summarizes whether the two principal mechanisms actually enter the runtime control flow\. The Candidate Ledger is a high\-coverage path: 497 of 500 trajectories attempt at least one candidate record, and 490 record one successfully\. The first attempt occurs after 2\.61 environment interactions on average\. By contrast, Search Skills are selectively activated on 135 trajectories, usually after the search has already encountered difficulty\.
Table 11:Runtime audit of the frozen Optimized Harness on 500 VDR\-testmini trajectories\.ComponentStatisticValueCandidate LedgerTrajectories attempting a candidate record497/500 \(99\.4%\)Trajectories with a successful candidate record490/500 \(98\.0%\)Successful candidate records per trajectory1\.522Mean interactions before the first record attempt2\.61First attempt immediately after two interactions353/497 \(71\.0%\)Mean rendered candidate context2,455 charsSearch SkillsTrajectories loading at least one skill135/500 \(27\.0%\)Successful skill loads137Loads ofrepair\-search\-query134/137 \(97\.8%\)Mean interactions before the first skill load5\.51
Therepair\-search\-queryskill is a recovery procedure triggered after a deterministic search failure; it guides the model to diagnose why the previous query was unproductive and formulate a materially revised query before retrieval resumes\. No trajectory reaches the ledger context\-truncation or update\-count cap, indicating that state capacity is not the current bottleneck\. The audit instead reveals a useful division of labor: the Candidate Ledger is the routine evidence\-state mechanism, whereas Search Skills primarily form a sparse query\-recovery path\. The broader skill catalog is available but is not yet reliably exercised\. These activation statistics are descriptive rather than causal: difficult trajectories are more likely to trigger recovery, so lower accuracy among skill\-using trajectories would not imply that the skill itself causes failure\.
One possible explanation for the limited use of the broader skill catalog is that these multimodal\-search tasks share recurring solution procedures that the policy may already have learned during training, leaving limited room for additional guidance in the form of standard operating procedures \(SOPs\)\. However, the present audit does not establish that skills provide little benefit: sparse activation may also reflect limitations of the triggering and routing mechanisms\. Controlled skill ablations would be needed to distinguish these explanations and quantify the marginal contribution of skills\.
The audit also clarifies the semantics of ledger verification\. Averifiedcandidate denotes multi\-source support, not proof that visual identity, relation direction, every question premise, and the requested terminal field are all correct\. Increasing record frequency or source count alone is therefore not an appropriate optimization objective\.
### A\.2A Candidate Ledger Trajectory
Table[12](https://arxiv.org/html/2609.28554#A1.T12)presents a compact VDR\-testmini example in which the query asks how a dress pattern influenced a 1980s horror subgenre and which materials were used for robotic antagonists in a representative film\. The ledger does not discard the initial useful style hypothesis when its film hypothesis fails; it preserves the supported portion, exposes the missing material constraint, and allows later evidence to repair the entity\.
Table 12:Candidate evolution in a successful trajectory\. Evidence counts denote source\-bound items recorded by the ledger\.StageCandidate stateEvidence gap and updateRoleInitial hypothesisComic\-book/pop\-art style;*X\-Tro*as a tentative filmTwo items support the style direction, but the film cannot satisfy the robotic\-material constraint\.Preserve style; reject entityEntity repairSwitch to*Chopping Mall*; record fiberglass, foam, and supporting production detailsThree items establish the robotic antagonists and their construction, while the style\-to\-subgenre relation remains incomplete\.Repair entity; fill fieldEvidence closureGraphic, action\-oriented techno\-/sci\-fi horror; robots primarily made from fiberglass and foamSix items jointly cover the style, subgenre, film identity, antagonists, and requested materials without direct contradiction\.Support final answer
This trajectory illustrates the intended role of the ledger: it makes hypothesis revision and constraint coverage explicit while leaving every evidence action and the final answer to the same frozen reasoning model\.
### A\.3Transfer across Frozen Policy Models
The Pistis\-27B\-Agentic row in Table[13](https://arxiv.org/html/2609.28554#A1.T13)reproduces the primary matched\-budget comparison in Table[6](https://arxiv.org/html/2609.28554#S3.T6)\. We then retain the same frozen Optimized Harness while replacing the reasoning policy with Seed\-2\.1\-turbo and GPT\-5\.5\. No harness revision or model\-specific version selection is performed\. All three policies improve relative to their corresponding Baseline Harness, while their average environment interactions remain close to the matched baselines\.
Table 13:Transfer of the same frozen Optimized Harness across reasoning policies on VDR\-testmini\.Frozen policyBaseline acc\.Optimized acc\.GainAvg\. interactions \(base/opt\.\)Pistis\-27B\-Agentic26\.828\.6\+1\.88\.68 / 8\.60Seed\-2\.1\-turbo27\.629\.0\+1\.47\.31 / 7\.69GPT\-5\.528\.632\.6\+4\.06\.44 / 6\.49
For GPT\-5\.5, the paired comparison contains 47 positive and 27 negative flips among 74 discordant examples\. Its larger gain suggests that stronger reasoning policies may use structured state and conditional procedures more effectively, but the present experiments do not isolate this interaction causally\.
### A\.4Design Lessons and Limitations
Design lessons\.The optimization history yields four practical lessons\. First, mechanism activation is only a diagnostic: a revision is accepted only when the complete development\-set task metric improves\. Second, guidance should be conditionally exposed from replayable states; globally persistent prompts can perturb trajectories that were already correct\. Third, upstream identity and relation errors should be addressed before strengthening downstream field extraction, since stricter completion around a wrong entity can reinforce the wrong answer\. Fourth, each candidate version should be independently derived from the current best harness and remain fully reversible, so rejected mechanisms do not silently accumulate\.
Limitations\.The frozen Optimized Harness can still over\-commit to an early incumbent because it lacks a systematic evidence\-backed challenger test\. Apparent source diversity can also be overstated when several evidence identifiers derive from the same underlying material\. Relation endpoints and the direct binding between the input image and a textual entity remain incompletely certified\. Finally, budget\-aware termination can occasionally prevent one last decisive verification, while most skills other than query repair are not yet reliably activated\. These limitations concern the current Optimized Harness; whether PAH discovers different and stronger mechanisms for other agent domains, such as coding or Claw\-Style interaction, requires separate development/test studies\.
## Appendix BDetailed Failure Cases
This appendix expands the process\-level failure analysis summarized in Section[4](https://arxiv.org/html/2609.28554#S4)\. We present representative cases from multimodal reasoning, tool\-integrated reasoning, and agentic search\. Across these settings, low\-level perception is often adequate; the central weakness is how the reasoning process interprets the task, gathers decision\-relevant evidence, and binds that evidence to the final answer\.
### B\.1Multimodal Reasoning: Salient Cues Are Read, Intent Is Not
Figure 8:Representative failure cases of Pistis\-Agentic in multimodal reasoning\. \(a\) The age\-group labels are read correctly, but the question intent \(the gap between the highest and lowest plotted rates\) is parsed literally as a gap between ages\. \(b\) The drawn arrow, a salient symbolic cue, overrides the cross\-frame evidence that the circle never moves\. Model excerpts are abridged; red marks the faulty steps and final answers, and green marks the reference answers\.In Figure[8](https://arxiv.org/html/2609.28554#A2.F8)\(a\), the model identifies the age\-group labels correctly but interprets “difference in the age” literally as the gap between group ages, rather than the intended gap between the plotted rates, and therefore answers a different question\. In Figure[8](https://arxiv.org/html/2609.28554#A2.F8)\(b\), when asked whether the circle moves to the right across an ordered sequence of frames, the model latches onto the right\-pointing arrow drawn inside the circle and never compares the circle’s position across frames\. Neither failure is due to missing the salient visual elements: in both cases, the model anchors on the most literal cue instead of verifying its interpretation against the visual evidence\.
### B\.2Tool\-Integrated Reasoning: Tools Confirm Instead of Measure
Figure 9:Representative failure cases of Pistis\-Agentic in tool\-integrated reasoning\. In both rollouts, the code interpreter is used to confirm a preselected hypothesis rather than to measure: \(a\) only the hypothesized region is cropped and the candidate comparison is never performed; \(b\) the script prints a hard\-coded count, and the printed output is then cited as confirmation\. Excerpts are abridged; red marks the faulty steps\.Figure[9](https://arxiv.org/html/2609.28554#A2.F9)shows two rollouts with a code interpreter\. In \(a\), the model first forms the hypothesis that the feet are closest to the target object, then writes code that crops only the target and the presumed feet region; the waist, head, and back are never examined, so the visualization merely reinforces the prior\. In \(b\), the question requires a strict greater\-than\-90 comparison; the model eyeballs the bars as “just over 90”, writes a script that prints a hard\-coded count, and then cites the printed output as confirmation\. In both cases, the tool call is formulated to support a preselected answer rather than produce decision\-relevant measurements, and an executed cell that adds no information is treated as independent evidence\.
### B\.3Agentic Search: Retrieval and Reasoning Are Loosely Coupled
Figure 10:Representative failure cases of Pistis\-Agentic in agentic search\. \(a\) Retrieval succeeds and the correct judgment \(green\) surfaces in the reasoning, yet the final answer contradicts it\. \(b\) The model never invokes the available search tool and produces a hallucinated date with fabricated attribution\. Excerpts are abridged; red marks the faulty steps\.Figure[10](https://arxiv.org/html/2609.28554#A2.F10)illustrates two complementary failures of the search–reasoning loop\. In \(a\), visual search succeeds and the correct judgment even surfaces in the reasoning trace, yet the final answer contradicts it: the model over\-differentiates between closely related concepts and overrides its own intermediate conclusion, so the retrieved evidence does not bind the final answer\. In \(b\), the converse occurs: facing a time\-sensitive factual question, the model talks itself out of using the available search tool, enumerates candidate dates from memory while explicitly acknowledging uncertainty, and finally emits a hallucinated date with fabricated attribution\. The uncertainty is verbalized but never operationalized into a tool call\.
These observations help explain, rather than retrospectively motivate, the PAH design in Section[2\.3](https://arxiv.org/html/2609.28554#S2.SS3)\. Its Candidate Ledger keeps retrieved facts bound to explicit candidates with traceable sources, its checkpoints ask the model to update this state before answering, and its Search Skills supply operating procedures when retrieval is blocked\. The same audit also exposes a remaining limitation: if the correct entity never enters the upstream candidate set, downstream evidence organization cannot recover it\.
### B\.4Implications for Future Work
The cases suggest three complementary training directions\. First,*intent grounding*should teach the model to state and verify its interpretation of an underspecified question against visual evidence before committing to a computation\. Second,*measurement\-grounded tool use*should reward calls that generate decision\-relevant evidence, such as comparing all candidate regions or calibrating chart axes before a threshold judgment, while penalizing confirmatory no\-op calls\. Third,*retrieval–reasoning coupling*should reward candidate recall, consistency between intermediate conclusions and final answers, and the conversion of expressed uncertainty into targeted retrieval\. These rule\-checkable trajectory signals can be incorporated into the RL phases of IDRL, while its distillation phases transfer the corresponding behaviors from a stronger teacher\. We leave systematic quantitative evaluation of these failure modes to future work\.Similar Articles
Motif 3: Technical Report
Motif 3 is a 314B-parameter Mixture-of-Experts language model with 13.2B active parameters per token, featuring Grouped Differential Latent Attention and trained on 12.5T tokens, demonstrating competitive performance across reasoning, coding, and long-context tasks.
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
This technical report introduces Ling and Ring 2.6, a family of large language models at the trillion-parameter scale designed for efficient and instant agentic intelligence.
@_akhaliq: paper:
This technical report presents Ling-2.6 and Ring-2.6, a family of trillion-parameter models designed for efficient and instant agentic intelligence, featuring architectural upgrades like hybrid linear attention and specialized training methods including KPop reinforcement learning. All checkpoints are open-sourced.
Index SLM Technical Report
Bilibili releases Index-1.9B, a series of open small language models pre-trained on 2.8 trillion tokens, achieving competitive performance on benchmarks. The four models include base, pure (no instruction data), chat, and a character model with retrieval-augmented generation for role-playing.
P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
This paper introduces P3D-Bench, a benchmark for evaluating multimodal large language models on parametric 3D generation tasks, including text-to-3D, image-to-3D, and assembly-3D, with metrics for geometric precision, semantic alignment, and part-level structure.