HARDEN: 约束进化搜索用于更难且答案保留的评估案例
摘要
HARDEN 引入约束进化搜索,以生成更难且保留预期输出的语言模型评估案例,从而在多个基准测试中显著降低了准确性。
arXiv:2609.30571v1 Announce Type: new
Abstract: Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches along generated domain-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity. Across FinQA, PubMedQA, and ContractNLI and three Qwen3.5 model scales (35B-A3B, 122B-A10B, and 397B-A17B), HARDEN reduces task-model accuracy by 22.7% on average and by up to 49.9% relative to single-pass baselines using the same feasibility checks. These results show that evolutionary search can produce substantially harder valid evaluation cases.
查看缓存全文
缓存时间: 2026/09/28 09:40
# Constrained Evolutionary Search for Harder, Answer-Preserving Evaluation Cases
Source: [https://arxiv.org/html/2609.30571](https://arxiv.org/html/2609.30571)
###### Abstract
Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments\. We introduceHARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed\.HARDENsearches along generated domain\-specific complexity axes while enforcing feasibility constraints such as preserving task semantics, realism, and execution validity\. Across FinQA, PubMedQA, and ContractNLI and three Qwen3\.5 model scales \(35B\-A3B, 122B\-A10B, and 397B\-A17B\),HARDENreduces task\-model accuracy by 22\.7% on average and by up to 49\.9% relative to single\-pass baselines using the same feasibility checks\. These results show that evolutionary search can produce substantially harder valid evaluation cases\.
## 1Introduction
Language models \(LMs\) are increasingly evaluated on curated or synthetic benchmarks[Chen et al\. \(2021\)](https://arxiv.org/html/2609.30571#bib.bib14);[Guha et al\. \(2023\)](https://arxiv.org/html/2609.30571#bib.bib17);[Jin et al\. \(2019\)](https://arxiv.org/html/2609.30571#bib.bib15);[Koreeda and Manning \(2021\)](https://arxiv.org/html/2609.30571#bib.bib16);[Li et al\. \(2024\)](https://arxiv.org/html/2609.30571#bib.bib1)\. While these benchmarks provide controlled and reproducible evaluations, their inputs often abstract away the ambiguity, grounding requirements, and contextual complexity of real\-world domains\. As a result, clean benchmarks provide a weak signal of reliability, and high effectiveness on them may not carry over to deployments\. We tackle the problem of automatically adapting existing evaluation cases into harder variants for real\-world AI systems\.
Making a benchmark more challenging, however, is insufficient\. A task modification can reduce model effectiveness by introducing contradictions or removing necessary evidence\. For example, changing a referenced fiscal period in a financial reasoning case may lower model effectiveness while also changing the correct answer, making the original evaluation target invalid\. As suggested by Goodhart’s law, optimizing a measure can encourage solutions that exploit the measure rather than pursue the intended objective\. Likewise, LM\-based case generators optimized only for model failure may increase difficulty by violating task semantics rather than exposing meaningful capability weaknesses\. We therefore ask:*how can we increase task difficulty \(measured by both model performance and model uncertainty\) while preserving the evaluation target?*
We formulate this problem as a constrained search for harder evaluation cases\. For each search, we start from an existing case whose task and expected output provide a fixed reference\. We modify only its input while keeping the expected output fixed\.
We introduceHARDEN, an evolutionary search algorithm that modifies existing benchmarks to enhance task complexity and model uncertainty\.HARDENiteratively proposes and selects increasingly challenging cases while enforcing three constraints: \(i\)*correctness*, requiring that each variant preserves the evidence needed to derive the expected output; \(ii\)*realism*, requiring that it remain plausible in the deployment domain; and \(iii\)*validity*, requiring that it remain well\-formed and executable by the original benchmark\.
Our contributions are as follows:
- •HARDEN, a constrained evolutionary search framework that adapts existing evaluation cases along domain\-specific complexity axes\.
- •LM\-based feasibility checks for correctness and realism that discard variants that fail either check, evaluated against human judgments\.
- •An evaluation across professional domains and model sizes shows that, relative to single\-pass baselines using the same feasibility checks,HARDENproduces valid cases that*reduce task\-model accuracy by 22\.7% on average and up to 49\.9%*, while increasing normalized output semantic entropy by 112% on average\.
## 2Related Work
Prior work generates harder evaluation cases while attempting to preserve their original semantics\. Answer\-preserving adversarial testing adds distractors to SQuAD without changing the correct answer\([Rajpurkar et al\., 2016](https://arxiv.org/html/2609.30571#bib.bib4);[Jia and Liang, 2017](https://arxiv.org/html/2609.30571#bib.bib2)\), while SEARs derive semantically equivalent replacement rules across multiple NLP tasks\([Ribeiro et al\., 2018](https://arxiv.org/html/2609.30571#bib.bib3)\)\. Population\-based search has similarly been used to find semantically and syntactically similar adversarial text, with TextAttack later formalizing attacks through transformations, constraints, objectives, and search\([Alzantot et al\.,](https://arxiv.org/html/2609.30571#bib.bib6);[Morris et al\., 2020](https://arxiv.org/html/2609.30571#bib.bib5)\)\.
More recent methods emphasize constrained and adaptive generation: SECA preserves semantic equivalence and coherence during adversarial prompt search\([Liang et al\., 2025](https://arxiv.org/html/2609.30571#bib.bib8)\); CETBench and SQLMorph apply semantics\-preserving transformations to code and Text\-to\-SQL evaluation\([Oza et al\., 2026](https://arxiv.org/html/2609.30571#bib.bib11);[Malekpour et al\., 2026](https://arxiv.org/html/2609.30571#bib.bib7)\); DARG adaptively modifies reasoning graphs while validating labels\([Zhang et al\., 2024](https://arxiv.org/html/2609.30571#bib.bib9)\); and AdvPrompter generates human\-readable adversarial suffixes that preserve instruction meaning\([Paulus et al\., 2025](https://arxiv.org/html/2609.30571#bib.bib10)\)\.
Unlike prior methods,HARDENtreats the correctness, realism, and validity constraints as separate without relying only on semantic similarity or predefined transformations\.
## 3Our Approach:HARDEN
We propose a constrained evolutionary method that modifies inputs while holding their evaluation targets fixed\. Let\(x,y\)\(x,y\)denote an original evaluation case as an input\-output pair\. TheHARDENsearch seeks to find a candidate inputx′x^\{\\prime\}that minimizes the fitness functions\(x′\)s\(x^\{\\prime\}\), defined as follows:
x⋆∈argminx′∈𝒳\(x\)\\displaystyle x^\{\\star\}\\;\\in\\;\\operatorname\*\{arg\\,min\}\_\{x^\{\\prime\}\\in\\mathcal\{X\}\(x\)\}s\(x′\)=1T∑t=1TGrade\(y^t,y\)\\displaystyle s\(x^\{\\prime\}\)\\;=\\;\\frac\{1\}\{T\}\\sum\_\{t=1\}^\{T\}\\operatorname\{Grade\}\\\!\\left\(\\hat\{y\}\_\{t\},\\,y\\right\)\(1\)s\.t\.\\displaystyle\\text\{s\.t\.\}C\(x,x′\)=1,R\(x′\)=1,V\(x′\)=1,s\(x′\)\>0\.\\displaystyle C\(x,x^\{\\prime\}\)=1,\\quad R\(x^\{\\prime\}\)=1,\\quad V\(x^\{\\prime\}\)=1,\\quad s\(x^\{\\prime\}\)\>0\.Here,𝒳\(x\)\\mathcal\{X\}\(x\)denotes the search space of candidate inputs derived fromxxthrough possible transformations𝒳\\mathcal\{X\}\. The fitness scores\(x′\)s\(x^\{\\prime\}\)estimates the model’s performance using the average benchmark score overTTsampled responsesy^t\\hat\{y\}\_\{t\}, whereGrade\\operatorname\{Grade\}is the benchmark evaluation function\. The search is subject to three binary feasibility checks: \(i\)C\(x,x′\)C\(x,x^\{\\prime\}\),*correctness*:x′x^\{\\prime\}preserves the evidence needed to derive the expected outputyy; \(ii\)R\(x′\)R\(x^\{\\prime\}\),*realism*:x′x^\{\\prime\}remains plausible in the deployment domain; \(iii\)V\(x′\)V\(x^\{\\prime\}\),*artifact validity*:x′x^\{\\prime\}remains well\-formed and executable by the original benchmark; and an additional score requirement:s\(x′\)\>0s\(x^\{\\prime\}\)\>0, which excludes cases that collapse the task entirely\.
Search Domain Analysis\.Before the evolutionary search,HARDENsamples the evaluation cases that it will evolve and derives two components used throughout the search: a set of complexity axes and a realism rubric\. First, an agent analyzes the sampled target cases to identify eight to eleven domain\-specific*complexity axes*and assign each an initial priority\. A complexity axis corresponds to a directional strategy that can be applied to the input to make it more difficult for a model to derive the output\. Second, a web\-search agent uses the benchmark details, including its description, schemas, and instructions, together with web sources outside the target benchmark, to construct a rubric for judging whether an input is realistic for the domain\. The rubric is calibrated on a held\-out set before advancement\. Examples of the realism rubric and complexity axes are provided in Appendices[A\.1](https://arxiv.org/html/2609.30571#A1.SS1)and[A\.2](https://arxiv.org/html/2609.30571#A1.SS2), respectively\.
Evolutionary Search\.HARDENperforms an independent search for each evaluation case selected for evolution\. For each case, the original inputxxinitializes the search, which proceeds forGGgenerations, generatingλ\\lambdavariants at each generation\. The search uses two distinct and fixed models throughout the search: a*mutation model*to generate input variants and a*task model*being evaluated on the variants through the fitness function in Eq\.[1](https://arxiv.org/html/2609.30571#S3.E1)\.
At generationg∈\{1,…,G\}g\\in\\\{1,\\ldots,G\\\}, letxgpx\_\{g\}^\{\\mathrm\{p\}\}denote the current parent input, wherex1p=xx\_\{1\}^\{\\mathrm\{p\}\}=x\. To generate thejj\-th variant, wherej∈\{1,…,λ\}j\\in\\\{1,\\ldots,\\lambda\\\},HARDENfirst selects a complexity axisag,ja\_\{g,j\}from the fixed set of complexity axes𝒜\\mathcal\{A\}identified during domain analysis and then uses it to guide the mutation:
xg,j′=Mutate\(xgp;ag,j\),ag,j∈𝒜\.x^\{\\prime\}\_\{g,j\}=\\operatorname\{Mutate\}\\left\(x\_\{g\}^\{\\mathrm\{p\}\};a\_\{g,j\}\\right\),\\qquad a\_\{g,j\}\\in\\mathcal\{A\}\.\(2\)The selected axis specifies the type of complexity to introduce, while the mutation model determines how to realize it\. Appendix[A\.5\.4](https://arxiv.org/html/2609.30571#A1.SS5.SSS4)provides the mutation prompt\.
Axis selection adapts across generations\. Before each generation,HARDENranks the axes using their initial priority from domain analysis, whether previous applications passed the feasibility constraints and lowered fitness, and whether an axis has already been used for the current case\. It favors high\-ranked axes while encouraging different axes across the generated variants\.
Each generated variant is then evaluated against the feasibility checks: correctness, realism, and validity in Eq\.[1](https://arxiv.org/html/2609.30571#S3.E1)\. The correctness and realism checks are LM\-based: each usesKKindependent model calls and passes when a majority accept the variant\. Candidates that fail any check are discarded\. Among the feasible candidates,HARDENretains only a providedμ\\mufrom the lowest fitness scores\. Ties are broken using normalized discrete semantic entropy over sampled outputs\([Farquhar et al\., 2024](https://arxiv.org/html/2609.30571#bib.bib13)\), favoring greater model uncertainty, with any remaining tie favoring a newly generated variant\. The best retained candidate is set asxg\+1px\_\{g\+1\}^\{\\mathrm\{p\}\}\. AfterGGgenerations, the best candidatex⋆x^\{\\star\}is returned only ifs\(x⋆\)<s\(x\)s\(x^\{\\star\}\)<s\(x\)and feasibility checks are satisfied; otherwise, the original input is retained\. Appendix[A\.3](https://arxiv.org/html/2609.30571#A1.SS3)shows an original evaluation case and itsHARDEN\-generated variant\.
## 4Experiments
We evaluate three research questions \(RQs\):
*RQ1\.**Does evolutionary search improve over single\-pass baselines using the same feasibility checks enough to justify its additional cost?*
*RQ2\.**DoesHARDENremain effective as task\-model capability increases?*
*RQ3\.**How doHARDEN’s correctness and realism checks compare with human judgment?*
### 4\.1Setup
Models\.We useQwen3\.5MoEs \(35B\-A3B, 122B\-A10B, and 397B\-A17B\) as task models served through OpenRouter with reasoning disabled\. We useClaudeOpus 4\.8for domain analysis and mutation at high reasoning effort\.
Benchmarks\.We use FinQA\([Chen et al\., 2021](https://arxiv.org/html/2609.30571#bib.bib14)\), PubMedQA\([Jin et al\., 2019](https://arxiv.org/html/2609.30571#bib.bib15)\), and ContractNLI\([Koreeda and Manning, 2021](https://arxiv.org/html/2609.30571#bib.bib16)\)\(part of the LegalBench suite\([Guha et al\., 2023](https://arxiv.org/html/2609.30571#bib.bib17)\)\) to represent three different professional domains where deployed solutions must handle messy inputs\. We selected these as near\-saturation tasks for the three task\-model scales\([Phogat et al\., 2023](https://arxiv.org/html/2609.30571#bib.bib18);[Singhal et al\., 2023](https://arxiv.org/html/2609.30571#bib.bib20);[Nori et al\., 2023](https://arxiv.org/html/2609.30571#bib.bib21);[Schuster et al\., 2022](https://arxiv.org/html/2609.30571#bib.bib19)\)\([A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4)\)\. We sample 200 examples as a representative set for each benchmark\.
Baselines\.We compare against two baselines representing LM\-based alternatives\.*Few\-Shot*makes one tool\-free mutation per evaluation case\.*Few\-Shot \(web\-search\)*makes the same mutation with read\-only web access, while blocking searches that identify the target benchmark\. Each benchmark case is paired with five representative challenging evaluation cases sampled from low\-scoring examples from a disjoint bank, identified as real, difficult few\-shot examples\. Appendix[A\.8](https://arxiv.org/html/2609.30571#A1.SS8)describes how these examples are selected\. The baselines use the same models and feasibility checks asHARDEN\.
HARDENparameters\.We runHARDENwithG=4G=4,λ=11\\lambda=11,μ=5\\mu=5,K=3K=3, andT=10T=10\.
Metrics\.We use each benchmark’s evaluation metric: execution accuracy for FinQA and exact label match for PubMedQA and ContractNLI\. We refer to both as*accuracy*for brevity\. We utilize*semantic entropy*\([Farquhar et al\., 2024](https://arxiv.org/html/2609.30571#bib.bib13);[Kuhn et al\., 2023](https://arxiv.org/html/2609.30571#bib.bib12)\)to measure a model’s predictive uncertainty\. This metric computes frequency\-based uncertainty over the output\-distribution of sampled task model outputs\. Calculation of additional post\-hoc uncertainty metrics is detailed in Appendix[A\.4](https://arxiv.org/html/2609.30571#A1.SS4)\.
### 4\.2Results
Figure 1:Task model accuracy by mutation method and model size\. Lower accuracy indicates more effective hardening\. Error bars are 95% case\-bootstrap CIs\([Efron and Tibshirani, 1993](https://arxiv.org/html/2609.30571#bib.bib22)\)based on ten responses per case\.Figure 2:Mean\-normalized discrete semantic entropy by mutation method and task model size\. Higher indicates greater task\-model output uncertainty\. Error bars show 95% case\-bootstrap confidence intervals\.Figure 3:Successfully hardened cases \(out of 200\) by mutation method and task\-model size\. A case is counted when it passes both LM\-based checks and reduces sampled task\-model accuracy relative to its original\.Figure[1](https://arxiv.org/html/2609.30571#S4.F1)reports task\-model accuracy over all 200 cases in each of FinQA, PubMedQA, and ContractNLI for the original data and for the datasets produced by both baselines andHARDEN\. Figure[2](https://arxiv.org/html/2609.30571#S4.F2)reports uncertainty \(semantic entropy\) for the same cases\. Figure[3](https://arxiv.org/html/2609.30571#S4.F3)reports the number of successfully hardened cases for each method\. Appendix[A\.10](https://arxiv.org/html/2609.30571#A1.SS10)reports accuracy and uncertainty metrics, including additional post\-hoc uncertainty metrics\.
##### RQ1\.HARDENconsistently outperforms single\-pass baselines\.
Across all benchmark\-model combinations,HARDENreduces accuracy for52\.4%52\.4\\%of the original cases, compared with12\.7%12\.7\\%for Few\-Shot and13\.3%13\.3\\%for Few\-Shot \(web\-search\)\. On average, for FinQA, PubMedQA, and ContractNLI,HARDENreduces the accuracy by16\.1%16\.1\\%,36\.3%36\.3\\%, and15\.6%15\.6\\%, respectively, across all three Qwen task models\. In contrast, Few\-Shot changes the accuracy by−1\.4%\-1\.4\\%\(FinQA\),\+3\.9%\+3\.9\\%\(PubMedQA\), and−0\.3%\-0\.3\\%\(ContractNLI\) across the three domains, while Few\-Shot \(web\-search\) changes it by−2\.4%\-2\.4\\%,\+5\.6%\+5\.6\\%, and\+0\.3%\+0\.3\\%\. The correctness and realism checks discard42\.0%42\.0\\%,3\.5%3\.5\\%, and23\.0%23\.0\\%of Few\-Shot mutations on FinQA, PubMedQA, and ContractNLI, respectively, and27\.0%27\.0\\%,2\.0%2\.0\\%, and22\.5%22\.5\\%of Few\-Shot \(web\-search\) mutations\. Semantic entropy shows the same separation: averaged across model scales,HARDENincreases entropy by 0\.19 on FinQA, 0\.24 on PubMedQA, and 0\.13 on ContractNLI, while both single\-pass baselines remain within 0\.02 of the original average in each domain\.HARDENconsistently identifies feasible mutations that yield lower task\-model accuracy and higher output uncertainty, indicating greater predictive difficulty across all three domains\. In comparison, single\-pass mutation, even with web access, rarely achieves both objectives of feasibility \(correctness and realism\) and increased difficulty\.
##### RQ2\.HARDENremains effective as task\-model size increases\.
Across the three benchmarks,HARDENreduces accuracy by an average of21\.5%21\.5\\%,25\.2%25\.2\\%, and21\.2%21\.2\\%for the 35B, 122B, and 397B Qwen task models, respectively\. The reduction remains substantial and consistent as model size increases\. At the largest model size, HARDEN successfully hardens an average of 81\.7 out of 200 cases for a benchmark, nearly 2\.5 times the 33\.0 achieved by the strongest baseline at the smallest model size \(Figure[3](https://arxiv.org/html/2609.30571#S4.F3)\)\. Across the three benchmarks,HARDENalso increases semantic entropy from 0\.19 to 0\.39 at 35B, 0\.17 to 0\.38 at 122B, and 0\.11 to 0\.26 at 397B, averaging 112% above the constraint\-filtered single\-pass baselines \(Figure[2](https://arxiv.org/html/2609.30571#S4.F2)\)\. The baseline approaches produce minimal shifts in task performance or predictive uncertainty\. These results show thatHARDENcontinues to identify and introduce challenging complexity as task\-model size and capability grow\.
##### RQ3\. Agreement of feasibility checks with human judgments\.
We compareHARDEN’s correctness and realism checks with human judgments on the same cases \([A\.11](https://arxiv.org/html/2609.30571#A1.SS11)\)\. For realism, human participants andHARDENindependently evaluate 60 cases per benchmark: 20 original cases and 40 synthetic cases, comprising 20 generated byGPT\-5\.6 Soland 20 byGPT\-3\.5, representing stronger and weaker generators, respectively\. For correctness, they evaluate 20 original–mutated pairs fromHARDENruns for each of FinQA and PubMedQA, balanced between 10 correct and 10 incorrect mutations as measured by the auditor\. In both studies, the human label is two\-of\-three majority vote\.
On FinQA, the realism check accepts90%90\\%,95%95\\%, and0%0\\%of the original,GPT\-5\.6 Sol\-generated, andGPT\-3\.5\-generated cases, versus60%60\\%,70%70\\%, and60%60\\%for participants\. Direct agreement is51\.7%51\.7\\%\. PubMedQA shows directional alignment and85\.0%85\.0\\%agreement, although the check accepts90%90\\%ofGPT\-3\.5\-generated cases\. Thus,HARDENis more selective than participants when filtering weaker\-generator variants on FinQA, and more closely tracks their judgments on PubMedQA\.
For correctness, across the 40 mutations from both benchmarks, participant majorities accept 39, including 19 of the 20 rejected by the check, providing little discrimination\. To distinguish indiscriminate rejection from conservative handling of ambiguous cases, we evaluated the check on a separate FinQA set specifically to test whether it could identify incorrect examples\. We generated counterfactuals by altering critical evidence to invalidate the original output, included correctness\-preserving controls, and retained only mutations with agreed co\-author labels; the check achieved 85\.7% agreement\. Overall, the study shows greater participant agreement for realism on PubMedQA than on FinQA\. The correctness check appears to identify incorrect examples, albeit conservatively, but participant judgments provide limited evidence for validation\.
## 5Limitations
The experiments cover benchmarks representing three professional domains, and use one task model family and one mutation model\. Trends across Qwen3\.5 model scales do not establish transfer across architectures\. All task models are evaluated with reasoning disabled, so whether reasoning changes robustness to the mutated examples remains open\. Cross\-model transfer of difficulty remains to be measured\. The feasibility check validation studies are small and provide mixed evidence\. Participants were role\-screened, but their domain expertise was not additionally verified as a part of the studies\. The initial correctness study provides limited evidence about the check’s ability to distinguish correct from incorrect mutations, and the secondary adjudicated study covers only 21 FinQA pairs\. Additional evidence is required for conclusive proof of validity\. Addressing these limitations requires larger studies, a redesigned participant correctness interface, evaluation of naturally occurring invalid mutations in every domain, and recalibration of check thresholds against those judgments\. Finally, our approach is substantially more expensive than single\-pass mutation, due to evolutionary expansion of mutated case candidates across multiple generations\. We encourage practitioners to gauge whether the benefits provided by the approach are worthwhile prior to leveraging the approach\.
## 6Conclusion
Curated benchmarks used for the evaluation of AI systems often fail to emulate the complexity and ambiguity of real\-world data\. We introduceHARDEN, a constrained evolutionary search method to inject domain\-realistic complexity into benchmarks while preserving their evaluation targets\. Across FinQA, PubMedQA, and ContractNLI,HARDENsignificantly increases task difficulty, as measured by accuracy and uncertainty, relative to single\-pass agentic baselines\. The effect remains substantial as Qwen3\.5 task model capability increases from 35B to 122B to 397B, showing that model\-conditioned evolutionary search retains efficacy as task models strengthen\. These results establishHARDENas an effective method for constructing harder, answer\-preserving evaluations across the studied domains and model scales\.
## Acknowledgments and Disclosure of Funding
We thank our colleagues Benjamin Fowlersmith, Deep Patel, Durga Sandeep Saluru, and Philippe Wyder at Distyl AI for their contributions to the greater system from which the design forHARDENwas inspired\. The Metapod artwork in the title is sourced from PNGKey\([PNGKey,](https://arxiv.org/html/2609.30571#bib.bib23)\)under the site’s stated noncommercial\-use terms\. The original artist is not identified on the source page\.
## References
- \[1\]M\. Alzantot, Y\. Sharma, A\. Elgohary, B\. Ho, M\. Srivastava, and K\. ChangGenerating natural language adversarial examples\.InEMNLP,pp\. 2890–2896\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p1.1)\.
- \[2\]Z\. Chen, W\. Chen, C\. Smiley, S\. Shah, I\. Borova, D\. Langdon, R\. Moussa, M\. Beane, T\. Huang, B\. Routledge, and W\. Y\. Wang\(2021\)FinQA: a dataset of numerical reasoning over financial data\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 3697–3711\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.300)Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px1.p1.1),[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30571#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[3\]B\. Efron and R\. J\. Tibshirani\(1993\)An introduction to the bootstrap\.Chapman & Hall/CRC\.Cited by:[§A\.10](https://arxiv.org/html/2609.30571#A1.SS10.p1.1),[Figure 1](https://arxiv.org/html/2609.30571#S4.F1)\.
- \[4\]S\. Farquhar, J\. Kossen, L\. Kuhn, and Y\. Gal\(2024\)Detecting hallucinations in large language models using semantic entropy\.Nature630\(8017\),pp\. 625–630\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-07421-0)Cited by:[§A\.4](https://arxiv.org/html/2609.30571#A1.SS4.SSS0.Px1.p2.1),[§3](https://arxiv.org/html/2609.30571#S3.p6.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p5.1)\.
- \[5\]N\. Guha, J\. Nyarko, D\. E\. Ho, C\. Ré, A\. Chilton, A\. Narayana, A\. Chohlas\-Wood, A\. Peters, B\. Waldon, D\. N\. Rockmore,et al\.\(2023\)LegalBench: a collaboratively built benchmark for measuring legal reasoning in large language models\.InAdvances in Neural Information Processing Systems 36, Datasets and Benchmarks Track,Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30571#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[6\]R\. Jia and P\. Liang\(2017\)Adversarial examples for evaluating reading comprehension systems\.InEMNLP,pp\. 2021–2031\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p1.1)\.
- \[7\]Q\. Jin, B\. Dhingra, Z\. Liu, W\. Cohen, and X\. Lu\(2019\)PubMedQA: a dataset for biomedical research question answering\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing,pp\. 2567–2577\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1259)Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px2.p1.1),[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2609.30571#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[8\]Y\. Koreeda and C\. Manning\(2021\)ContractNLI: a dataset for document\-level natural language inference for contracts\.InFindings of the Association for Computational Linguistics: EMNLP 2021,pp\. 1907–1919\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.findings-emnlp.164)Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2609.30571#S1.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[9\]L\. Kuhn, Y\. Gal, and S\. Farquhar\(2023\)Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VD-AYtP0dve)Cited by:[§A\.4](https://arxiv.org/html/2609.30571#A1.SS4.SSS0.Px1.p2.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p5.1)\.
- \[10\]J\. Li, B\. Hui, G\. Qu, J\. Yang, B\. Li, B\. Li, B\. Wang, B\. Qin, R\. Geng, N\. Huo,et al\.\(2024\)Can llm already serve as a database interface? a big bench for large\-scale database grounded text\-to\-sqls\.Advances in Neural Information Processing Systems36\.Cited by:[§1](https://arxiv.org/html/2609.30571#S1.p1.1)\.
- \[11\]B\. Liang, L\. Peng, J\. Luo, D\. Thaker, K\. H\. R\. Chan, and R\. Vidal\(2025\)SECA: semantically equivalent and coherent attacks for eliciting llm hallucinations\.InNeurIps,pp\. 142059–142099\.External Links:[Document](https://dx.doi.org/10.52202/085713-4753)Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p2.1)\.
- \[12\]M\. Malekpour, M\. Riahi, M\. Lamothe, and A\. Mhedhbi\(2026\)SQLMorph: query mutation and fine\-grained metrics for text\-to\-sql evaluation\.InICDE,pp\. 2628–2640\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p2.1)\.
- \[13\]J\. Morris, E\. Lifland, J\. Y\. Yoo, J\. Grigsby, D\. Jin, and Y\. Qi\(2020\)TextAttack: a framework for adversarial attacks, data augmentation, and adversarial training in NLP\.InEMNLP: System Demonstrations,pp\. 119–126\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p1.1)\.
- \[14\]H\. Nori, Y\. T\. Lee, S\. Zhang, D\. Carignan, R\. Edgar, N\. Fusi, N\. King, J\. Larson, Y\. Li, W\. Liu,et al\.\(2023\)Can generalist foundation models outcompete special\-purpose tuning? case study in medicine\.arXiv preprint arXiv:2311\.16452\.Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[15\]N\. Oza, I\. Govil, P\. Gupta, D\. Khandelwal, D\. Garg, and P\. Singla\(2026\)LLMs are brittle to simple code transformations: introducing CETBench – a benchmark for code\-equivalence checking\.InFindings of the Association for Computational Linguistics: ACL,pp\. 41653–41685\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p2.1)\.
- \[16\]A\. Paulus, A\. Zharmagambetov, C\. Guo, B\. Amos, and Y\. Tian\(2025\)AdvPrompter: fast adaptive adversarial prompting for LLMs\.InICML,pp\. 48439–48469\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p2.1)\.
- \[17\]K\. S\. Phogat, C\. Harsha, S\. Dasaratha, S\. Ramakrishna, and S\. A\. Puranam\(2023\)Zero\-shot question answering over financial documents using large language models\.arXiv preprint arXiv:2311\.14722\.Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px1.p1.1),[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[18\]PNGKey011metapod dream – pokemon metapod\.Note:[https://www\.pngkey\.com/maxpic/u2q8u2o0e6t4t4o0/](https://www.pngkey.com/maxpic/u2q8u2o0e6t4t4o0/)Accessed September 18, 2026Cited by:[Acknowledgments and Disclosure of Funding](https://arxiv.org/html/2609.30571#Sx1.p1.1)\.
- \[19\]P\. Rajpurkar, J\. Zhang, K\. Lopyrev, and P\. Liang\(2016\)SQuAD: 100,000\+ questions for machine comprehension of text\.InEMNLP,pp\. 2383–2392\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p1.1)\.
- \[20\]M\. T\. Ribeiro, S\. Singh, and C\. Guestrin\(2018\)Semantically equivalent adversarial rules for debugging NLP models\.InACL,pp\. 856–865\.Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p1.1)\.
- \[21\]T\. Schuster, S\. Chen, S\. Buthpitiya, A\. Fabrikant, and D\. Metzler\(2022\)Stretching sentence\-pair NLI models to reason over long documents and clusters\.InFindings of the Association for Computational Linguistics: EMNLP 2022,Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px3.p1.1),[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[22\]K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.\(2023\)Large language models encode clinical knowledge\.Nature620,pp\. 172–180\.Cited by:[§A\.7](https://arxiv.org/html/2609.30571#A1.SS7.SSS0.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.30571#S4.SS1.p2.1)\.
- \[23\]Z\. Zhang, J\. Chen, and D\. Yang\(2024\)DARG: dynamic evaluation of large language models via adaptive reasoning graph\.InNeurIps,pp\. 135904–135942\.External Links:[Document](https://dx.doi.org/10.52202/079017-4317)Cited by:[§2](https://arxiv.org/html/2609.30571#S2.p2.1)\.
## Appendix AHARDEN artifacts and check validation
This appendix records concrete artifacts from our experiments\. We include them to make the information boundaries, calibration decisions, and check behavior auditable rather than describing them only at the level of prompts\.
### A\.1Example target\-blind realism rubric
Table[1](https://arxiv.org/html/2609.30571#A1.T1)summarizes an actual rubric created by HARDEN’s realism check for FinQA\. The domain\-specific rubric was first researched, finding TAT\-QA and FinanceBench as external benchmark families and SEC EDGAR filings as a primary real\-world source family, after which it calibrated against real FinQA benchmark cases\. The table preserves the rubric’s operational thresholds and high\-score anchors while shortening the descriptions for presentation\.
Table 1:The target\-blind “Financial\-report numerical\-reasoning input realism rubric” used in our FinQA experiment\. Scores use a 1–5 scale\.A case passes only when its mean score is at least 4 out of 5, every dimension is at least 4 out of 5, and no blocking red flag is detected\. The two blocking flags are \(R1\) an absent or degenerate financial table and \(R2\) numeric cells with no authentic accounting convention\. The rubric also records two non\-blocking signals that lower the relevant dimension scores: generic placeholder vocabulary \(R3\) and a question ungrounded in the supplied report \(R4\)\.
Applied diagnostically to the original and synthetic \(20 each generated by GPT\-5\.6 Sol and GPT\-3\.5 Turbo respectively\) FinQA cases used in the check\-validation study, the frozen rubric’s per\-source acceptance rates are those reported in Table[10](https://arxiv.org/html/2609.30571#A1.T10)\(Appendix[A\.11](https://arxiv.org/html/2609.30571#A1.SS11)\); 19 of the 20 rejected GPT\-3\.5 Turbo cases triggered the degenerate\-table flag\.
### A\.2Example complexity\-axis catalog
Table[2](https://arxiv.org/html/2609.30571#A1.T2)shows the complete set of ten axes produced by the domain\-research stage for FinQA\. Each axis specifies a family of realistic, target\-preserving mutations rather than one fixed transformation\.
Table 2:Generated FinQA complexity\-axis catalog, with abbreviated examples\.
### A\.3Example original andHARDEN\-generated case
Table[3](https://arxiv.org/html/2609.30571#A1.T3)excerpts a ContractNLI case hardened byHARDENduring the completed full run under the legal\-boilerplate and recital\-inflation axis\. The evaluation target is unchanged: the hypothesis “Receiving Party shall destroy or return some Confidential Information upon the termination of Agreement\.” retains its gold labelNotMentioned, and the operative clauses of the agreement are preserved verbatim\. The added material is ceremonial and non\-determinative, yet the mean Qwen 3\.5 397B score over ten trials falls from 1\.0 to 0\.3\.
Table 3:Excerpt of a ContractNLI case before and after evolution\. The mutation prepends authentic recitals and pads the execution formalities of a real mutual NDA; ellipses mark elided unchanged text\.
### A\.4Search and task model details
Rubric research receives the use\-case description, schemas, execution plan, and system prompt, but no target cases, labels, mutations, evaluator results, or task model scores\. It draws on at least two independent external benchmark families and one primary real\-world source family, calibrates 3–5 realism dimensions on held\-out examples, and freezes the rubric by SHA\-256 hash\. Complexity axis discovery is separate and may inspect target inputs, immutable outputs, and the evaluator to identify label\-preserving difficulty mechanisms\. During evolution, HARDEN records mutation axes applications that pass both LM\-based checks and lower fitness; it rewards prior improvement and penalizes repeated non\-improvement\.
Each benchmark is evaluated with the Qwen 3\.5 35B, 122B, and 397B task models served through OpenRouter with reasoning disabled\. Ten sampled completions use pinned model\-card profiles in[A\.5\.1](https://arxiv.org/html/2609.30571#A1.SS5.SSS1); these samples are used for primary performance and uncertainty\. A temperature\-zero completion and its token log\-probabilities are retained for post\-hoc sequence NLL and mean\-token NLL\. OpenRouter returns output\-token log\-probabilities for all generations, which the post\-hoc likelihood\-based uncertainty measures require\.
In a preliminary evaluation using 15 official cases per benchmark and two independent calls per setting, enabling reasoning produced no consistent accuracy benefit\. On FinQA with the 35B model, accuracy decreased by 3\.3%, from 76\.7% to 73\.3%\. On PubMedQA with the 122B model, accuracy increased by 3\.3%, from 46\.7% to 50\.0%\. These mixed effects came with substantially greater token usage: approximately 3\.7×\\timeson FinQA \(864 to 3\.2k tokens\) and 8\.3×\\timeson PubMedQA \(228 to 1\.9k tokens\), as well as increased latency\. Further, the output token log\-probabilities referenced to calculate uncertainty as a secondary measure of task difficulty are incompatible with reasoning models\. This is due to the fact that the output token lob\-probabilities would be conditioned on input tokens as well as reasoning tokens, which differ across samples; this confounds the task uncertainty we aim to measure\.
##### Reproducibility materials\.
Appendix[A\.5](https://arxiv.org/html/2609.30571#A1.SS5)discloses the fixed method prompts, dynamic\-field schemas, search configuration, model settings, and tool boundaries used by the experiments\. No code, data, or exact run outputs are released\. The disclosure is intended to support an independent implementation and reproduction of the experimental procedure\.
Our post\-hoc likelihood\-weighted semantic entropy follows Kuhn et al\.\[[9](https://arxiv.org/html/2609.30571#bib.bib12)\]: samples are grouped into exact task\-level semantic classes, weighted by length\-normalized sequence likelihood, and aggregated before computing entropy\. Online selection instead uses a normalized adaptation of the discrete semantic entropy introduced by Farquhar et al\.\[[4](https://arxiv.org/html/2609.30571#bib.bib13)\], estimating class probabilities from sample frequencies\. For PubMedQA and ContractNLI, the exact semantic classes are the three answer labels,ui\(xi\)=−∑zp^zlogp^z/log3u\_\{i\}\(x\_\{i\}\)=\-\\sum\_\{z\}\\hat\{p\}\_\{z\}\\log\\hat\{p\}\_\{z\}/\\log 3\. For FinQA,zzis the distinct executed numerical result and the normalization islogT\\log T, whereT=10T=10\. Post\-hoc analysis also reports sequence NLL, mean\-token NLL, frequency entropy, and likelihood\-weighted entropy; these complementary measures characterize confidence without changing the primary task\-score objective\.
### A\.5Reproducibility prompt and configuration disclosure
This appendix discloses the method\-critical templates and dynamic\-field schemas used in the experiments\. Line wrapping and surrounding configuration syntax are normalized forLaTeX\. Internal terminology is rendered as “task model” for consistency\. The verbatim production templates retain the implementation term “gate”; it denotes the correctness or realism check described in the paper, not a separate mechanism\.
#### A\.5\.1Search, model, and tool configuration
The paper\-facing search configuration samples 200 cases without replacement with seed 42, preserves source order after sampling, and partitions the cases into ten deterministic worker runs\. The same mutation axis catalog is shared across worker runs\.
Rubric research, axis research, and mutation identify Claude Opus 4\.8 \(anthropic/claude\-opus\-4\-8\) with adaptive thinking and high effort\. No temperature or output\-token override is applied to those agent\-driven calls, so they use the agent runtime defaults\. Correctness \(both extraction and audit\) and realism \(audit\) calls use the same model through direct OpenRouter checks with temperature 0, at most 16,384 output tokens, JSON\-object responses, and high reasoning effort\. They fail closed unless the response metadata confirms that route\. The three\-member ensembles use majority voting, which requires two of three passes\.
All task\-model profiles request one greedy completion at temperature 0 and ten sampled completions\. Reasoning is disabled, repetition penalty is 1\.0, maximum output tokens are 16,384, and selected\-token log probabilities request the top five alternatives\. OpenRouter parameter enforcement is required\. The exact sampled profiles are:
- •qwen/qwen3\.5\-35b\-a3b: Parasail, temperature 1\.0, top\-pp0\.95, top\-kk20, and presence penalty 1\.5\.
- •qwen/qwen3\.5\-122b\-a10b: Novita, temperature 1\.0, top\-pp1\.0, top\-kk40, and presence penalty 2\.0\.
- •qwen/qwen3\.5\-397b\-a17b: Parasail, temperature 0\.7, top\-pp0\.8, top\-kk20, and presence penalty 1\.5\.
Rubric research may read, search, fetch, and use shell commands, while the named target benchmark family and its mirrors and derivatives are forbidden\. Axis research may use the web\. Before inspecting the benchmark cases, the axis researcher reads two fixed support files: a research specification requiring evidence\-backed, real\-world difficulty that preserves labels and avoids rubric\-targeting, and a catalog of generic transformation ideas used only as seeds for domain\-specific axis discovery\. Their paths are included in the single axis\-research request; they do not create additional model calls\. Mutation generation has no web, read, search, or shell access and works only from the constructed context\. Direct checks and task\-model calls use no tools\. Both single\-pass baselines use Claude Opus 4\.8 as the mutation model with adaptive thinking, high reasoning effort, default temperature, one user turn, and a maximum of 16,384 output tokens\.
##### Deterministic selection and artifacts\.
Parents and check\-approved children are sorted by ascending task\-model score\. Exact score ties prefer higher normalized sampled\-outcome uncertainty and then a new child\. The five best survive\. A score of zero is treated as collapsed and rejected\. Final promotion requires a strict score improvement over the immutable baseline\. Before scoring, deterministic checks require a valid mutation index and manifest, explicit passes from both correction and realism checks, edits confined to declared mutable inputs and unchanged immutable files\. Finalization reconstructs each selected case from its immutable snapshot, copies only declared mutable artifacts, parses selected JSON files, and runs the structural completion check\. These are deterministic code contracts, not model prompts\.
#### A\.5\.2Rubric research templates
Prompt A\.1 is the system instruction for the target\-blind rubric researcher\. Prompt A\.2 is the user request sent to that researcher once before case\-level evolution begins\. They are the two messages in one rubric\-research call\.
Youareadomain\-realismresearchagent\.Buildanevidence\-backedrealismrubricthatcanjudgewhetherbenchmarkinputsresembleauthenticrecordsfromtherelevantreal\-worlddomain\.
Theparentworkflowsuppliesabundlerootandusecase\.YouMUSTread\`grounding/domain\_realism/research\_constraints\.json\`\.Everydegradationrubricrunistarget\-blind:inferthedomainandtaskfamilyONLYfrommetadata,schemas,prompts,andthegenericusecase\.Theresearchbundleintentionallycontainsnotargetcases\.Donotsearchfor,retrieve,read,cite,normalize,orcalibrateonexamplesfromanyfamilynamedin\`excluded\_benchmark\_families\`,includingmirrorsandderivativedatasets\.
Neveruseexpectedoutputs,labels,goldanswers,evaluatorresults,mutationartifacts,ortask\-modelscoresasrealismevidence\.
Thispromptisintentionallydomain\-agnostic\.Youmustdiscovertherelevantreferencesourcesyourself\.Donotassumethatthetargetbenchmark'sformattingconventionsdefinedomainrealism\.
Researchrequirements:
1\.DiscoverandretrievesamplesfromatleasttwoEXTERNAL,INDEPENDENTbenchmarkfamiliesrelevanttotheinferreddomainandtask\.Thetargetbenchmarkdoesnotcounttowardthisminimum\.Datasetsderivedfromthesameunderlyingrecordscountasonefamily\.
2\.RetrievesamplesfromatleastonePRIMARYreal\-worldsourceforthedomain\.Abenchmark,blog,paperdescription,orsyntheticcorpusisnotprimaryreal\-worlddata\.
3\.Verifysourcerelevance,lineage,canonicalURL,revision/version,andlicenseorusageprovenance\.Rejectweaklyrelatedsources\.
4\.Saveimmutablesnapshotsofeverysampleused\.NormalizeINPUTartifactsintoacommonreadablerepresentationwhileretainingtherawsnapshots\.
5\.Excludelabels,expectedanswers,goldprograms,evaluatorrubrics,andmodeloutputsfromnormalizedresearchsamples\.
6\.Derivecommondomainpropertiesacrosssources\.Keepsource\-specificserializationquirksseparate\.Ablockingcriterionneedsevidencefromatleasttwoindependentsourcefamilies,oronebenchmarkfamilyplusprimaryreal\-worlddata\.
7\.Build3\-5authenticitydimensionswitha1\-5scale,concretehigh/lowanchors,ametthreshold,syntheticredflags,andanexplicitcase\-levelpassrule\.Donotscoretaskcorrectness\.
8\.Calibratetherubricondeterministic,disjointprobeandheld\-outsamplesfromeverysourcefamily\.Realheld\-outrecordsshouldmeettherubric;donotweakenitmerelytoforcecalibration\.
9\.Inatarget\-blindrun,everyregistrysourcemustset\`is\_target\_family:false\`;excludedtargetexamplesmustnotappearanywherein\`sources/\`,theregistry,referenceprofile,calibration,orreport\.
Writeeverythingbelow\`$\{BUNDLE\_ROOT\}/grounding/domain\_realism/\`:
\-\`source\_registry\.json\`
\-\`rubric\.json\`
\-\`reference\_profile\.json\`
\-\`research\_report\.md\`
\-\`sources/\`containingrawandnormalizedsamplesnapshots
\`source\_registry\.json\`mustbe:
\{
"inferred\_domain":"concisedomain",
"inferred\_task\_family":"concisetaskfamily",
"target\_benchmark\_family":"familynameorunknown",
"sources":\[
\{
"id":"stable\_source\_id",
"name":"sourcename",
"family":"independentlineagefamily",
"kind":"benchmarkorprimary\_real\_world",
"is\_target\_family":false,
"canonical\_url":"https://\.\.\.",
"revision":"commit,version,date,orimmutableidentifier",
"license\_or\_provenance":"licenseoraccessprovenance",
"retrieved\_at":"ISO\-8601timestamp",
"relevance":"whythissourcematchestheinferreddomainandtask",
"lineage":"underlyingrecordlineageandknownderivatives",
"raw\_sample\_paths":\["grounding/domain\_realism/sources/\.\.\."\],
"normalized\_sample\_path":"grounding/domain\_realism/sources/\.\.\.",
"sample\_count":10,
"sha256":"sha256ofthenormalizedsamplefile"
\}
\]
\}
\`rubric\.json\`mustbe:
\{
"rubric\_name":"domain\-specificname",
"domain\_summary":"holisticsynthesissharedacrosssources",
"scale\_min":1,
"scale\_max":5,
"met\_threshold":4,
"dimensions":\[
\{
"id":"D1",
"name":"dimensionname",
"description":"whatauthenticitypropertyisjudged",
"score\_high":"concreteanchorfor5",
"score\_low":"concreteanchorfor1",
"evidence\_source\_ids":\["source\_a","source\_b"\]
\}
\],
"synthetic\_red\_flags":\[
\{
"id":"R1",
"description":"blockingorstrongsyntheticsignal",
"blocking":true,
"evidence\_source\_ids":\["source\_a","source\_b"\]
\}
\],
"case\_pass\_rule":\{
"minimum\_mean\_score":4\.0,
"all\_dimensions\_must\_meet\_threshold":true,
"blocking\_red\_flags\_fail":true
\},
"overall\_guidance":"howtoapplythecommonrubricwithoutoverfittingonesource",
"calibration":\{
"probe\_paths":\["grounding/domain\_realism/sources/\.\.\."\],
"heldout\_paths":\["grounding/domain\_realism/sources/\.\.\."\],
"heldout\_results\_by\_source":\{
"source\_id":\{"n":5,"mean\_score":4\.2,"met\_rate":0\.8\}
\},
"target\_met":true
\}
\}
\`reference\_profile\.json\`mustidentifycommonpatterns,source\-specific
patterns,deterministicmeasurements,andthenormalizedexamplesthelater
independentauditormayuse\.EveryclaimmustcitesourceIDs\.
Useabsolutepathsforfileoperations\.Finishonlyafterallrequiredfiles
existandparse\.Replywithaconcisesource/rubricsummary;theworkflowreads
thefilesdirectly\.
The uppercase names in the request are dynamic fields\.
BUNDLE\_ROOT:$\{bundle\_dir\}
USE\_CASE:$\{use\_case\}
Deep\-researchadomain\-realismrubricusingthegenericrequirementsinyour
agentinstructions\.Inferthedomainandtaskfromthisbundle\.Discoverall
referencesourcesyourself;nosourcenamesordomain\-specificrubric
dimensionsarebeingsuppliedbytheworkflow\.Useabsolutepathsrootedat
BUNDLE\_ROOTandwriteeveryrequiredartifactunder
BUNDLE\_ROOT/grounding/domain\_realism/\.
The dynamic fields arebundle\_diranduse\_case\. Deterministic validation runs before the rubric is frozen\. If validation fails, its errors are returned to the researcher for correction before strict validation is repeated; that conditional repair instruction is not a separate experimental prompt\.
#### A\.5\.3Axis research templates
The following system instruction and user request form the axis\-research call that produces one catalog shared by all worker runs\.
YouarethedegradationresearcherforacompletedButtonstep01benchmarkbundle\.
Yourjobistoresearchhowtoaddrealisticnoiseandobscuritytothisuse
case'sinputs,thendesigndegradationaxes\.Everycaseinthebundleisin
scopefordegradation;donotclassify,prioritize,orexcludecasesas
non\-degradable\.Wheninvokedbytheshared\-axisresearchworkflow,thebundle
isthefullcorpusandyouraxesbecometheimmutablepoolforallexecution
lanes\.Donotmodifycases,executor\.py,evaluator\.py,prompts/,resources/,or
goldoutputs\.
DefinitionsforcurrentButtonbundles:
\-\`input\_data\`:\`cases/case\_N/input\.json\`plusfileslistedin
\`cases/case\_N/metadata\.json\`under\`input\`\.Thesearetheonlydegradation
targets\.
\-\`gold\_output\`:\`cases/case\_N/output\.json\`plusfileslistedinmetadata
under\`expected\_output\`\.Theseareimmutable\.
\-\`shared\_resources\`:\`resources/\`and\`prompts/\`\.Theseareimmutable\.
\-Degradabletextartifacts:\`\.json\`,\`\.csv\`,\`\.txt\`,\`\.html\`,\`\.md\`,\`\.xml\`,\`\.yaml\`,\`\.yml\`\.
\-Immutableartifacts:images,audio,video,PDFswithoutaneditabletextsidecar,archives,binaries\.
Hard\-but\-realfilterforeveryaxis:
\-Real\-worldmechanism:whatordinaryprocessproducesthiscomplexity?
\-Plausibility:wouldthisappearinrawrealinputsforthisdomain?
\-Rubricindependence:woulditstillbehardwithoutknowingtheevaluatorrubric?
\-Organicovermanufactured:prefermessyworkflows,institutionalvariation,
corrections,ambiguity,density,andimplicitcontext;rejectgotchas,
plantedcontradictions,andremovedfacts\.
Workflow:
1\.Beforereadingthebundle,readbothrequiredpathssuppliedbytheparentworkflow\.Thisismandatory:
\-\`RESEARCH\_DOSSIER\_PROMPT\_PATH\`:thehard\-caseresearchdossiermethod\.
Applyitsreal\-world\-mechanism,rubric\-independence,label\-validity,
evidence\-class,andanti\-input\-hackingrequirementstothisresearch\.
\-\`GENERIC\_AXIS\_SEED\_PATH\`:generictransformationideas\.Treatthemonly
asaseedvocabulary;donotcopytheirnames,definitions,orprompt
guidanceasfinalaxes\.Derivedomain\-specificaxesfromthebundleand
researchedevidence,andrejectanyseedthatlacksaplausibledomain
mechanism\.
\-Ifeitherpathisunsetorunreadable,stopandreportthemissingpath
ratherthanproceedingwithoutit\.Citebothpathsin
\`degradation/degradation\_research\_report\.md\`,includingwhichdossier
constraintsandseedconceptsinformedorwererejectedfromthefinal
axes\.
2\.Read\`bundle\_manifest\.json\`,\`grounding/authoring\_digest\.json\`,
\`grounding/realism\_brief\.json\`,\`grounding/realism\_criteria\.json\`,
\`executor\.py\`,\`evaluator\.py\`,\`prompts/\*\.j2\`,\`benchmark\_results\.json\`,and
all\`cases/case\_\*/metadata\.json\`,\`input\.json\`,and\`output\.json\`files\.Treat
everyinput/goldpairinthisresearchbundleastheevidencebasefor
selectingaxes:prioritizemechanismsthatfitoneormoreexampleswhile
preservingeveryimmutablegoldoutput,andrejectaxesthatareonlygeneric
domainideaswithnoplausibleapplicationtothecorpus\.
3\.Readanycaseinputsidecarfileslistedbymetadatathataretextartifacts\.
4\.Foreverycase,recordtheinputartifactstomutate,relativetothecase
directory,suchas\`input\.json\`or\`invoice\.csv\`\.Includeeverycasein
\`degradable\_cases\`;donotcreateresearcher\-levelskippedcases\.
5\.Researchdomain\-specificdifficultypatterns\.Usewebsearchifavailable;
persistusefulfetchedevidenceunder\`degradation/retrieved\_\*\.md\`\.
6\.Writeallofthesefiles:
\-\`degradation/degradation\_axes\.json\`:everyevidence\-backedreusableaxis
discoveredfromthiscorpus;donotimposeanarbitraryaxis\-countcap\.
Exactlyoneaxismustbenamed\`domain\_specific\_realism\`\.Eachaxismust
include\`axis\_name\`,\`definition\`,\`goal\`,\`distinctive\_focus\`,
\`target\_artifact\_types\`,atleast3\`subtypes\`,andatleast5
\`prompt\_guidance\`strings\.Everyaxismustexplicitlysayitoperatesonly
oncaseinputartifactsandmustpreservegoldoutputcorrectness\.Add
optional\`priority\`\(\`high\`,\`normal\`,or\`low\`\)basedonexpected
effectivenessforthecorpus\.Selectaxesthataddplausiblenoise,
ambiguity,obscurity,orretrievalburdentoconcreteexamples;donot
producegenericaxesdisconnectedfromtheinputs\.ForFinCoT\-like
financialtables,includehigh\-prioritytable\-labelcompressionand
table/lookupambiguityaxeswhenevidencesupportsthem;de\-prioritize
broadproseorOCR\-onlychangesunlesstheypreserveameaningful
table\-interpretationchallenge\.
\-\`degradation/degradation\_research\_report\.md\`
\-\`degradation/degradation\_research\_context\.json\`
\-\`degradation/manifest\.json\`
\-\`degradation/degradable\_cases\.json\`\-thisexactJSONshape:
\{
"degradable\_cases":\[
\{
"case\_id":"case\_1",
"baseline\_score":0\.92,
"degradable\_files":\["input\.json","invoice\.csv"\]
\}
\],
"skipped\_cases":\[\],
"summary":"Shortprosesummaryofwhatyoufound\."
\}
\`baseline\_score\`mustbenormalizedto\`\[0,1\]\`\.
Afterwritingeveryfile,replywithabriefplain\-textsummaryofwhatyou
produced\.Theworkflowreads\`degradation/degradable\_cases\.json\`fromdisk\.
The following request is sent once to produce the axis catalog shared by all worker runs\.
Bundleroot:$\{bundle\_dir\}
Usecase:$\{use\_case\}
Thisisthesingleshared\-axisresearchstageforapartitioneddegradation
run\.Thebundlecontainsthecompletecorpusthatalllaterlaneswillprocess\.
Readbothrequiredcontextfilesbeforeexaminingthebundle:
RESEARCH\_DOSSIER\_PROMPT\_PATH:$\{research\_dossier\_prompt\_path\}
GENERIC\_AXIS\_SEED\_PATH:$\{generic\_axis\_seed\_path\}
Curateonereusableglobaldegradation\-axiscatalogfromthefullcorpus\.
Returneveryevidence\-backedaxisthatisplausibleforoneormorecases;do
notimposeanarbitraryaxis\-countcap\.Writeallstandardresearchartifacts
under$\{bundle\_dir\}/degradation/\.Thisstageresearchesonly:donotinvoke
evolutionaryinitialization,mutation,evaluation,orfinalization\.
The dynamic fields arebundle\_dir,use\_case,research\_dossier\_prompt\_path, andgeneric\_axis\_seed\_path\. The dossier and seed are context files, not prompts attributed to the axis researcher\.
#### A\.5\.4Mutation template and context schema
##### Axis ranking\.
For an axisaawithAaA\_\{a\}prior applications,GaG\_\{a\}constraint\-passing applications, andIaI\_\{a\}fitness\-improving applications, the base priority is4Ia/Aa\+Ga/Aa4I\_\{a\}/A\_\{a\}\+G\_\{a\}/A\_\{a\}, or1\.51\.5whenAa=0A\_\{a\}=0\. The implementation adds33for axes assigned high priority during research\. It subtracts44after at least two applications without improvement,22when the constraint\-passing rate is below0\.30\.3after at least two applications, and11when the axis already appears in the retained lineage\. It may also add benchmark\-specific bonuses: in the FinQA runs, it additionally adds55for table\-label compression axes and44for table\-lookup\. Axes are sorted by this score\. The mutation agent receives the ranking and underlying counts, then selects one axis per offspring while preferring distinct, previously unused axes\.
The following system instruction and user request form each mutation\-generation call\.
Youarethemutationgeneratorinanindependentlygatedevolutionary
degradationpipeline\.
Generateandwritecandidatedegradedversionsofonebenchmarkcase\.Youdo
notjudgecorrectness,realism,orfitness\.Separateindependentgatesownall
acceptancedecisions\.
Theworkflowpromptsupplies\`MUTATION\_CONTEXT\_JSON\`withthecurrentcase,
candidatedirectories,writablefiles,gold\-outputsummary,degradationaxes,
evolutionaryfeedback,andamanifestpath\.Useonlythatcontext\.Writeonly
insidethesuppliedcandidatedirectoriesandmanifestpath\.
Generationrequirements:
1\.Identifyeveryload\-bearingfact,value,relationship,andconstraint
neededfortheimmutableexpectedoutput\.
2\.Read\`candidates\_per\_generation\`andproduceexactlythatmanycandidates\.
3\.Useexactlyonesupplieddegradationaxispercandidate\.Preferdistinct,
high\-performingaxesnotalreadyinthemutationhistory\.
4\.Applyplausiblereal\-worldcomplexity,notplantedtraps,contradictions,
ordeletionofnecessaryinformation\.
5\.Writecompletereplacementcontenttoeachassigned\`case\_dir\`,touching
onlyits\`writable\_files\`\.
6\.Donotruntheexecutor/evaluatoranddonotpredictwhetheranygatewillpass\.
Writethemanifestto\`manifest\_path\`:
\{
"mutations":\[
\{
"individual\_index":0,
"axes":\["axis\_name"\],
"subtypes":\["subtype\_name"\],
"description":"Whatchangedanditsplausiblereal\-worldorigin\.",
"written\_files":\["input\.json"\]
\}
\]
\}
Useintegerindexesfromzerothrough\`candidates\_per\_generation\-1\`,exactly
onceeach\.Nevermodifyexpectedoutputs,metadata,prompts,resources,
executor/evaluatorcode,orcanonicalcasefiles\.
Afterwritingallcandidatesandthemanifest,replyonlywithaconcisecountandtheaxesused\.
Each mutation call suppliesMUTATION\_CONTEXT\_JSON: \{context\}followed by:
Useonlythiscontext\.Generateexactly\`candidates\_per\_generation\`mutations,
writecompletereplacementstotheassignedcandidatedirectories,andwrite
themanifestto\`manifest\_path\`\.Donotjudgecorrectness,realism,orfitness\.
The deterministic mutation context supplies the following exhaustive top\-level fields:case\_id,generation,baseline\_score,candidates\_per\_generation,population\_size,max\_generations,degradable\_files,source\_case\_dir,candidates,manifest\_path,case\_state,candidate\_axes,axis\_feedback,realism\_criteria\_summary,metadata,gold\_output\_summary, andfiles\. Eachcandidatesentry containsindividual\_index,case\_dir, andwritable\_files\. When present,case\_statecontainscurrent\_generation,best\_score,best\_individual\_path, andpopulation; each population entry containsscore,path,axes,mutation\_chain, anddescription\. Each candidate axis containsaxis\_name,priority,definition,goal,distinctive\_focus,subtypes, andprompt\_guidance\. Each axis\-feedback entry containsapplied,survived\_gates, andimproved\_fitness\.
#### A\.5\.5Correctness extraction and audit prompts
The correctness check makes one extraction call for the parent case and one audit call for each candidate\. The boxes below show the complete user requests after fixed instructions and dynamic input fields are combined\.
RespondwithasinglevalidJSONobjectonly\.DonotincludeMarkdown,prose,oranalysis\.
<TASK\>
Identifytheminimalsetofinformationnecessarytoproducetheexpectedoutput\.
</TASK\>
<BASELINE\_EXECUTION\_CONTEXT\>
Thebaselineexecutiontaskis:
\{\{baseline\_execution\_prompt\}\}
</BASELINE\_EXECUTION\_CONTEXT\>
<INSTRUCTIONS\>
Givenaninput,supplementarydata,andexpectedoutput,identifytheMINIMALsetofinformation
fromtheINPUTthatisNECESSARYtoreasonaboutandproducetheexpectedoutput\.
IMPORTANT:ExtractinformationONLYfromtheINPUT,notfromsupplementaryoroutput\.
\-Thesupplementarydataprovidescontext\(e\.g\.,policies,guidelines\)fordecision\-making
\-Theexpectedoutputshowswhatdecisionwasmade
\-Yourjob:identifywhichpiecesofINPUTinformationwerecriticalforreachingthatoutput
CRITICALREQUIREMENT\-OutputEvidence:
Beforemarkinganyinformationas"necessary",youMUST:
1\.FindtheEXACTtextintheexpectedoutputthatdependsonthisinformation
2\.Quotethespecificwords/phrasesfromtheoutput
3\.Explainthedirectcausallink:inputinformation\-\>reasoning\-\>outputtext
DONOTusevaguetermslike"implies","suggests","references","indicates","conveys"\.
YouMUSTpointtoexplicittextintheoutputthatwouldbecomewrongiftheinformationchanged\.
CRITICALREQUIREMENT\-RequestedTaskSemantics:
TheINPUT'srequestedtaskisalwaysnecessaryinformation\.Extractitasa
\`requested\_task\_semantics\`itemevenwhentheexpectedoutputdoesnotquote
thetaskwording\.Itsvaluemustcapture:
\-therequestedfield/entityorquantity;
\-theoperation\(forexampleextraction,ratio,difference,percentagechange,
endingvalue,orgrowth\);
\-comparisondirection,baseline,period,scope,andunitswhereapplicable;
\-therequiredanswerform\(forexampleavalue,listofentities,or
key=valueextractionpairs\)\.
Itsdecision\_criteriamuststatethatversion2canparaphrasethetaskbut
mustaskfortheidenticaloutputsemantics\.Aversion2requestthatasksfor
adifferentarithmeticoperation,reversesacomparison,changesthetarget
field/entity,asksforanendingvalueinsteadofareturn,orrequestsan
unsupportedcomparisonMUSTbelistedasinvalid\.
ExampleofVALIDreasoning:
BAD:"patient\_nameisnecessarybecauseit'sreferencedintheoutput"
GOOD:"patient\_name:'JohnSmith'isnecessarybecausetheoutputexplicitlystates:'PatientJohnSmithiseligible\.\.\.'\.Ifthenamechanged,thisexacttextwouldbeincorrect\."
ExampleofwheninformationisNOTnecessary:
\-Outputsays"Patientiseligible"withoutmentioningname\-\>patient\_nameisNOTnecessary
\-Outputsays"Diagnosiscodeindicatesdiabetes"butdoesn'tstatethecode\-\>specificdiagnosis\_codevalueisNOTnecessary\(onlypresenceofsomediagnosis\)
Thinkcritically:
1\.Whatspecificfacts,values,orreferencesarerequired?
2\.Whatinformationwouldmakethetaskimpossibleifremoved?
3\.Whatisthecausalchainfrominput\-\>reasoning\-\>output?
Foreachpieceofnecessaryinformation,specify:
\-\*\*source\*\*:Always"input"\(allextractedinformationmustcomefrominputdata\)
\-\*\*semantic\_label\*\*:Ameaningfulidentifierforthispieceoffactualinformation
\-CreateclearlabelsdescribingWHATtheinformationIS,notwhereit'sstored
\-ForstructuredJSON:UsetheJSONpath\(e\.g\.,"patient\.age","diagnosis\.code"\)
\-Forunstructuredtext:Createsemanticidentifiers\(e\.g\.,"diagnosis\_code","hba1c\_result","physician\_name","patient\_age"\)
\-Makelabelsspecificenoughthatvalidationcanfindthesameinformationeveniftextisreorganized
\-\*\*value\*\*:Theextractedfactualvaluefromtheinputdata
\-Extractclean,specificfacts\(e\.g\.,"E11\.9","8\.2%","dog","63"\)
\-MUSTextracttheactualvaluefromthedata\-neverusenull
\-Ifyoucan'tfindavalue,don'tincludethisitem
\-Forlongquotedtextpassages:capturethekeysemanticconcepts,notverbatimquotes
\-\*\*decision\_criteria\*\*:Describewhatmakesthisvalueacceptableforproducingthesameoutput,includingexamplesofothervalidvalues
\-Statethedecisionrule/threshold/categorythattheoutputdependson
\-ListspecificexamplesofvaluesthatWOULDwork\(producesameoutput\)
\-ListspecificexamplesofvaluesthatWOULDNOTwork\(producedifferentoutput\)
\-Keyquestion:"Whatothervaluesinv2wouldproducethesameoutput?"
\-Whenoutputquotesevidencefrominput,specifythatparaphrasedversionspreservingsemanticmeaningareacceptable
\-Examples:
THRESHOLD\-BASED:
\-Output:"Age63,qualifiesforprogram\(under65requirement\)"
\-\>value:"63"
\-\>decision\_criteria:"Qualificationrequiresage<65\.Validvalues:Anyage0\-64\(e\.g\.,50,60,63,64\)\.Invalidvalues:65,66,70,etc\.\(wouldnotqualify\)\."
\-Output:"HbA1c8\.2%exceeds7\.0%threshold\-\>Approved"
\-\>value:"8\.2%"
\-\>decision\_criteria:"ApprovalrequiresHbA1c\>7\.0%\.Validvalues:7\.1%,8\.0%,8\.2%,9\.5%,12%\(any\>7\.0%\)\.Invalidvalues:6\.8%,7\.0%,5\.5%\(wouldbedenied\)\."
CATEGORY\-BASED:
\-Output:"Hasdog,qualifiesaspetownerfordiscount"
\-\>value:"dog"
\-\>decision\_criteria:"Qualificationrequirespetownership\.Validvalues:dog,cat,fish,bird,hamster,anypet\.Invalidvalues:nopet,none\(wouldnotqualify\)\."
\-Output:"E11\.9\(Type2Diabetes\)qualifiesforCGMcoverage"
\-\>value:"E11\.9"
\-\>decision\_criteria:"Qualificationrequiresdiabetesdiagnosis\.Validvalues:E10\.0\-E10\.9,E11\.0\-E11\.9,E13\.0\-E13\.9\(anydiabetescode\)\.Invalidvalues:J45\.0\(asthma\),I10\(hypertension\),etc\.\(wouldnotqualify\)\."
\-\*\*description\*\*:Human\-readableexplanationofwhatthisfactualinformationrepresents
\-\*\*reason\*\*:QuotetheEXACTtextfromexpectedoutputthatprovesthisisnecessary
\*\*RELATIONSHIPSBETWEENENTITIES\*\*
Afteridentifyingallnecessaryinformation,analyzethelogicalrelationshipsbetweentheseentitiesinthecontextofdecision\-making\.Relationshipsdescribehowmultiplepiecesofinformationinteracttoproducetheoutput\.
Keyrelationshiptypestoidentify\(NOTEXHAUSTIVE\):
1\.\*\*ORRelationships\*\*\(Alternative/Disjunctive\):
\-OutputdependsonatleastONEofmultipleconditionsbeingmet
\-Ifanyentitysatisfiesthecriteria,outputremainsthesame
\-Example:"Qualifiesfordiscount:petownerORseniorcitizen"
\-\>entities:\["pet\_ownership","age"\]
\-\>relationship:"OR\-qualificationrequireseithercondition"
\-\>description:"Ifv2losespet\_ownershipbutmaintainsage\>=65,outputpreserved\.Ifv2losesage\>=65butmaintainspet\_ownership,outputpreserved\.LosingBOTHbreaksoutput\."
2\.\*\*ANDRelationships\*\*\(Conjunctive/Required\):
\-OutputdependsonALLconditionsbeingmetsimultaneously
\-Ifanyentityfails,outputchanges
\-Example:"Approved:diabetesdiagnosisANDHbA1c\>7\.0%"
\-\>entities:\["diagnosis\_code","hba1c\_level"\]
\-\>relationship:"AND\-bothconditionsrequiredforapproval"
\-\>description:"v2mustpreserveBOTHdiabetesdiagnosisandHbA1c\>7\.0%\.Losingeitheronechangesoutputfromapprovedtodenied\."
3\.\*\*ConditionalRelationships\*\*\(If\-ThenDependencies\):
\-Oneentitytriggersrequirementforanother
\-Example:"Ifprocedure\_code=99213,thendiagnosis\_coderequiredforbilling"
\-\>entities:\["procedure\_code","diagnosis\_code"\]
\-\>relationship:"CONDITIONAL\-diagnosis\_coderequiredonlywhenprocedure\_code=99213"
\-\>description:"Ifv2changesprocedure\_codeto99214,diagnosis\_codebecomesunnecessary\.Ifv2keepsprocedure\_code=99213butremovesdiagnosis\_code,outputbreaks\."
4\.\*\*CompensatoryRelationships\*\*\(Trade\-offs\):
\-Lowvalueinoneentitycanbeoffsetbyhighvalueinanother
\-Example:"Riskscore:lowincomecompensatedbyhighcreditscore"
\-\>entities:\["income","credit\_score"\]
\-\>relationship:"COMPENSATORY\-highcredit\_scorecanoffsetlowincome"
\-\>description:"v2withincome$30k\-\>$25kacceptableifcredit\_score750\-\>800\.Bothdecreasingtogethermaychangeoutput\."
5\.\*\*HierarchicalRelationships\*\*\(Priority/Ordering\):
\-Orderorprioritymattersfordecisionlogic
\-Example:"Primarydiagnosisoverridessecondarydiagnosisforcoveragedetermination"
\-\>entities:\["primary\_diagnosis","secondary\_diagnosis"\]
\-\>relationship:"HIERARCHICAL\-primary\_diagnosistakesprecedence"
\-\>description:"v2swappingprimaryandsecondarydiagnoseschangesoutput\.Removingsecondary\_diagnosismaynotaffectoutputifprimaryissufficient\."
\*\*Whentoidentifyrelationships:\*\*
\-Multipleentitiescontributetoasingledecisionpointintheoutput
\-Changingoneentityaffectswhetheranotherentityisnecessary
\-Outputexplicitlymentionslogicaloperators\(or,and,either,both,unless\)
\-Decisioninvolvesthresholds/boundariesthatdependonmultiplevalues
\*\*Howtodescriberelationshipimpacts:\*\*
\-Specifywhatchangestoentityvalueswouldpreserveoutput
\-Specifywhatchangeswouldbreakoutput
\-Explaindependencychains\(ifAchanges,doesBmattermore/less?\)
\-Referencespecificdecision\_criteriafromtheinvolvedentities
\*\*Formatforeachrelationship:\*\*
\-\*\*entities\*\*:Listsemantic\_labelsofallrelatedentities\(mustmatchsemantic\_labelsfromnecessary\_information\)
\-\*\*relationship\*\*:Shortclassificationofrelationshiptype\(OR,AND,CONDITIONAL,COMPENSATORY,HIERARCHICAL,orcustomdescription\)
\-\*\*description\*\*:Detailedexplanationofhowchangestotheseentitiesinteractandaffectoutputpreservation
Leaverelationshipsarrayemptyifallnecessaryinformationitemsareindependent\(nologicalrelationships\)\.
\*\*CRITICAL\*\*:Iftheoutputiscomprisedofmultiplecomponents,e\.g\.multipleanswersorsections,makesureyouextracttheinformationrequiredtoproduce\*\*ALL\*\*componentsoftheoutput\.
Therequestedtasksemanticsmustalsoappearinarelationshipwithevery
factorentityneededtoanswerit\.Describewhythesamefactsarenotenough
whenthetaskasksforadifferentoperationoroutputform\.
Beprecise\!OnlyincludeinformationthatisDIRECTLYusedtoproducetheoutput\.
Don'tincludenice\-to\-haveorcontext\-onlyinformation\.
Examplesofnecessaryinformation:
\-"patient\_id:12345"\-\>Neededbecauseoutputreferencesthispatient
\-"diagnosis\_code:E11\.9"\-\>Neededbecauseoutputmakescoveragedecisionbasedonthis
\-"procedure\_date:2024\-01\-15"\-\>Neededbecauseoutputcheckstimingrequirements
ExamplesofNOTnecessary:
\-"phone\_number"\-\>Onlyusedforcontact,notfordecisionlogic
\-"formattingmetadata"\-\>Doesn'taffectthereasoning
</INSTRUCTIONS\>
<OUTPUT\_FORMAT\>
ReturnaJSONobjectwithnecessaryinformationitemsandtheirrelationships:
\{
"necessary\_information":\[
\{
"source":"input",
"semantic\_label":"meaningfulidentifierforthisinformation\(e\.g\.,patient\_age,diagnosis\_code\)",
"value":"extractedfactualvaluefrominputdata\(nevernull\)",
"decision\_criteria":"whatmakesthisvalueacceptable,withexamplesofvalid/invalidvalues",
"description":"humanreadabledescriptionofthisinformation",
"reason":"quoteexacttextfromexpectedoutputthatprovesthisisnecessary"
\},
\.\.\.
\],
"relationships":\[
\{
"entities":\["semantic\_label1","semantic\_label2"\],
"relationship":"OR\|AND\|CONDITIONAL\|COMPENSATORY\|HIERARCHICAL\|CUSTOM",
"description":"detailedexplanationofhowchangestotheseentitiesinteractandaffectoutputpreservation"
\},
\.\.\.
\],
"reasoning":"overallexplanationofhowthesepiecesconnectinputtooutput"
\}
</OUTPUT\_FORMAT\>
<BASELINE\_EXECUTION\_PROMPT\>
$\{baseline\_prompt\}
</BASELINE\_EXECUTION\_PROMPT\>
<PARENT\_INPUT\>
$\{parent\_input\}
</PARENT\_INPUT\>
<EXPECTED\_OUTPUT\>
$\{expected\_output\}
</EXPECTED\_OUTPUT\>
RespondwithasinglevalidJSONobjectonly\.DonotincludeMarkdown,prose,oranalysis\.
<TASK\>
Validatethatanewversionofacasepreservesallnecessaryinformation\.
</TASK\>
<INSTRUCTIONS\>
Youaregiven:
1\.AlistofNECESSARY\_INFORMATIONextractedfromversion1ofacase
2\.Version2'sINPUTandSUPPLEMENTARYdata
YourjobistoverifythatALLnecessaryinformationisstillpresentinversion2\.
ForeachiteminNECESSARY\_INFORMATION:
1\.Usethesemantic\_labeltounderstandWHATinformationtolookfor
2\.Searchforthisinformationinv2'sINPUTdata\(sourceisalways"input"\)
3\.Reviewthedecisioncriteriatounderstandwhatvaluesareacceptableforproducingthesameoutput
4\.Checkifv2'svaluesatisfiesthedecisioncriteria
5\.Markas"preserved"or"missing"
Bethorough\!FocusonwhethertheFACTUALINFORMATIONispresentinv2'sinput,regardlessofhowit'sstructuredorphrased\.
Thesemantic\_labeltellsyouwhattolookfor\-findthatinformationanywhereinv2'sinputdata\.
CRITICAL\-RequestedTaskSemantics:
When\`necessary\_information\`contains\`requested\_task\_semantics\`,youMUST
readtheexplicitversion\-2requestandcompareitagainsttheextracted
field/entity,operation,direction,periods,scope,units,andrequiredanswer
form\.Donotinfertherequestfromsurroundingfacts\.Failthecandidatewhen
itchangesapercentagechangeintoaratio,anendingvalueintoareturn,a
field/entityintoarelatedfield/entity,comparisondirectionorperiod,ora
supportedquestionintoanunsupportedone\.Factsremaininginthecontextdo
notpreservecorrectnessiftherequestedtaskchanges\.
Thesuppliedauditinputisacomplete,line\-wrappedrenderingoftheoriginal
JSONartifact\.Readthroughitstrailingrequestratherthanassumingthata
longinputisintactorthatthetaskmatchesversion1\.
HANDLINGQUOTESTEXT:HandlingEvidenceandQuotedTextinOutput
Whentheexpectedoutputincludesdirectquotesorextractiveevidencefromtheinput:
\-TheSEMANTICCONTENTandDECISION\-CRITICALINFORMATIONmustbepreserved
\-TheEXACTWORDINGdoesnotneedtomatchifthemeaningisidentical
\-Paraphrasedtextthatconveysthesamefacts,concepts,anddecisioncriteriaisACCEPTABLE
\-Onlymarkasmissingiftheparaphrasingchangesthesemanticmeaningorfailsdecision\_criteria
CRITICAL:ValueMatchingLogicusingdecision\_criteria
\-Eachitemhas\*\*value\*\*\(fromv1\)and\*\*decision\_criteria\*\*\(explainswhatv2valueswouldproducesameoutput\)
\-Usedecision\_criteriatoevaluateifv2'svaluepreservestheinformation:
\-decision\_criteriadescribesthedecisionruleandprovidesexamplesofvalid/invalidvalues
\-Checkifv2'svaluemeetsthecriteriadescribed
\-Examples:
THRESHOLD\-BASED:
\-v1:"63",decision\_criteria:"age<65required\.Valid:0\-64\.Invalid:65\+"
\-\>v2has"64"\(64<65,meetscriteria\)
\-\>v2has"66"\(66\>=65,failscriteria\)
\-v1:"8\.2%",decision\_criteria:"HbA1c\>7\.0%required\.Valid:\>7\.0%\.Invalid:<=7\.0%"
\-\>v2has"8\.5%"\(8\.5%\>7\.0%,meetscriteria\)
\-\>v2has"6\.8%"\(6\.8%<=7\.0%,failscriteria\)
CATEGORY\-BASED:
\-v1:"dog",decision\_criteria:"petownershiprequired\.Valid:dog,cat,fish,bird,anypet\.Invalid:nopet"
\-\>v2has"cat"\(catisapet,meetscriteria\)
\-\>v2has"nopet"\(nopet,failscriteria\)
\-v1:"E11\.9",decision\_criteria:"diabetescoderequired\.Valid:E10\.\*,E11\.\*,E13\.\*\.Invalid:non\-diabetescodes"
\-\>v2has"E10\.5"\(E10\.5isdiabetes,meetscriteria\)
\-\>v2has"J45\.0"\(J45\.0isasthma,failscriteria\)
CRITICAL:SemanticSpecificity\-TheExactValueMustBePreservedorMeetDecisionCriteria
DONOTassumesemanticlabelsarepreservedjustbecausesimilarconceptsexistinv2\.YoumustverifytheSPECIFICvalueanditssemanticmeaning\.
\*\*CommonFalsePositiveErrorstoAvoid:\*\*
1\.\*\*Over\-generalizationofconcepts\*\*:
\-semantic\_label="loan\_type",v1:"mortgageloan"\-\>v2has"loan"
\*"loan"istoogeneral\-couldbepersonalloan,autoloan,etc\.
\*TheSPECIFICtype\(mortgage\)ismissing
\-semantic\_label="systolic\_blood\_pressure",v1:"140mmHg"\-\>v2has"bloodpressuremeasured"
\*Presenceof"bloodpressure"conceptisinsufficient
\*TheSPECIFICVALUE\(140\)andTYPE\(systolicvsdiastolic\)aremissing
2\.\*\*Matchingonkeywordswithoutverifyingsemantics\*\*:
\-semantic\_label="approval\_date",v1:"approvedon2024\-01\-15"\-\>v2has"applicationdate2024\-01\-15"
\*Bothhavedates,butDIFFERENTsemanticmeanings\(approvaldate\!=applicationdate\)
\-semantic\_label="primary\_author",v1:"Dr\.Smith"\-\>v2has"reviewedbyDr\.Smith"
\*Samepersonname,butDIFFERENTroles\(author\!=reviewer\)
3\.\*\*Ignoringqualifiersandmodifiers\*\*:
\-semantic\_label="annual\_revenue",v1:"$500K/year"\-\>v2has"$500Kinvestmentround"
\*Sameamount,butDIFFERENTfinancialconcepts\(revenue\!=investment\)
4\.\*\*Partialvaluematches\*\*:
\-semantic\_label="medication\_name\_and\_dosage",v1:"Metformin500mg"\-\>v2has"Metformin"
\*Medicationnamepresentbutdosagemissing\-incompletematch
\*\*ValidationProtocol:\*\*
Beforemarkingas"preserved",askyourself:
1\.Doesv2containtheEXACTsemanticconceptidentifiedbythesemantic\_label?
2\.Isthevalueinv2theSAMEvalueordoesitmeetthedecision\_criteria?
3\.AreALLqualifiers,modifiers,andcontextpreserved\(type,scope,role,etc\.\)?
4\.Wouldtheoutputtextremainaccuratewithv2'svalue?
5\.For\`requested\_task\_semantics\`,doesv2explicitlyaskfortheidentical
taskandanswerformratherthanmerelycontainingthesamesupportingfacts?
IftheanswertoANYoftheseisNO,markas"missing"withclearexplanationofthesemanticmismatch\.
ExamplesofwhattomarkasINVALID\(differentsemanticmeaning\):
\-semantic\_label="patient\_age",v1:"65"\-\>v2has"ageatdiagnosis:63"\(currentagevshistoricalage\)
\-semantic\_label="current\_medications"\-\>v2onlyhas"past\_medications"\(temporalscopediffers\)
\-semantic\_label="primary\_diagnosis"\-\>v2onlyhas"secondary\_diagnosis"\(prioritydiffers\)
CRITICAL:ValidatingRelationshipsBetweenEntities
Aftervalidatingindividualentities,youMUSTvalidatethatrelationshipsbetweenentitiesarepreserved:
\*\*Keyvalidationprincipleforrelationships:\*\*
\-Eachrelationshipdescribeshowentitychangesinteracttoaffectoutput
\-Individualentitiesmaychange,buttherelationshiplogicmuststillproducethesameoutput
\-Referencetherelationshipdescriptiontounderstandwhatchangesareacceptable
\-Arelationshipfailsifthecombinedeffectofentitychangeswouldaltertheoutput
Examplesbyrelationshiptype:
\-\*\*ORrelationship\*\*\(pet\_ownershipORage\>=65\):v2mustsatisfyatleastONEcondition
\-\*\*ANDrelationship\*\*\(diagnosisANDhba1c\>7%\):v2mustsatisfyBOTHconditions
\-\*\*CONDITIONAL\*\*\(IFprocedure=XTHENdiagnosisrequired\):iftriggerpresentinv2,dependentmustbepresent
\-\*\*COMPENSATORY\*\*\(incomevscredit\_score\):combinedeffectmustpreserveoutputdecision
\-\*\*HIERARCHICAL\*\*\(primary\>secondary\):priority/orderingmustbemaintained
\-\*\*CUSTOM\*\*\(e\.g\.,"total\_cost=quantity\*unit\_price"\):mathematical/logicalrelationshipmustholdinv2
Classificationrules:
\-\*\*preserved\*\*:
\-TheSPECIFICinformationidentifiedbysemantic\_labelexistsinv2withcorrectsemanticmeaning
\-v2'svaluemeetsthedecision\_criteria\(producessameoutput\)
\-ALLqualifiers,modifiers,types,andcontextmatchthesemantic\_label
\-Structure/phrasingcandiffer,butthesemanticconceptmustbeIDENTICAL
\-\*\*missing\*\*:
\-Informationidentifiedbysemantic\_labelcompletelyabsentfromv2
\-Informationexistsbutv2'svaluedoesn'tmeetdecision\_criteria\(wouldproducedifferentoutput\)
\-Informationexistsinwrongsemanticcontext\(e\.g\.,theentityinv1speaksaboutemotionalageandtheentityinv2hasbeenmodifiedtospeakaboutphysicalage\)
\-Similarconceptexistsbutlacksspecificity\(e\.g\.,semantic\_label="mortgage\_loan"butv2onlyhas"loan"\)
Finalverdict:
\-\*\*pass\*\*:ALLnecessaryinformationispreservedANDALLrelationshipsarepreserved
\-\*\*fail\*\*:ANYnecessaryinformationismodifiedinadecision\-criticalmanner
CRITICAL:The\`pass\`fieldMUSTmatchyourreasoningandmissing\_informationlist:
\-Ifmissing\_informationisNOTempty\-\>passMUSTbefalse
\-Ifyourreasoningsays"fails","doesnotpreserve","missing",or"broken"\-\>passMUSTbefalse
\-IfyouidentifyANYmissinginformationorbrokenrelationships\-\>passMUSTbefalse
\-ONLYsetpass=trueifmissing\_informationisemptyANDallentitiespreservedANDallrelationshipspreserved
IMPORTANT:BegenerouswithstructuralequivalencebutEXTREMELYstrictwithsemanticequivalence\.
\-Thephrasing/wordingcandiffer\.Verifythattheinformationpresentedwithinthenecessaryentityisnotpresentelsewhereinasaparaphrase\.
\-BUTthesemanticmeaningmustbeIDENTICAL\-sameconcept,samespecificity
\-Whenindoubt,markasmissingandexplainthesemanticdifference
\-Itisbettertohaveafalsenegative\(incorrectlymarkingasmissing\)thanafalsepositive\(incorrectlymarkingaspreserved\)
</INSTRUCTIONS\>
<OUTPUT\_FORMAT\>
ReturnaJSONobject:
\{
"pass":true\|false,
"preserved\_information":\[
\{
"semantic\_label":"entityidentifierfromnecessary\_information",
"found\_at":"locationordescriptionofwherefoundinv2",
"status":"exact\_match"\|"structural\_equivalent"\|"value\_modified"
\},
\.\.\.
\],
"missing\_information":\[
\{
"semantic\_label":"entityidentifierthatismissing",
"description":"whatinformationismissingorhowitfailscriteria",
"impact":"whythisbreakstheabilitytoproduceexpectedoutput"
\},
\.\.\.
\],
"rationale":"overallassessmentofwhetherv2preservesallnecessaryinformationandrelationships",
"validation\_reasoning":"detailedexplanationofhoweachentityandrelationshipwasvalidated"
\}
</OUTPUT\_FORMAT\>
<NECESSARY\_INFORMATION\>
$\{necessary\_information\_json\}
</NECESSARY\_INFORMATION\>
<IMMUTABLE\_BASELINE\_EXECUTION\_PROMPT\>
$\{baseline\_prompt\}
</IMMUTABLE\_BASELINE\_EXECUTION\_PROMPT\>
Theimmutablebaselineexecutionpromptaboveissuppliedseparatelybecauseitisnotpartofthemutablebenchmarkinput\.Itisguaranteedunchangedbetweentheoriginalandcandidate\.Treattaskinstructions,outputformat,andnormalizationrulesfoundthereaspreserved;doNOTrequirethemtoappearagaininVERSION\_2\_INPUT\.AuditonlywhetherVERSION\_2\_INPUTpreservesthecase\-specificinformationneededunderthatunchangedtask\.
<VERSION\_2\_INPUT\>
$\{candidate\_input\}
</VERSION\_2\_INPUT\>
The extraction fields arebaseline\_prompt,parent\_input, andexpected\_output\. The audit fields arenecessary\_information\_json,baseline\_prompt, andcandidate\_input\. The deterministic context builder additionally suppliescase\_id,generation,parent\_input\_path,parent\_input\_json\_path,expected\_output\_path,baseline\_execution\_prompt\_path,extraction\_path,audits\_path, andcandidates; each candidate containsindividual\_index,case\_input\_json\_path,case\_input\_path, andaudit\_path\.
#### A\.5\.6Realism audit\-call template
Each correctness\-approved candidate is evaluated with the following realism instruction\.
RespondwithonevalidJSONobjectonly\.DonotincludeMarkdown\.
Youareanindependentdomain\-realismauditor\.Themutationgeneratorand
correctnessauditorareseparateandhavenotsuppliedanyreasoningtoyou\.
JudgeonlywhetherCANDIDATE\_INPUTresemblesanauthenticinputfromthedomain
describedbytheimmutablecross\-sourcerubricandreferenceprofile\.
Applyeveryrubricdimensionindependently\.Returnanintegerscoreforevery
dimensionID\.Reportonlyred\-flagIDsdefinedbytherubric\.Donotjudge
answercorrectness,solvabilityagainstagoldanswer,mutationdifficulty,or
modelfitness\.Donotrevisetherubric\.
Returnexactly:
\{"pass":true\|false,"dimension\_scores":\{"D1":1\},"detected\_red\_flags":\[\],"rationale":"conciseevidence\-basedexplanation"\}
<IMMUTABLE\_RUBRIC\>
$\{rubric\}
</IMMUTABLE\_RUBRIC\>
<CROSS\_SOURCE\_REFERENCE\_PROFILE\>
$\{reference\_profile\}
</CROSS\_SOURCE\_REFERENCE\_PROFILE\>
<CANDIDATE\_INPUT\>
$\{candidate\_input\}
</CANDIDATE\_INPUT\>
The dynamic fields are the completerubric,reference\_profile, andcandidate\_input\. The model’spassis diagnostic\. Deterministic code verifies each member’s pass from the rubric: every required dimension score must be an integer in the rubric’s 1–5 range; the arithmetic mean must meetminimum\_mean\_score; every dimension must meetmet\_thresholdwhen required; and no defined blocking red flag may appear when blocking flags fail the case\. Missing or malformed scores fail\. Only correctness\-approved candidates are audited\. The ensemble verdict is then recomputed from member passes under majority or unanimity voting\.
#### A\.5\.7Single\-pass baseline template
Few\-Shot runs the following template without tools, while Few\-Shot \(web\-search\) runs it with the bounded read\-only web permissions described above\.
ReturnonevalidJSONobjectonly\.DonotincludeMarkdownoranytextoutsidethatobject\.
Createonerevisedversionof\`TARGET\_INPUT\`\.Therevisedinputmustbe
realistic,self\-contained,morechallengingtosolvecarefully,andstillhave
theexactsame\`IMMUTABLE\_EXPECTED\_OUTPUT\`\.
Usethefive\`REFERENCE\_EXAMPLES\`onlyasexamplesofnaturallychallenging
inputsinthistaskfamily\.Theirnumericscoresaredescriptivecontext\.Do
notcopyfacts,entities,numbers,wording,oranswersfromareferenceexample
intothetarget\.
Youmayusetheavailableread\-onlywebtoolstocheckdomainconventions,
terminology,andrealisticdocumentstructures\.Donotsearchforthetarget
record,itssourceID,itsexacttext,oritsanswer\.
Keeptheschemaandfieldtypesunchanged\.Preserveeveryfact,value,
relationship,qualifier,unit,date,andtaskconstraintneededtoderivethe
expectedoutput\.Donotintroducecontradictions,falsefacts,arbitrary
corruption,answerleaks,ambiguity,orinstructionsdirectedatthesolver\.
<TASK\_CONTRACT\>
\{task\_contract\}
</TASK\_CONTRACT\>
<INPUT\_SCHEMA\>
\{input\_schema\}
</INPUT\_SCHEMA\>
<REFERENCE\_EXAMPLES\>
\{hard\_reference\_examples\}
</REFERENCE\_EXAMPLES\>
<TARGET\_INPUT\>
\{target\_input\}
</TARGET\_INPUT\>
<IMMUTABLE\_EXPECTED\_OUTPUT\>
\{immutable\_expected\_output\}
</IMMUTABLE\_EXPECTED\_OUTPUT\>
Returnthisshape:
\{
"rewritten\_input":\{\},
"change\_summary":"shortdescriptionoftheaddedrealisticcomplexity",
"realism\_check":"whytherevisedinputisplausible",
"correctness\_retention\_check":"whytheexpectedoutputisunchanged"
\}
\`rewritten\_input\`mustbethecompletereplacementinputandconformexactlyto\`INPUT\_SCHEMA\`\.
The dynamic fields aretask\_contract,input\_schema,hard\_reference\_examples,target\_input, andimmutable\_expected\_output\. Each of the five reference entries containsscoreandinput\. The rendering implementation also acceptsbenchmark\_task\_contractas an alias fortask\_contract\.
### A\.6Benchmark details
FinQA is graded by the official evaluation script on the executed answer and on program equivalence to the gold calculation program\. PubMedQA is graded by exact match over the yes/no/maybe label and is explicitly described as plateaued in the literature for models of the scale we evaluate\. ContractNLI crosses 123 real non\-disclosure agreements with 17 standard diligence hypotheses \(2,091 pairs\), graded by exact match over the Entailment/Contradiction/NotMentioned label; the task mirrors first\-pass NDA review\. The three benchmarks also exercise different output structures: an executed numerical program versus fixed three\-way labels\.
##### Assets and licenses\.
### A\.7Task model prompts
All three benchmarks predate instruction\-tuned language models: their official baselines are fine\-tuned systems \(FinQANet, BioBERT\-based classifiers, and BERT/Span NLI\), so no official chat prompt exists for any of them\. We therefore author one task prompt per benchmark, reproduced verbatim below \(line wrapping added for presentation\)\.
Each task model evaluation sends one user message and no separate system message\. Each task prompt has three common parts: \(i\) the official task definition, transcribed from the benchmark’s paper; \(ii\) a machine\-parsable final line \(Program:/Answer:/Label:\), required because fitness and uncertainty are computed over ten sampled generations per evaluation, so every sample must parse deterministically\. We use regex answer extraction over a reason\-then\-answer format to score every sample consistently\. \(iii\) delimited input fields\. ContractNLI additionally instructs the model to treat the contract and hypothesis as quoted data rather than instructions\.
##### FinQA\.
The program\-language block transcribes the official FinQA grammar: six arithmetic operations, four table\-aggregation operations,const\_\*constants, and\#Nstep references from Chen et al\.\[[2](https://arxiv.org/html/2609.30571#bib.bib14)\], and predicted programs are executed by a transcription of the benchmark’s MIT\-licensed official evaluation script\. Prompting an LLM to emit an externally executed FinQA\-style program follows ZS\-FinDSL\[[17](https://arxiv.org/html/2609.30571#bib.bib18)\]; unlike ZS\-FinDSL, we keep the official table operators and constants\. The conventions block restates output conventions the official scorer already enforces \(a percent change is scored as a ratio; changes aresubtract\(later, earlier\)\); this is format disambiguation rather than a task hint, and removing it produced spurious format failures rather than reasoning failures\.
Youareafinancialanalystansweringaquestionaboutacompanyfilingby
writingareasoningprogram\.
Youaregiventhetextbeforeatable,thetableitself,thetextafterthe
table,andaquestion\.
Task:writeaprogramintheFinQAprogramlanguagethatcomputestheanswer
tothequestion\.
Programlanguage:
\-Arithmetic:add\(a,b\),subtract\(a,b\),multiply\(a,b\),divide\(a,b\),exp\(a,b\)
\-Comparison:greater\(a,b\)evaluatestoyesorno
\-Tableaggregation:table\_max\(row,none\),table\_min\(row,none\),
table\_sum\(row,none\),table\_average\(row,none\)
where\`row\`istheexactlabelinthetable'sfirstcolumn\.
\-Argumentsarenumberscopiedfromthereport\(writethembare,without$or
commas\),constantswrittenasconst\_100,const\_1000,const\_1000000,const\_m1,
or\#N,whichreferstotheresultofstepNcountingfrom0\.
\-Chainmultiplestepsbyseparatingthemwith","\.Stepscannotbenested:
add\(1,add\(2,3\)\)isinvalid,sowriteadd\(2,3\),add\(1,\#0\)instead\.
Conventionsthisbenchmarkexpects:
\-Leavepercentagesandratesasdecimalratios\.For"whatpercentage\.\.\."or
"whatwasthepercentchange\.\.\.",stopatthedivision\.Write
divide\(849,5424\),notdivide\(849,5424\),multiply\(\#0,const\_100\)\.
\-Achangefromanearliervaluetoalatervalueissubtract\(later,earlier\)\.
Apercentchangedividesthatdifferencebytheearliervalue\.
\-Alwaysanswerwithaprogram,neverwithabarenumber\.
Examples:
Program:subtract\(5829,5735\)
Program:subtract\(5829,5735\),divide\(\#0,5735\)
Program:table\_average\(netrevenue,none\)
Outputstructure:reasonstepbystepfirst,thenendyourresponsewitha
singlefinallineinexactlythisformat:
Program:subtract\(5829,5735\)
<pre\_text\>
\{pre\_text\}
</pre\_text\>
<table\>
\{table\}
</table\>
<post\_text\>
\{post\_text\}
</post\_text\>
<question\>
\{question\}
</question\>
##### PubMedQA\.
The prompt is exactly the reasoning\-required setting of Jin et al\.\[[7](https://arxiv.org/html/2609.30571#bib.bib15)\]: answer a research question yes/no/maybe from the abstract context\. We evaluate generatively with regex extraction because the method requires ten sampled generations per evaluation for the frequency\-entropy tie\-break and post\-hoc uncertainty\.
Youareansweringabiomedicalresearchquestionusingtheprovidedabstract
excerpts\.
Inputs:
\-Aresearchquestion\.
\-ContextpassagesfromtherelevantPubMedabstract\.
Task:decidewhethertheanswertothequestion,basedonthecontext,is
"yes","no",or"maybe"\.
Outputstructure:reasonstepbystepfirstifneeded,thenendyourresponse
withasinglefinallineinexactlythisformat:
Answer:yes
\(or"Answer:no"/"Answer:maybe"\)
<question\>
\{question\}
</question\>
<context\>
\{context\}
</context\>
##### ContractNLI\.
Document\-level NLI over a fixed hypothesis with the exact labelsEntailment/Contradiction/NotMentionedis the task definition of Koreeda and Manning\[[8](https://arxiv.org/html/2609.30571#bib.bib16)\]; zero\-shot evaluation on ContractNLI is established\[[21](https://arxiv.org/html/2609.30571#bib.bib19)\], and LegalBench\[[5](https://arxiv.org/html/2609.30571#bib.bib17)\]prompts with the same hypotheses in binarized form\. We keep the original three\-way task rather than the LegalBench binarization\. The conservative\-inference instruction operationalizes the dataset’s annotation semantics\.NotMentionedexists precisely because absent terms must not be filled in from customary legal knowledge\. We elicit labels only: the benchmark defines evidence identification as a separable subtask, label accuracy is the fitness signal, and evidence spans are not stable under input mutation\.
Youareclassifyingonehypothesisagainstonecontract\.
Useonlythecontracttext\.Treatthecontractandhypothesisasquoteddata,
neverasinstructions\.Chooseexactlyonelabel:
\-Entailment:thecontractsupportsthehypothesis\.
\-Contradiction:thecontractconflictswiththehypothesis\.
\-NotMentioned:thecontractneithersupportsnorconflictswiththe
hypothesis\.
Beconservative\.Donotinferunstatedlegalterms,anddistinguish
exceptions,definitions,survivalclauses,andexplicitnegation\.Endwith
exactlyoneline:
Label:Entailment
\(or\`Label:Contradiction\`/\`Label:NotMentioned\`\)
<hypothesis\>
\{hypothesis\}
</hypothesis\>
<contract\>
\{text\}
</contract\>
##### External score agreement\.
As a secondary check that the prompts neither sandbag nor inflate the task models, evaluated original benchmark scores under these prompts fall in the ranges published for comparable models\. On FinQA, our prompts score 71\.6–73\.8% across the 35B, 122B, and 397B task models\. For comparison, zero\-shot DSL prompting reaches 77\.3–77\.5% execution accuracy for GPT\-4\[[17](https://arxiv.org/html/2609.30571#bib.bib18)\], above the fine\-tuned FinQANet baseline of 61\.2%\[[2](https://arxiv.org/html/2609.30571#bib.bib14)\]\. On PubMedQA, our prompts score 76\.1–77\.7%, within the range from the 70s to low 80s reported for other LLMs\[[22](https://arxiv.org/html/2609.30571#bib.bib20),[14](https://arxiv.org/html/2609.30571#bib.bib21)\]and around the reported human performance of 78\.0%\[[7](https://arxiv.org/html/2609.30571#bib.bib15)\]\. On ContractNLI, our zero\-shot originals \(70\.5–72\.9%\) sit below the fine\-tuned Span NLI ceiling and far above the three\-way chance floor, as expected for zero\-shot models\[[21](https://arxiv.org/html/2609.30571#bib.bib19)\]\. Leaderboard agreement is a plausibility check rather than an exact calibration\. The load\-bearing controls remain the benchmark\-native task and the paired within\-prompt design\.
### A\.8Baseline configuration
Both baselines use the mutation model with one request per benchmark case and no multi\-turn loop\. Each request contains the task instructions, the allowed input fields, five reference examples with their task model scores, the target input, and the immutable expected output\. Appendix[A\.5\.7](https://arxiv.org/html/2609.30571#A1.SS5.SSS7)reproduces the template and its dynamic fields\.
Reference examples are drawn from a shared per\-benchmark bank: a 100\-case pool disjoint from the targets is scored with the reasoning\-off largest task model using ten sampled generations per case, and the 20 lowest\-scoring cases are retained\. Each benchmark case receives five bank references sampled without replacement using a case\-specific deterministic seed\. These references serve only as difficulty inspiration and must not supply content\.
Both baselines receive the same task instructions, template prompt, and allowed input fields\. The instructions describe the task without naming the benchmark and state what must remain true\. For example, FinQA must still support the same executable program, and ContractNLI must retain the fixed\-hypothesis label\. Few\-Shot has no web, filesystem, or other tools\. Few\-Shot \(web\-search\) adds read\-only web search and fetch for domain conventions and realistic structures\. The source identifier, exact target text, and answer are prohibited from web lookup\.
Correctness and realism checks are applied post hoc to both baselines using the same three\-member ensembles as the evolutionary runs\. Failed mutations revert to their original cases for constraint\-filtered aggregation\.
### A\.9Why the single\-pass baselines inject complexity weakly
Qualitative inspection shows that Few\-Shot primarily lengthens or paraphrases the context while leaving the question, expected output, and answer\-bearing evidence unchanged\. Representative PubMedQA edits expand “15 degrees head\-down tilt” into a longer description of the same Trendelenburg position or replace “lower limbs elevated for prolonged periods” with an equivalent phrase\. Similar edits add hedging and non\-decisive context; they increase surface complexity without changing the evidence that determines the answer\.
Web access makes Few\-Shot \(web\-search\) edits more domain\-authentic, but not necessarily more difficult\. Observed FinQA edits add thousands separators, period headers, and filing\-style row labels while preserving every operative number\. PubMedQA edits introduce clinical register and equivalent units around the same findings\. ContractNLI edits add legal boilerplate around unchanged operative clauses\. These changes improve realism, but the decisive quantities, findings, and clauses remain directly available\. This distinction is especially pronounced for the fixed three\-way labels of PubMedQA and ContractNLI, which are robust to paraphrase and padding\. FinQA is somewhat more sensitive to formatting and indirect phrasing, but its executable program still depends on preserved numbers and question intent\.
Neither baseline method observes task model performance or searches over alternatives\. It therefore cannot select the rare edit that is simultaneously valid, realistic, and difficult for a particular model\. The evolutionary method can compound several valid changes\. The qualitative and quantitative results therefore have the same explanation: web research improves authenticity, while model\-conditioned selection produces difficulty\.
### A\.10Accuracy and uncertainty results
Table[4](https://arxiv.org/html/2609.30571#A1.T4)gives the numerical values plotted in Figure[1](https://arxiv.org/html/2609.30571#S4.F1)\. Every entry uses the same 200 cases per benchmark and task model\.To quantify case\-sampling uncertainty, we use percentile bootstrap intervals\[[3](https://arxiv.org/html/2609.30571#bib.bib22)\]\. We first average the ten model responses for each case\. For each condition, we then samplenncase\-level means with replacement 50,000 times \(seed 42\), recompute overall accuracy, and report the 2\.5th and 97\.5th percentiles\. Model outputs remain fixed during resampling\.
Table 4:Task model accuracy \(%\) under each mutation method\.Few\-ShotandFew\-Shot \(web\-search\)entries report unfiltered \(constraint\-filtered\) accuracy: first on all 200 mutations, then after cases that fail either LM\-based check revert to the original case \(Appendix[A\.8](https://arxiv.org/html/2609.30571#A1.SS8)\)\.Table[5](https://arxiv.org/html/2609.30571#A1.T5)accompanies the accuracy results of Table[4](https://arxiv.org/html/2609.30571#A1.T4), reporting normalized discrete semantic entropy over the ten sampled completions for the same cases and methods\.
Tables[6](https://arxiv.org/html/2609.30571#A1.T6)–[8](https://arxiv.org/html/2609.30571#A1.T8)report the complementary post\-hoc uncertainty measures\. Higher values indicate greater predictive uncertainty or lower model\-assigned likelihood\. Relative to Original,HARDENincreases all three measures in every benchmark\-model combination\. Sequence NLL is sensitive to completion length; mean\-token NLL controls for this by dividing by the greedy completion’s token count\.
Table 5:Mean normalized discrete semantic entropy under each mutation method\. Sample counts follow Table[4](https://arxiv.org/html/2609.30571#A1.T4)\.Few\-ShotandFew\-Shot \(web\-search\)entries use the same unfiltered \(constraint\-filtered\) convention as Table[4](https://arxiv.org/html/2609.30571#A1.T4)\.Table 6:Mean likelihood\-weighted semantic entropy under each mutation method\.Few\-ShotandFew\-Shot \(web\-search\)entries report unfiltered \(constraint\-filtered\) values\.Table 7:Mean sequence negative log\-likelihood of the greedy completion under each mutation method\. Higher values indicate lower sequence likelihood and are sensitive to completion length\.Few\-ShotandFew\-Shot \(web\-search\)entries report unfiltered \(constraint\-filtered\) values\.Table 8:Mean token negative log\-likelihood of the greedy completion under each mutation method\. This is sequence NLL divided by token count; it is not exponentiated perplexity\. Higher values indicate lower mean token likelihood\.Few\-ShotandFew\-Shot \(web\-search\)entries report unfiltered \(constraint\-filtered\) values\.##### Hardened\-subset effects\.
Because cases the search cannot mutate within budget revert to their originals, the dataset\-level numbers of Tables[4](https://arxiv.org/html/2609.30571#A1.T4)and[5](https://arxiv.org/html/2609.30571#A1.T5)dilute the per\-case effect of a successful HARDEN mutation\. Table[9](https://arxiv.org/html/2609.30571#A1.T9)restricts each run to cases with successful post\-filtered mutations only\. Accuracy falls by 25–64 points and normalized semantic entropy rises by 0\.28–0\.43, while each immutable evaluation target remains unchanged\.
Table 9:Original\-versus\-HARDENcomparison restricted to successfully hardened cases \(non\-hardened cases excluded\)\.
### A\.11Human\-participant check\-validation study
We compare the agentic checks with human judgments on matched FinQA and PubMedQA artifacts\. The study did not independently test participants’ domain competence, so we refer to them as role\-screened participants\.
##### Protocol\.
Participants were recruited on Prolific and screened for English fluency and occupational role: accounting or finance roles for FinQA, and nursing, general\-practice, or other healthcare roles for PubMedQA\. Every case received three independent participant judgments\. Collection was split into phases with disjoint case sets, so a participant could complete several tasks without ever rating the same case twice\. Pilot runs with two participants per study calibrated task length before the main collection\. Participation was voluntary and minimal risk, and participants could withdraw at any time\. No directly identifying information was collected; retained data comprised pseudonymous identifiers, realism/correctness judgments, written rationales, and task duration\. All study descriptions were configured with an estimated completion time of 12 and 15 minutes per session for realism and correctness respectively, adjusted up from 10 minutes each after the first phase revealed higher average completion times\. These conditions were disclosed in the study information and consent materials\. The application stored pseudonymous platform participant, study, and session identifiers, structured judgments, written rationales, and duration under server\-side access controls\. The production\-study materials below evidence the participant\-facing instructions and screenshots; they do not independently evidence the recruitment screening, compensation, consent and withdrawal process, researcher training, employer compliance, or IRB status reported here\.
##### Participant\-facing interfaces\.
Participants had to finish a study\-introduction video and acknowledge the instructions before beginning\. Figures[4](https://arxiv.org/html/2609.30571#A1.F4)and[5](https://arxiv.org/html/2609.30571#A1.F5)reproduce the benchmark\-specific correctness and realism introductions\. The response controls are shown in Figure[6](https://arxiv.org/html/2609.30571#A1.F6)\. The captures are redacted and contain no participant, platform, or deployment identifiers\.
##### Realism task\.
For each benchmark, the study contains 60 cases: 20 official benchmark cases, 20 synthetic cases authored by GPT\-5\.6 Sol, and 20 synthetic cases authored by GPT\-3\.5 Turbo\. Each task presents one case from each source in randomized order\. Three participants label every case realistic or unrealistic and provide rationales; the human label is the two\-of\-three majority\. We compare these labels with the three\-member feasibility ensemble on matched case banks\.
##### Correctness task\.
Each task presents two original–mutated pairs drawn from real HARDEN runs\. Participants first explain the original case’s solution, having been informed of its correctness, then assess whether the highlighted mutation preserves correctness and justify the decision\. Each benchmark contributes 20 gold\-labeled pairs, balanced between ten correct and ten incorrect mutations, with three independent participant judgments per pair\. The human label is again the participant majority\.


Figure 4:Redacted correctness\-study introduction screens for FinQA \(top\) and PubMedQA \(bottom\)\. The crops omit the embedded video preview and show no participant or platform identifiers\.

Figure 5:Redacted realism\-study introduction screens for FinQA \(top\) and PubMedQA \(bottom\), including the benchmark\-specific realism definitions and task descriptions\. No participant or platform identifiers are shown\.


Figure 6:Participant\-facing response controls\. The original\-case explanation screen \(top\) is followed by the modified\-case correctness judgment and justification screen \(middle\); the realism judgment and rationale screen is shown at bottom\. These controls are identical across FinQA and PubMedQA, so only one capture of each is included\.
##### Analysis\.
For realism, we report source\-wise realistic rates in Table[10](https://arxiv.org/html/2609.30571#A1.T10); direct check–human agreement, Cohen’sκ\\kappa, Gwet’s AC1, positive and negative agreement, andϕ\\phi/MCC in Table[11](https://arxiv.org/html/2609.30571#A1.T11)\. We additionally test the pre\-specified official\-versus\-GPT\-3\.5 source comparison with a two\-sided Fisher exact test\. For correctness, we report confusion counts with correct as the positive class, accuracy, specificity,κ\\kappa, and an exact McNemar test with Holm adjustment across the two benchmarks\.
Table 10:Percentage of cases labeled realistic by the agentic check and the two\-of\-three participant majority\. Each row contains 20 cases\.Table 11:Realism agreement between the agentic check and participant majority over 60 cases per benchmark\.
##### Realism results\.
Tables[10](https://arxiv.org/html/2609.30571#A1.T10)and[11](https://arxiv.org/html/2609.30571#A1.T11)provide the detailed statistics behind the summary in Section[4\.2](https://arxiv.org/html/2609.30571#S4.SS2)\. On FinQA, the check attains 95% provenance accuracy \(38/40\), but direct agreement with participant majorities is 51\.7%\. The observed participant acceptance rate is 60% for both official and GPT\-3\.5 cases \(p=1\.000p=1\.000, two\-sided Fisher exact test\)\. With 20 cases per source, this estimate is imprecise and does not establish equal population rates or imply that finance experts generally cannot distinguish the sources\.
On PubMedQA, AC1 is 0\.820, participant majorities separate official from GPT\-3\.5 cases \(p=0\.0197p=0\.0197\), and feasibility check provenance accuracy is 55% \(22/40\)\. Together with the 85\.0% overall agreement and 18\.2% agreement regarding unrealistic cases in Table[11](https://arxiv.org/html/2609.30571#A1.T11), these results support the conclusion: the check is well aligned with humans on realistic examples but over\-permissive on unrealistic ones, motivating threshold recalibration\.
Table 12:Correctness agreement on 20 gold\-labeled mutated pairs per benchmark\. Confusion counts are TP/FN/FP/TN with correct as positive\.
##### Correctness results\.
Table[12](https://arxiv.org/html/2609.30571#A1.T12)quantifies the low agreement summarized in Section[4\.2](https://arxiv.org/html/2609.30571#S4.SS2): participant specificity is 10\.0% on FinQA and 0\.0% on PubMedQA, with a significant correct\-rate mismatch on both benchmarks after Holm adjustment \(p=0\.0039p=0\.0039\)\. These results do not by themselves determine whether strictness comes from superior check sensitivity or an interface that makes semantic errors difficult for participants to identify\. The secondary study below therefore evaluates the check directly against a more informative labeled bank\.
### A\.12Secondary correctness\-check validation
We therefore conduct a second validation against correctness labels assigned by the co\-authors rather than participant\-majority labels\. Only pairs on which the co\-authors agree enter the evaluation set; split judgments are excluded from all reported metrics, and the agreed co\-author verdict serves as the gold label\. We report the completed FinQA adjudication and leave a corresponding PubMedQA evaluation to future work\. The submitted FinQA bank contains 25 pairs: ten manually mutated with the intent of making them incorrect, five manually mutated while preserving correctness, five generated by a single\-model counterfactual operation, and five generated by a single\-model label\-flip operation\. Four pairs receive split co\-author judgments, leaving 21 for evaluation\. The counterfactual prompt requests the smallest realistic change that invalidates the original output; the label\-flip prompt requests stronger changes to decision\-critical source evidence while preserving the question\. The production correctness check uses three independent Claude Opus 4\.8 calls at temperature zero and returns the two\-of\-three majority\.
Table 13:Correctness\-check agreement on the co\-author\-agreement evaluation set, by construction method\. Numerators are matches to the agreed co\-author verdict; denominators exclude split judgments\.Table 14:Secondary correctness performance on the co\-author\-agreement evaluation set, with correct as the positive class\. Confidence intervals are Wilson 95% intervals for accuracy\.As summarized in Table[14](https://arxiv.org/html/2609.30571#A1.T14), the check matches 18 of 21 agreed FinQA labels, with 81\.6% balanced accuracy, Cohen’sκ=0\.577\\kappa=0\.577, andϕ\\phi/MCC=0\.583\{\}=0\.583\. Table[13](https://arxiv.org/html/2609.30571#A1.T13)shows that it rejects all ten single\-model counterfactual and label\-flip examples and accepts all three agreed manually degraded correct controls\. These results show that the check can distinguish explicit output\-invalidating mutations from the tested valid controls; the first study’s low participant specificity should therefore not be interpreted as evidence that the check merely rejects indiscriminately\.
The three disagreements expose errors in both directions\. The check rejects one manually constructed pair that both co\-authors judge correct, treating a changed table header as decision\-critical despite compensating narrative context\. It accepts two pairs that both co\-authors judge incorrect, overlooking a location mismatch when the numeric operands remain and a shifted temporal bucket in a financial table\. All three feasibility check decisions are unanimous\.
This secondary validation is deliberately small and class\-imbalanced, and four of the 25 submitted FinQA pairs are excluded after split judgments\. Its model\-generated negatives were prompted with the original output and may be more explicit than naturally occurring invalid mutations; the three agreed correct controls are also all manually authored\. We therefore interpret the results as evidence that the check is a useful but incomplete validity filter: it reliably detects the tested model\-generated changes and preserves the agreed controls, but can both over\-reject benign inconsistencies and over\-accept subtler errors involving entity alignment or temporal semantics\. Larger adjudicated studies, including a corresponding PubMedQA evaluation, are needed before treating a check acceptance as conclusive proof of validity\.相似文章
BenchEvolver: 基于解决方案进化的前沿任务合成
BenchEvolver 是一个进化框架,能够自动从现有编程问题中生成更难的题目,创建保持有效性和多样性的挑战性基准,同时支持模型自我改进和提升训练性能。
Hybrid Open-Ended Tri-Evolution 打造更好的深度研究者
本文提出混合开放式三方进化(HOTE)框架,该框架使用混合模式强化学习协同进化提议者、求解者和评判者,用于深度研究任务,以8B模型实现了超越更大静态模型的最优结果。
重新审视智能体框架演进的评估
本文重新评估了 LLM 智能体自动框架演进的方法论,指出其收益可能源于额外的测试时搜索而非改进的框架设计,并且在相同基准上的评估存在过拟合风险。实验表明,框架演进并不始终优于更简单的测试时扩展方法。
Evo-Bench:语言模型能否改进智能体工具框架?
介绍 Evo-Bench,这是首个用于评估语言模型在搜索、办公和通用智能体领域中内在工具框架进化能力的基准,表明顶级模型取得了显著进步,但在办公工作流上仍面临挑战。
PACEvolve++:提升进化搜索代理的测试时学习能力
本文介绍了 PACEvolve++,这是一种强化学习框架,通过将假设生成与执行解耦,提高了进化搜索代理在测试时的策略适应能力。