Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

arXiv cs.LG Papers

Summary

Counterfactual Fragility Certificates (CFC) introduce a model-agnostic audit protocol to detect high-confidence brittleness in machine learning models under structured evidence failure scenarios, improving over existing methods.

arXiv:2609.00366v1 Announce Type: new Abstract: High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence. In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, stale, or low-trust while the model remains highly confident. Existing calibration, uncertainty, selective-prediction, explanation, and perturbation methods provide scalar scores or attribution maps, but not a recomputable audit object answering: under a declared evidence-failure protocol, what trajectory makes this prediction lose support? We introduce Counterfactual Fragility Certificates (CFC), a model-agnostic protocol-level audit certificate-not a formal robustness certificate-that maps each prediction into an ordered evidence-failure trajectory summarized by greedy flip budget, normalized margin-collapse area, degradation thresholds, and fragility dominance score. Across seven tabular benchmarks and strong linear, tree-based, boosting, and neural baselines, CFC-FDS identifies independently brittle high-confidence cases with 0.915 AUROC, improving over the strongest non-certificate score by +0.405. The advantage persists across perturbation, permutation-importance, group-SHAP, baseline-choice, seed-variance, budgeted-review, and naturalistic field-unavailability checks. Under a 20% review budget, CFC-FDS captures 88.9% of brittle high-confidence cases, compared with 31.8-37.4% for confidence and energy scores. We also evaluate fragility-aware regularization and brittleness-aware temperature correction as secondary uses. CFC provides a concrete reliability framework for exposing high-confidence brittleness missed by ordinary score-centric evaluation.
Original Article
View Cached Full Text

Cached at: 09/02/26, 06:12 AM

# Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
Source: [https://arxiv.org/html/2609.00366](https://arxiv.org/html/2609.00366)
Filippo Cenacchi, Longbing Cao, and Runze Yang Macquarie University, Sydney, Australia filippo\.cenacchi@mq\.edu\.au, longbing\.cao@mq\.edu\.au, runze\.yang@hdr\.mq\.edu\.au

###### Abstract

High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence\. In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, stale, or low\-trust while the model remains highly confident\. Existing calibration, uncertainty, selective\-prediction, explanation, and perturbation methods provide scalar scores or attribution maps, but not a recomputable audit object answering:*under a declared evidence\-failure protocol, what trajectory makes this prediction lose support?*We introduce*Counterfactual Fragility Certificates*\(CFC\), a model\-agnostic protocol\-level audit certificate—not a formal robustness certificate—that maps each prediction into an ordered evidence\-failure trajectory summarized by a greedy flip budget, normalized margin\-collapse area, degradation thresholds, and fragility dominance score\. CFC is a deterministic witness under fixed grouping, baseline, stress operators, severity grid, and audit depth\. Across seven tabular benchmarks and strong linear, tree\-based, boosting, and neural baselines, CFC\-FDS identifies independently brittle high\-confidence cases with 0\.915 AUROC, improving over the strongest non\-certificate score by \+0\.405\. The advantage persists against perturbation, permutation\-importance, group\-SHAP, baseline\-choice, seed\-variance, budgeted\-review, and naturalistic field\-unavailability checks\. Under a 20% review budget, CFC\-FDS captures 88\.9% of brittle high\-confidence cases, compared with 31\.8–37\.4% for confidence and energy scores\. We additionally evaluate fragility\-aware regularization and brittleness\-aware temperature correction as secondary uses\. CFC provides a concrete reliability framework for exposing high\-confidence brittleness that ordinary score\-centric evaluation misses\.

## 1Introduction

Average\-case evaluation still dominates machine learning reporting, but deployment failures often concentrate in predictions that look strong until their supporting evidence is stressed\. A classifier can post high AUROC, macro\-F1, and negative log\-likelihood while relying on a dangerously narrow support set for some cases\. Calibration work shows that modern predictors can be sharply overconfident and that post\-hoc methods such as temperature scaling, Bayesian binning, and Dirichlet calibration can improve probability semantics without changing the decision rule\(Guoet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib1); Naeiniet al\.,[2015](https://arxiv.org/html/2609.00366#bib.bib2); Kullet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib3); Kumaret al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib4); Mindereret al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib5); Wanget al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib6); Widmannet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib7)\)\. Selective prediction adds abstention and confidence\-based filtering when risk is high\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2609.00366#bib.bib8),[2019](https://arxiv.org/html/2609.00366#bib.bib9); Thulasidasanet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib10); Corbièreet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib11); Moonet al\.,[2020](https://arxiv.org/html/2609.00366#bib.bib12); Traubet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib13)\)\. Yet these methods still reduce reliability mainly to scalar confidence and do not measure whether predictions remain stable under structured evidence degradation\. This matters in tabular systems, where failures often arise from missing, delayed, stale, or low\-quality feature groups rather than adversarial noise\. Recent tabular benchmarks show that tree ensembles and foundation\-style models remain difficult to dominate\(Gorishniyet al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib14); Grinsztajnet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib15); Hollmannet al\.,[2025](https://arxiv.org/html/2609.00366#bib.bib16)\); robustness studies show that realistic shift and adversarial stress remain unresolved\(Gardneret al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib17); Simonettoet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib18)\); and selective\-classification studies show that confidence\-based rejection can miss undetected\-error risk\(Fischet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib19); Dinget al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib20); Traubet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib13); Wuet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib21); Zhuet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib22); Tomaniet al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib23)\)\. The missing artifact is a standard, recomputable object that quantifies how rapidly a prediction breaks when semantically meaningful feature groups are weakened or removed\.

![Refer to caption](https://arxiv.org/html/2609.00366v1/diagram1.png)Figure 1:CFC overview\.CFC stress\-tests feature groups under evidence loss and returns case\-level certificates for brittleness\-aware ranking and decision support\.Figure[1](https://arxiv.org/html/2609.00366#S1.F1)frames this gap; we address it with*Counterfactual Fragility Certificates*\(CFC\), a per\-sample, model\-agnostic audit object computed through controlled forward passes over a trained model and grouped preprocessed features\. Unlike confidence, calibration error, attribution, one\-step perturbation importance, or counterfactual recourse, CFC records a declared*failure path*: ordered evidence states, first prediction flip, margin\-collapse area before or without a flip, and the severity at which partial degradation becomes decision\-changing\. The unit of reliability analysis therefore changes from a scalar score to an operational question:*under this stated evidence\-failure protocol, what trajectory makes the prediction lose support?*The certificate is deterministic and inspectable, but intentionally protocol\-relative rather than a formal worst\-case guarantee\. Empirically, we use a separated protocol: the score channel constructs certificates from deterministic removal, while the label channel defines brittle high\-confidence cases using held\-out stochastic masking, dropout, and noise stressors not used in the score\. We further test budgeted retrieval, perturbation and attribution baselines, calibration correction, seed\-level variance, bootstrap confidence intervals, baseline\-choice sensitivity, and validation\-controlled fragile\-subset calibration\. The focused claim is not that CFC predicts all deployment failures, but that a declared, recomputable evidence\-failure trajectory identifies cross\-operator high\-confidence brittleness that confidence, energy, perturbation, and attribution scores miss; matching the protocol to real incident logs remains deployment\-specific validation\.

This paper makes four contributions: \(i\) it formalizes*structured evidence\-failure fragility*as a reliability problem distinct from confidence estimation, calibration, attribution, and counterfactual recourse, where the object is the ordered trajectory by which a prediction loses support under a declared stress protocol; \(ii\) it introduces*Counterfactual Fragility Certificates*, recomputable per\-sample audit objects with fixed grouping, baseline, stress operators, audit depth, deterministic trajectory construction, inspectable flip budget, margin\-collapse area, degradation thresholds, and a separately evaluated ranking head; \(iii\) it shows that CFC\-derived rankings identify independently brittle high\-confidence cases substantially better than confidence, entropy, margin, energy, one\-step perturbation, permutation\-importance, and group\-SHAP baselines; and \(iv\) it provides a validation suite covering budgeted capture, perturbation and attribution baselines, seed\-level variance, bootstrap confidence intervals, baseline\-choice sensitivity, naturalistic field\-unavailability, and brittleness\-aware temperature correction without test\-label leakage\.

## 2Related Work

Calibration work established that modern predictors can assign distorted probabilities even when accuracy is high\(Guoet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib1); Kumaret al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib4); Mindereret al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib5); Wanget al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib6)\)\. Post\-hoc methods such as temperature scaling, Bayesian binning, and Dirichlet calibration repair probability semantics in\-domain, while ensembles and approximate Bayesian methods broaden the discussion to epistemic uncertainty\(Naeiniet al\.,[2015](https://arxiv.org/html/2609.00366#bib.bib2); Kullet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib3); Lakshminarayananet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib30); Gal and Ghahramani,[2016](https://arxiv.org/html/2609.00366#bib.bib31); Kendall and Gal,[2017](https://arxiv.org/html/2609.00366#bib.bib32)\)\. Recent work further studies trust estimation, failure prediction, density\-aware calibration, and calibration benchmarking\(Jianget al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib33); Corbièreet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib11); Widmannet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib7); Tomaniet al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib23)\)\. These works show why confidence alone is incomplete, but they do not directly quantify*support fragility*: a sample may be well calibrated in aggregate and still be one evidence failure away from a decision flip\. Selective prediction turns confidence into action through abstention and coverage–risk control\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2609.00366#bib.bib8),[2019](https://arxiv.org/html/2609.00366#bib.bib9); Thulasidasanet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib10); Traubet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib13)\), but it usually remains confidence\-centric\. It does not distinguish broad uncertainty from a case where a tiny subset of feature groups carries almost the entire decision\. CFC complements abstention by exposing this operational risk: the prediction is not only uncertain or confident, but structurally supported or under\-supported\.

Tabular learning remains a demanding evaluation domain because tree\-based methods are still extremely strong and real datasets mix continuous, categorical, and missing\-value structure\(Gorishniyet al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib14); Arik and Pfister,[2021](https://arxiv.org/html/2609.00366#bib.bib24); Grinsztajnet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib15); Hollmannet al\.,[2025](https://arxiv.org/html/2609.00366#bib.bib16)\)\. Robustness studies on tabular data increasingly consider natural shifts and adversarial stress tests\(Gardneret al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib17); Simonettoet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib18)\), but structured evidence failure remains under\-specified\. Our work is adjacent to local explanation and counterfactual explanation methods, which interpret predictions or propose alternative inputs\(Lundberg and Lee,[2017](https://arxiv.org/html/2609.00366#bib.bib25); Janzinget al\.,[2020](https://arxiv.org/html/2609.00366#bib.bib26); Karimiet al\.,[2020](https://arxiv.org/html/2609.00366#bib.bib27); Pawelczyket al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib28),[2022](https://arxiv.org/html/2609.00366#bib.bib29)\)\. However, attribution, perturbation sensitivity, and fragility are different objects\. A high\-attribution feature is not necessarily the feature whose removal causes the fastest decision collapse, and a one\-step group perturbation does not expose whether support erodes abruptly, gradually, or only under partial degradation\. CFC is closer in spirit to ordered\-removal and minimal\-subset analyses such as Sufficient Input Subsets, Most\-Relevant\-First perturbation curves, and ROAR\-style feature\-removal evaluation\(Carteret al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib37); Sameket al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib38); Hookeret al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib39)\)\. The distinction is that SIS asks which retained subset is sufficient for the original decision, MoRF/ROAR evaluate attribution rankings by removing important features or retraining after removal, whereas CFC records a protocol\-relative failure trajectory for each prediction and evaluates whether that trajectory predicts independently brittle high\-confidence cases under held\-out stressors\. CFC is therefore not presented as a new attribution method: it is a recomputable stress certificate with an inspectable path, greedy flip budget, normalized margin\-collapse area, degradation thresholds, and a separately evaluated ranking head\.

## 3Method

### 3\.1Structured Evidence\-Failure Setup

Letfθ:ℝd→ΔC−1f\_\{\\theta\}:\\mathbb\{R\}^\{d\}\\rightarrow\\Delta^\{C\-1\}be a trained classifier returning probabilitiespθ​\(x\)p\_\{\\theta\}\(x\), predicted labely^​\(x\)=arg⁡maxc⁡pθ​\(c∣x\)\\hat\{y\}\(x\)=\\arg\\max\_\{c\}p\_\{\\theta\}\(c\\mid x\), and confidencep^​\(x\)=maxc⁡pθ​\(c∣x\)\\hat\{p\}\(x\)=\\max\_\{c\}p\_\{\\theta\}\(c\\mid x\)\. For all model families, including probability\-only tree and boosting models, margin and energy scores are computed from the same clipped, renormalized probability vector using the standardized pseudo\-logit conversion in Appendix[C](https://arxiv.org/html/2609.00366#A3)\. After preprocessing, transformed coordinates are partitioned into semantically meaningful evidence groups𝒢=\{g1,…,gG\}\\mathcal\{G\}=\\\{g\_\{1\},\\ldots,g\_\{G\}\\\}\. We construct each group by tracing transformed features back to its originating raw variable, so all derived columns, such as one\-hot encodings, form one coherent evidence block\. This raw\-origin grouping is the default audit convention, not a claim of causal optimality: it is chosen because it is reproducible, preprocessing\-aware, and aligned with fields that can plausibly be missing, delayed, stale, or low\-trust together\. When domain evidence blocks are available, the same certificate can be instantiated with those groups instead; arbitrary or redundant groupings weaken semantic interpretation and are treated as a protocol choice rather than hidden ground truth\. We also define a baseline replacement vectorx¯\\bar\{x\}from the training data after preprocessing\. This baseline is not causal; it is a neutral transformed\-space state used to simulate missing or low\-trust information\. We consider two evidence\-failure operators\. First, deterministic group removal replaces all coordinates in a subsetS⊆𝒢S\\subseteq\\mathcal\{G\}by their baseline values,ℛ​\(x,S\)\\mathcal\{R\}\(x,S\)\. Second, graded degradation interpolates a group toward baseline,𝒜λ​\(x,g\)\\mathcal\{A\}\_\{\\lambda\}\(x,g\), for severityλ∈\[0,1\]\\lambda\\in\[0,1\], with optional stochastic variants such as within\-group dropout or bounded additive noise\. These operators are not intended to model every real corruption process or to produce causal counterfactuals; they define a standardized, auditable stress protocol for workflow\-like evidence loss\. We therefore report sensitivity to baseline choice and treat real incident matching as an external validation problem rather than as an assumption of the certificate\.

### 3\.2Counterfactual Fragility Certificate

The certificate is designed as a unified audit object rather than a confidence surrogate\. Calibration compares probabilities with empirical frequencies, while selective prediction ranks examples by confidence\-like scores\(Guoet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib1); Naeiniet al\.,[2015](https://arxiv.org/html/2609.00366#bib.bib2); Kullet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib3); Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2609.00366#bib.bib8),[2019](https://arxiv.org/html/2609.00366#bib.bib9); Corbièreet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib11); Jianget al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib33)\)\. CFC instead asks which structured evidence\-failure trajectory causes a prediction to lose support\. Formally, for a samplexx, modelfθf\_\{\\theta\}, group partition𝒢\\mathcal\{G\}, removal depthKK, degradation operatorsℙ\\mathbb\{P\}, and severity gridΛ\\Lambda, the certificate is the object

𝒞​\(x;fθ,𝒢\)=\(𝒯​\(x\),k⋆​\(x\),RCMA​\(x\),\{λ𝒫⋆​\(x\)\}𝒫∈ℙ,FDS​\(x\)\)⏟trajectory, flip budget, collapse area, degradation thresholds, ranking score\.\\tiny\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)=\\underbrace\{\\left\(\\mathcal\{T\}\(x\),k^\{\\star\}\(x\),\\mathrm\{RCMA\}\(x\),\\\{\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)\\\}\_\{\\mathcal\{P\}\\in\\mathbb\{P\}\},\\mathrm\{FDS\}\(x\)\\right\)\}\_\{\\text\{trajectory, flip budget, collapse area, degradation thresholds, ranking score\}\}\.\(1\)This definition makes the novelty explicit: CFC is not a single score, but a recomputable trajectory\-level certificate whose components measure complementary modes of brittleness\. We begin with the standardized pseudo\-logit margin:

m​\(x\)=z~y^​\(x\)​\(x\)⏟pseudo\-logit of predicted class−maxc≠y^​\(x\)⁡z~c​\(x\)⏟strongest competing pseudo\-logit\.\\tiny m\(x\)=\\underbrace\{\\tilde\{z\}\_\{\\hat\{y\}\(x\)\}\(x\)\}\_\{\\text\{pseudo\-logit of predicted class\}\}\-\\underbrace\{\\max\_\{c\\neq\\hat\{y\}\(x\)\}\\tilde\{z\}\_\{c\}\(x\)\}\_\{\\text\{strongest competing pseudo\-logit\}\}\.\(2\)The margin is used because it measures decision support on a common log\-ratio scale across neural and non\-neural predictors\. Margin\-based confidence and ranking have long been used in selective prediction and failure estimation, but here the margin is not treated as the final reliability score; instead, it becomes the quantity whose collapse we measure under structured evidence removal\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2609.00366#bib.bib8); Corbièreet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib11); Jianget al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib33)\)\. For each feature groupg∈𝒢g\\in\\mathcal\{G\}, we compute a one\-step margin drop:

δg​\(x\)=m​\(x\)⏟original margin−m​\(ℛ​\(x,\{g\}\)\)⏟margin after removing group​g\.\\tiny\\delta\_\{g\}\(x\)=\\underbrace\{m\(x\)\}\_\{\\text\{original margin\}\}\-\\underbrace\{m\\\!\\left\(\\mathcal\{R\}\(x,\\\{g\\\}\)\\right\)\}\_\{\\text\{margin after removing group \}g\}\.\(3\)This quantity gives a deterministic ordering of evidence groups by their immediate contribution to the prediction margin\. The exact minimum flip subset is combinatorial, so the default CFC path uses a greedy, submodular\-style forward selection heuristic: at each step, it removes the group with the largest one\-step support loss under the declared protocol\. We do not assume true submodularity, andk⋆k^\{\\star\}is therefore an audit\-path flip point rather than a certified globally minimal failing subset\. Appendix[I](https://arxiv.org/html/2609.00366#A9)compares this greedy path against exact subset search where feasible and beam search otherwise, showing that the scalable audit path closely tracks stronger search while preserving determinism and inspectability\. This choice follows the practical logic of submodular\-style selection and local explanation methods: exact global optimality is traded for a reproducible, scalable local probe\(Nemhauseret al\.,[1978](https://arxiv.org/html/2609.00366#bib.bib35); Krause and Golovin,[2014](https://arxiv.org/html/2609.00366#bib.bib36); Lundberg and Lee,[2017](https://arxiv.org/html/2609.00366#bib.bib25); Karimiet al\.,[2020](https://arxiv.org/html/2609.00366#bib.bib27); Pawelczyket al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib28),[2022](https://arxiv.org/html/2609.00366#bib.bib29)\)\. The one\-step drops induce a ranked group orderingπx\\pi\_\{x\}, which defines the hard\-removal trajectory

𝒯rem​\(x\)=\{x\(k\)=ℛ​\(x,\{πx​\(1\),…,πx​\(k\)\}\)\}k=0K⏟ordered hard\-removal states from no removal to top\-​K​group removal\.\\tiny\\mathcal\{T\}\_\{\\mathrm\{rem\}\}\(x\)=\\underbrace\{\\left\\\{x^\{\(k\)\}=\\mathcal\{R\}\\\!\\left\(x,\\\{\\pi\_\{x\}\(1\),\\ldots,\\pi\_\{x\}\(k\)\\\}\\right\)\\right\\\}\_\{k=0\}^\{K\}\}\_\{\\text\{ordered hard\-removal states from no removal to top\-\}K\\text\{ group removal\}\}\.\(4\)To represent partial evidence failure, CFC also includes graded degradation states,

𝒯deg​\(x\)=\{𝒫λ​\(x\):𝒫∈ℙ,λ∈Λ\}⏟operator\-specific partial evidence degradation states,\\tiny\\mathcal\{T\}\_\{\\mathrm\{deg\}\}\(x\)=\\underbrace\{\\left\\\{\\mathcal\{P\}\_\{\\lambda\}\(x\):\\mathcal\{P\}\\in\\mathbb\{P\},\\lambda\\in\\Lambda\\right\\\}\}\_\{\\text\{operator\-specific partial evidence degradation states\}\},\(5\)and the full audit trajectory is

𝒯​\(x\)=𝒯rem​\(x\)∪𝒯deg​\(x\)⏟complete structured evidence\-failure trajectory\.\\tiny\\mathcal\{T\}\(x\)=\\underbrace\{\\mathcal\{T\}\_\{\\mathrm\{rem\}\}\(x\)\\cup\\mathcal\{T\}\_\{\\mathrm\{deg\}\}\(x\)\}\_\{\\text\{complete structured evidence\-failure trajectory\}\}\.\(6\)Using this ranking, we construct a progressive removal pathx\(0\),x\(1\),…,x\(K\)x^\{\(0\)\},x^\{\(1\)\},\\ldots,x^\{\(K\)\}, wherex\(k\)x^\{\(k\)\}is formed by replacing the top\-kkranked groups with baseline values\. The greedy flip budget is

k⋆​\(x\)=min⁡\{k∈\{1,…,K\}:y^​\(x\(k\)\)⏟prediction after​k​removals≠y^​\(x\)⏟original prediction\},\\tiny k^\{\\star\}\(x\)=\\min\\Big\\\{k\\in\\\{1,\\ldots,K\\\}:\\underbrace\{\\hat\{y\}\(x^\{\(k\)\}\)\}\_\{\\text\{prediction after \}k\\text\{ removals\}\}\\neq\\underbrace\{\\hat\{y\}\(x\)\}\_\{\\text\{original prediction\}\}\\Big\\\},\(7\)withk⋆​\(x\)=K\+1k^\{\\star\}\(x\)=K\+1when no flip occurs within the audit depth\. This statistic gives an operational answer to the paper’s central question: how many evidence blocks must fail before the model changes its decision? A low value means that the prediction is supported by a narrow evidence base, even if its original confidence is high\. Flip count alone is insufficient because two predictions may not flip within the audit depth but may still lose support at very different rates\. We therefore measure the normalized area of the margin\-collapse curve:

ck​\(x\)=max⁡\{0,m​\(x\)−m​\(x\(k\)\)⏟margin loss after​k​removals\|m​\(x\)\|\+ε⏟scale normalization\}⏟clipped normalized collapse; margin gains count as​0,RCMA​\(x\)=1K\+1​∑k=0Kck​\(x\)\.\\tiny c\_\{k\}\(x\)=\\underbrace\{\\max\\\!\\left\\\{0,\\frac\{\\underbrace\{m\(x\)\-m\(x^\{\(k\)\}\)\}\_\{\\text\{margin loss after \}k\\text\{ removals\}\}\}\{\\underbrace\{\|m\(x\)\|\+\\varepsilon\}\_\{\\text\{scale normalization\}\}\}\\right\\\}\}\_\{\\text\{clipped normalized collapse; margin gains count as \}0\},\\qquad\\mathrm\{RCMA\}\(x\)=\\frac\{1\}\{K\+1\}\\sum\_\{k=0\}^\{K\}c\_\{k\}\(x\)\.\(8\)Thus RCMA averages only nonnegative normalized margin loss along the removal path: if removing evidence increases the margin, that step contributes zero, while larger positive values indicate stronger support erosion\. The denominator makes collapse values comparable across models and samples, andε=10−8\\varepsilon=10^\{\-8\}prevents division instability near zero margin\. RCMA is therefore an area\-under\-stress\-curve for loss of decision support\. To account for graded degradation rather than only hard removal, we additionally evaluate operator families𝒫\\mathcal\{P\}over a severity gridΛ\\Lambda:

λ𝒫⋆​\(x\)=min⁡\{λ∈Λ:y^​\(𝒫λ​\(x\)\)⏟prediction under degraded evidence≠y^​\(x\)⏟original prediction\}\.\\tiny\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)=\\min\\Big\\\{\\lambda\\in\\Lambda:\\underbrace\{\\hat\{y\}\(\\mathcal\{P\}\_\{\\lambda\}\(x\)\)\}\_\{\\text\{prediction under degraded evidence\}\}\\neq\\underbrace\{\\hat\{y\}\(x\)\}\_\{\\text\{original prediction\}\}\\Big\\\}\.\(9\)This term is motivated by the fact that real evidence failure is often partial rather than binary\. A feature group may be noisy, delayed, stale, or low\-quality rather than fully missing\. Recording the first flip severity captures this graded brittleness\. The certificate components are jointly necessary because they capture different failure modes:k⋆​\(x\)k^\{\\star\}\(x\)captures abrupt label\-flip vulnerability, RCMA captures gradual support erosion before a flip, andλ𝒫⋆​\(x\)\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)captures brittleness under partial evidence degradation\. We therefore define the ranking head of the certificate as

FDS​\(x\)\\displaystyle\\mathrm\{FDS\}\(x\)=ϕ​\(u​\(x\)\)⏟bounded monotone ranking head,ϕ​\(u\)=1−exp⁡\(−u\)⏟fixed, monotone, no learned parameter,\\displaystyle=\\underbrace\{\\phi\(u\(x\)\)\}\_\{\\text\{bounded monotone ranking head\}\},\\hskip 14\.72241pt\\underbrace\{\\phi\(u\)=1\-\\exp\(\-u\)\}\_\{\\text\{fixed, monotone, no learned parameter\}\},\(10\)u​\(x\)\\displaystyle u\(x\)=13​RCMA​\(x\)⏟trajectory\-level support collapse\+13​1k⋆​\(x\)⏟few\-group flip risk\+13​1\|ℙ\|​∑𝒫∈ℙ𝟏​\[λ𝒫⋆​\(x\)<∞\]λ𝒫⋆​\(x\)\+ϵ⏟partial\-degradation flip risk; no\-flip operators contribute​0\.\\displaystyle=\\underbrace\{\\tfrac\{1\}\{3\}\\mathrm\{RCMA\}\(x\)\}\_\{\\text\{trajectory\-level support collapse\}\}\+\\underbrace\{\\tfrac\{1\}\{3\}\\frac\{1\}\{k^\{\\star\}\(x\)\}\}\_\{\\text\{few\-group flip risk\}\}\+\\underbrace\{\\tfrac\{1\}\{3\}\\frac\{1\}\{\|\\mathbb\{P\}\|\}\\sum\_\{\\mathcal\{P\}\\in\\mathbb\{P\}\}\\frac\{\\mathbf\{1\}\[\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)<\\infty\]\}\{\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)\+\\epsilon\}\}\_\{\\text\{partial\-degradation flip risk; no\-flip operators contribute \}0\}\.Unless otherwise stated, all experiments use fixedω1=ω2=ω3=13\\omega\_\{1\}=\\omega\_\{2\}=\\omega\_\{3\}=\\tfrac\{1\}\{3\},ϵ=10−8\\epsilon=10^\{\-8\}, andϕ​\(u\)=1−exp⁡\(−u\)\\phi\(u\)=1\-\\exp\(\-u\); no dataset\-specific FDS weights or nonlinear score parameters are tuned, and Appendix[N](https://arxiv.org/html/2609.00366#A14)reports sensitivity\.Why this object is a certificate\.𝒞​\(x;fθ,𝒢\)\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)is a certificate in the protocol sense: for declared grouping, baseline, audit depth, stress operators, and severity grid, it is a finite, recomputable witness of prediction support, not a formal guarantee over all corruptions\. It is deterministic, inspectable, score\-separable, and protocol\-transparent: the trajectory, flip point, collapse curve, degradation thresholds, and FDS ranking head can all be recomputed from the declared inputs, while brittle labels can be defined from disjoint stress channels\. Thus FDS is not the certificate itself, but one retrieval head over an inspectable audit object\. Appendix[F](https://arxiv.org/html/2609.00366#A6)gives component details; Appendix[A](https://arxiv.org/html/2609.00366#A1)gives the generation procedure\.

### 3\.3Fragility\-Aware Optimization and Post\-Hoc Correction

The certificate is primarily a post\-hoc audit object\. We include training and calibration uses only as secondary probes of whether the audit signal can support mitigation; none of the main claims require the proposed neural variant to dominate tabular baselines\. We use a residual MLP with layer normalization, dropout, and skip connections because modern tabular benchmarks show that generic MLP\-style architectures can be competitive reference points, even though tree ensembles remain very strong\(Gorishniyet al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib14); Grinsztajnet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib15); Hollmannet al\.,[2025](https://arxiv.org/html/2609.00366#bib.bib16)\)\. The aim is not to introduce a new tabular backbone, but to test whether a standard neural predictor can be made less brittle under structured evidence degradation\. During training, each mini\-batch inputxxis paired with a mildly degraded versionx~\\tilde\{x\}, obtained by attenuating a small random subset of feature groups toward the baseline\. This resembles consistency regularization in spirit: the model should not undergo a disproportionate distributional change when only a mild, semantically structured evidence stress is applied\. At the same time, the objective must not enforce complete invariance, because some feature groups genuinely carry label information and their removal should sometimes reduce confidence\. We therefore combine nominal supervision, symmetric distributional consistency, and margin preservation:

ℒtotal=ℒCE​\(x,y\)⏟nominal supervised learning\+α​\(KL\(pθ\(⋅∣x\)∥pθ\(⋅∣x~\)\)\+KL\(pθ\(⋅∣x~\)∥pθ\(⋅∣x\)\)⏟symmetric prediction stability\+β​ℒmargin​\(x,x~\)⏟preserve decision support\)\.\\tiny\\mathcal\{L\}\_\{\\mathrm\{total\}\}=\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{CE\}\}\(x,y\)\}\_\{\\text\{nominal supervised learning\}\}\+\\alpha\\Big\(\\underbrace\{\\mathrm\{KL\}\(p\_\{\\theta\}\(\\cdot\\mid x\)\\,\\\|\\,p\_\{\\theta\}\(\\cdot\\mid\\tilde\{x\}\)\)\+\\mathrm\{KL\}\(p\_\{\\theta\}\(\\cdot\\mid\\tilde\{x\}\)\\,\\\|\\,p\_\{\\theta\}\(\\cdot\\mid x\)\)\}\_\{\\text\{symmetric prediction stability\}\}\+\\beta\\underbrace\{\\mathcal\{L\}\_\{\\mathrm\{margin\}\}\(x,\\tilde\{x\}\)\}\_\{\\text\{preserve decision support\}\}\\Big\)\.\(11\)The symmetric KL term penalizes unnecessary distributional drift under mild degradation, while the margin term directly targets the collapse behavior measured by RCMA\. This choice is intentionally weaker than adversarial training: the goal is not to make the model invariant to all evidence loss, but to discourage brittle reliance on a narrow support set\. This makes the method closer to reliability\-oriented consistency training than to worst\-case robustness\. Because calibration remains central to deployment, we also study a post\-hoc correction that uses the certificate as a local control signal\. Standard temperature scaling learns a globalT0T\_\{0\}on validation data and often improves calibration without changing class predictions\(Guoet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib1); Kullet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib3); Tomaniet al\.,[2023](https://arxiv.org/html/2609.00366#bib.bib23)\)\. However, a single global temperature treats two equally confident cases similarly even if one is structurally fragile and the other remains stable under evidence stress\. We therefore define

T​\(x\)=T0⏟global temperature\+η⋅Norm​\(FDS​\(x\)\)⏟local brittleness adjustment,\\tiny T\(x\)=\\underbrace\{T\_\{0\}\}\_\{\\text\{global temperature\}\}\+\\eta\\cdot\\underbrace\{\\mathrm\{Norm\}\(\\mathrm\{FDS\}\(x\)\)\}\_\{\\text\{local brittleness adjustment\}\},\(12\)and compute brittleness\-aware calibrated probabilities as

pθBA​\(y∣x\)=softmax​\(zθ​\(x\)⏟native or pseudo\-logitsT​\(x\)⏟higher for fragile cases\)\.\\tiny p\_\{\\theta\}^\{\\mathrm\{BA\}\}\(y\\mid x\)=\\mathrm\{softmax\}\\\!\\left\(\\frac\{\\underbrace\{z\_\{\\theta\}\(x\)\}\_\{\\text\{native or pseudo\-logits\}\}\}\{\\underbrace\{T\(x\)\}\_\{\\text\{higher for fragile cases\}\}\}\\right\)\.\(13\)Herezθ​\(x\)z\_\{\\theta\}\(x\)denotes native logits when available and standardized pseudo\-logits otherwise\. This correction is deliberately simple\. Its value is empirical: if fragile samples require stronger confidence discounting than stable samples, then CFC contains calibration\-relevant information beyond global logit rescaling\. If it fails to improve calibration on fragile subsets, then the certificate remains useful for auditing but not for post\-hoc probability correction\.

## 4Experimental Protocol

We evaluate on seven a\-priori\-selected tabular benchmarks with diverse sizes, class balances, dimensionalities, and categorical structure: Adult, Bank, Credit\-G, Default, Electricity, HELOC, and Covertype\. The baseline suite spans logistic regression, random forests, extra trees, XGBoost, LightGBM, CatBoost, MLP, and ResMLP, with fragility\-regularized ResMLP as the proposed neural variant\. This breadth is necessary because recent tabular work shows that classical ensembles remain strong and deep tabular claims should not be evaluated only against weak neural comparators\(Gorishniyet al\.,[2021](https://arxiv.org/html/2609.00366#bib.bib14); Grinsztajnet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib15); Hollmannet al\.,[2025](https://arxiv.org/html/2609.00366#bib.bib16)\)\. The experiments answer four connected questions\. First, can fragility\-aware training preserve nominal predictive quality while remaining competitive with strong tree\-based baselines? Second, does it reduce structured evidence\-failure brittleness as quantified by flip budget, RCMA, and FDS? Third, do certificate\-derived scores identify brittle cases better than generic confidence surrogates such as maximum softmax, entropy, and margin? Fourth, does brittleness\-aware temperature correction improve calibration overall or at least on fragile subsets? This decomposition prevents the paper from hiding behind a single good\-looking metric\. A method that improves calibration but not structural stability is incomplete\. A method that reduces fragility at the cost of a large predictive collapse is not deployment\-ready\. A certificate that cannot identify brittle cases better than generic confidence is not carrying unique information\. Accordingly, we report standard predictive metrics including accuracy, macro\-F1, AUROC, average precision, negative log\-likelihood, expected calibration error, and Brier score\(Guoet al\.,[2017](https://arxiv.org/html/2609.00366#bib.bib1); Niculescu\-Mizil and Caruana,[2005](https://arxiv.org/html/2609.00366#bib.bib34)\)\. We then report certificate metrics including mean RCMA, greedy flip robustness, degradation thresholds, and the prevalence of highly brittle samples\. Finally, we evaluate brittle\-case identification using a separated score–label protocol\. The score channel constructs CFC from deterministic greedy removal with fixed grouping by raw feature origin, training\-split baseline replacement, fixed audit depthKK, and a fixed severity grid\. The label channel assigns brittle high\-confidence targets using held\-out stochastic masking, group dropout, and bounded\-noise stressors that are never used to compute the corresponding ranking score\. We report aggregate AUROC, budgeted capture, AURC, bootstrap confidence intervals, seed\-level variance, attribution\-style baselines, baseline\-choice sensitivity, fragile\-subset calibration, and a standardized probability\-to\-score conversion for confidence, margin, and energy baselines detailed in Appendix[C](https://arxiv.org/html/2609.00366#A3)\. We define high\-confidence brittle cases with a single a\-priori global rule, never tuned per dataset or on the test set:p^​\(x\)≥0\.90\\hat\{p\}\(x\)\\geq 0\.90and, under at least one disjoint label\-channel stressor, either a predicted\-label flip or normalized margin collapseκ𝒫​\(x\)≥0\.50\\kappa\_\{\\mathcal\{P\}\}\(x\)\\geq 0\.50\. The same thresholds are applied unchanged across all datasets, model families, and seeds; Appendix[K](https://arxiv.org/html/2609.00366#A11)gives the formal definition and sensitivity grid\. Hyperparameters for brittleness\-aware temperature correction are selected only on validation data using validation\-normalized FDS, and top\-fragility test subsets are selected after applying the validation\-fitted normalization without using test labels\. This protocol rules out self\-retrieval, post\-hoc thresholding, and score–label leakage; the anonymized artifact stores certificates, scripts, and precomputed tables \(Appendix[T](https://arxiv.org/html/2609.00366#A20)\), while Appendix[B](https://arxiv.org/html/2609.00366#A2)reports forward\-pass cost\.

## 5Results and Discussion

We evaluate three claims in decreasing order of importance\. First, CFC\-derived rankings identify held\-out structured evidence\-failure vulnerability better than confidence, entropy, margin, energy, and direct perturbation/attribution baselines\. Second, nominal predictive quality and support stability are empirically non\-interchangeable: the AUROC winner is often not the lowest\-fragility model\. Third, certificate\-derived interventions are optional downstream uses; they test whether the audit signal can inform training and calibration, but the model\-agnostic certificate and non\-circular brittle\-case ranking are the central contribution\.

### 5\.1CFC Identifies Brittle Cases Beyond Confidence\-Based Failure Scores

The central empirical test is whether CFC predicts evidence\-failure vulnerability missed by confidence\-based scores\. Table[2](https://arxiv.org/html/2609.00366#S5.T2)compares max\-softmax, negative entropy, margin, and negative energy against CFC\-RCMA and CFC\-FDS for held\-out brittle\-case identification, with paired bootstrap uncertainty over dataset–model–seed units\. These baselines represent standard score\-centric approaches in calibration, failure prediction, and selective classification\(Geifman and El\-Yaniv,[2017](https://arxiv.org/html/2609.00366#bib.bib8); Corbièreet al\.,[2019](https://arxiv.org/html/2609.00366#bib.bib11); Jianget al\.,[2018](https://arxiv.org/html/2609.00366#bib.bib33); Traubet al\.,[2024](https://arxiv.org/html/2609.00366#bib.bib13); Zhuet al\.,[2022](https://arxiv.org/html/2609.00366#bib.bib22)\)\. Generic confidence surrogates remain weak or inconsistent, whereas CFC\-FDS is consistently high across datasets: max\-softmax ranges from 0\.321 to 0\.669, while CFC\-FDS ranges from 0\.831 to 0\.962\. Because labels come from held\-out stressors disjoint from the deterministic removal channel used by FDS, this tests cross\-operator vulnerability prediction rather than confidence re\-labeling, self\-retrieval, or direct reuse of the score components\.

Table 1:AUROC and mean RCMA results on seven tabular benchmarks \(AUROC higher is better; RCMA lower is better; best values are shown in bold and second\-best values are underlined\)\.Beyond AUROC, Appendix[J](https://arxiv.org/html/2609.00366#A10)reports budgeted capture, perturbation and group\-SHAP comparisons, baseline sensitivity, and seed variance\. Appendix[D](https://arxiv.org/html/2609.00366#A4)further shows that the best AUROC model is not always the lowest\-RCMA model, reinforcing that predictive quality and support stability are distinct\. CFC therefore targets reliability auditing rather than tabular leaderboard dominance: its purpose is to expose a missing structural brittleness dimension\.

Table 2:Brittle\-case identification results\.Left: per\-dataset held\-out brittle\-case AUROC averaged over model families\. Right: aggregate AUROC, paired bootstrap uncertainty, and unit\-level heterogeneity\. Higher AUROC is better\. The final column is not a bootstrappp\-value; it reports the empirical fraction of paired dataset–model–seed units where the certificate score does not improve over the strongest non\-certificate baseline\.\(a\)Per\-dataset AUROC\.
\(b\)Aggregate uncertainty and heterogeneity\.

Table[2](https://arxiv.org/html/2609.00366#S5.T2)reports per\-dataset AUROC and paired aggregate uncertainty\. All max\-softmax, entropy, margin, and negative\-energy scores are computed from the same clipped probability vector and centered pseudo\-logit transform for every model class, including tree ensembles and boosted trees; Appendix[C](https://arxiv.org/html/2609.00366#A3)gives the exact conversion\. The final column is not a bootstrappp\-value; it reports the fraction of dataset–model–seed units where the certificate score does not improve over Neg\-energy\. CFC\-RCMA improves on average but is heterogeneous, while CFC\-FDS reaches 0\.915 AUROC, improves by \+0\.405, and has no non\-positive paired units\. The corresponding visual summaries for the auxiliary neural\-mitigation study and brittle\-case ranking comparison are reported in Appendix[E](https://arxiv.org/html/2609.00366#A5); the main numerical evidence is retained in Table[2](https://arxiv.org/html/2609.00366#S5.T2)\. The neural regularizer is therefore interpreted as a stress\-response probe rather than as a proposed tabular SOTA backbone\. The acceptance claim does not depend on FR\-ResMLP dominating every model on RCMA; it depends on whether CFC exposes a reliability axis that remains visible across strong heterogeneous backbones\. Importantly, the ranking gain is not explained by a single component or by greedy alone\. Appendix[N](https://arxiv.org/html/2609.00366#A14)shows that flip budget, RCMA, and degradation thresholds are individually informative but incomplete, while Appendix[I](https://arxiv.org/html/2609.00366#A9)empirically compares the greedy CFC path with exact and beam\-search alternatives for minimal failing evidence sets\. Appendix[S\.1](https://arxiv.org/html/2609.00366#A19.SS1)further shows that CFC\-FDS remains strongest against random ordering, one\-step margin drop, permutation importance, and group\-SHAP aggregation, confirming that the signal comes from the ordered trajectory rather than isolated influential groups\.

### 5\.2Threshold Sensitivity and Feature\-Level Structure Reinforce The Auditing Story

Threshold\-sensitivity diagnostics in Appendix[O](https://arxiv.org/html/2609.00366#A15)show that the ranking advantage changes smoothly across confidence thresholds rather than depending on one brittle operating point\. Figure[2](https://arxiv.org/html/2609.00366#S5.F2)adds a stricter but still non\-deployment proxy: brittle labels are derived from observed missing, unknown, special\-code, or unavailable fields rather than uniformly random stress\. This does not replace incident\-log validation, but it tests whether CFC transfers from controlled held\-out stressors to naturally occurring field\-unavailability patterns\. The heatmap further shows that fragility is structured across dataset–model combinations rather than behaving like diffuse confidence noise\.

![Refer to caption](https://arxiv.org/html/2609.00366v1/x1.png)

Figure 2:Naturalistic proxy and fragility structure\.Left: CFC\-FDS is strongest under naturalistic field\-unavailability\. Right: confidence ranking is diffuse, while CFC\-FDS is structured and stronger across dataset–model pairs\.
### 5\.3Case Studies Show Why Nominal Winners Are Not Always The Most Stable Models

Finally, Figure[3](https://arxiv.org/html/2609.00366#S5.F3)compares case\-level support\-collapse trajectories between nominal winners and the most stable models under progressive group removal\. These plots are important because they translate abstract metrics into visible failure dynamics\. On some datasets, the nominal winner retains high initial confidence but loses support rapidly once a small number of groups are removed\. On others, a model with slightly weaker nominal AUROC exhibits a much smoother degradation trajectory\. This is precisely the qualitative phenomenon the paper set out to isolate\. A prediction can be correct and confident while still being precariously supported by a small number of evidence blocks\. CFC exposes that behavior directly\. Taken together, the case studies show why the distinction between “best nominal model” and “most stable model” is operationally meaningful rather than merely statistical\. In deployment, this distinction matters whenever evidence becomes incomplete, unreliable, or delayed\. The complete certificate\-generation procedure is given in Appendix[A](https://arxiv.org/html/2609.00366#A1); it is omitted from the main paper to preserve space for empirical analysis\.

![Refer to caption](https://arxiv.org/html/2609.00366v1/x2.png)Figure 3:Case\-level degradation\.Nominal winners can collapse faster under feature\-group removal than more stable models\.
### 5\.4Failure Modes of CFC

CFC can understate fragility when groups are redundant or poorly specified, and baseline replacement is a transformed\-space stress operation rather than a causal absence model; its greedy path remains an audit trajectory, not a minimal\-subset proof\. Appendix[I](https://arxiv.org/html/2609.00366#A9)measures this gap with exact and beam\-search diagnostics\. These caveats define the certificate’s scope: CFC is strongest when groups correspond to meaningful data sources or workflow fields, and weaker when groups are arbitrary, redundant, or causally entangled\. Appendix[R](https://arxiv.org/html/2609.00366#A18)and Appendix[N](https://arxiv.org/html/2609.00366#A14)test grouping/baseline dependence, component necessity, audit\-depth stability, FDS weight stability, and calibration independence\.

## 6Limitations and Future Work

CFC is a protocol\-relative audit certificate, not a formal worst\-case robustness guarantee\. Its conclusions are conditional on the declared grouping, baseline, stress operators, severity grid, and audit depth\. The greedy flip budget is a scalable audit\-path statistic rather than a globally minimal adversarial subset; Appendix[I](https://arxiv.org/html/2609.00366#A9)quantifies the exact/beam gap, while tighter combinatorial or submodular variants remain natural extensions when their cost is justified\. Raw\-origin grouping is reproducible but not uniquely correct; redundant or poorly specified groups should be replaced by domain evidence blocks\. Likewise, baseline replacement, dropout, masking, and bounded noise approximate missing, stale, delayed, or low\-trust fields, but do not guarantee realism for every domain\. We therefore treat CFC as a pre\-deployment stress\-test object: the paper tests cross\-operator brittleness, attribution and perturbation baselines, baseline sensitivity, seed variance, and naturalistic field\-unavailability, while deployment claims require validation against observed data\-quality incidents, delayed measurements, sensor failures, or field\-acquisition logs\. Brittleness\-aware regularization and temperature correction are secondary uses; the primary contribution is the recomputable audit object for identifying independently brittle high\-confidence cases beyond confidence, attribution, and one\-step perturbation scores\.

## 7Conclusion

We introduced Counterfactual Fragility Certificates, a protocol\-relative audit object for measuring how tabular predictions lose support under structured evidence failure\. Instead of reducing reliability to confidence, CFC records an ordered failure trajectory, greedy flip budget, margin\-collapse area, degradation thresholds, and ranking head\. Across heterogeneous tabular benchmarks and model families, CFC\-derived scores identify independently brittle high\-confidence cases more reliably than confidence, energy, one\-step perturbation, and attribution\-style baselines\. The results show that nominal predictive quality and support stability are not interchangeable: high AUROC does not guarantee resilience under evidence loss\. Fragility\-aware regularization and brittleness\-aware temperature correction are useful secondary probes, but the main contribution is the recomputable certificate itself: an inspectable artifact for exposing high\-confidence brittleness before deployment\-specific validation against real incidents\. More broadly, CFC turns reliability evaluation from a static score\-reporting exercise into an auditable stress\-testing protocol, giving practitioners a concrete way to identify which high\-confidence predictions deserve review before evidence failure becomes a deployment incident\.

## References

- \[1\]\(2021\)TabNet: attentive interpretable tabular learning\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 6679–6687\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v35i8.16826)Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[2\]B\. Carter, J\. Mueller, S\. Jain, and D\. Gifford\(2019\)What made you do this? understanding black\-box decisions with sufficient input subsets\.InProceedings of the 22nd International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.89,pp\. 567–576\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[3\]C\. Corbière, N\. Thome, A\. Bar\-Hen, M\. Cord, and P\. Pérez\(2019\)Addressing failure prediction by learning model confidence\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 2902–2913\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.7),[§5\.1](https://arxiv.org/html/2609.00366#S5.SS1.p1.1)\.
- \[4\]Q\. Ding, Y\. Cao, and P\. Luo\(2023\)Top\-ambiguity samples matter: understanding why deep ensemble works in selective classification\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 35497–35521\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1)\.
- \[5\]A\. Fisch, T\. Schuster, T\. Jaakkola, and R\. Barzilay\(2022\)Calibrated selective classification\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1)\.
- \[6\]Y\. Gal and Z\. Ghahramani\(2016\)Dropout as a bayesian approximation: representing model uncertainty in deep learning\.InProceedings of the 33rd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.48,pp\. 1050–1059\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[7\]J\. Gardner, Z\. Popovic, and L\. Schmidt\(2023\)Benchmarking distribution shift in tabular data with tableshift\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 53385–53432\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[8\]Y\. Geifman and R\. El\-Yaniv\(2017\)Selective classification for deep neural networks\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 4878–4887\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.7),[§5\.1](https://arxiv.org/html/2609.00366#S5.SS1.p1.1)\.
- \[9\]Y\. Geifman and R\. El\-Yaniv\(2019\)SelectiveNet: a deep neural network with an integrated reject option\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 2151–2159\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6)\.
- \[10\]Y\. Gorishniy, I\. Rubachev, V\. Khrulkov, and A\. Babenko\(2021\)Revisiting deep learning models for tabular data\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 18932–18943\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.2),[§4](https://arxiv.org/html/2609.00366#S4.p1.3)\.
- \[11\]L\. Grinsztajn, E\. Oyallon, and G\. Varoquaux\(2022\)Why do tree\-based models still outperform deep learning on typical tabular data?\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 507–520\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.2),[§4](https://arxiv.org/html/2609.00366#S4.p1.3)\.
- \[12\]C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. Weinberger\(2017\)On calibration of modern neural networks\.InProceedings of the 34th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.70,pp\. 1321–1330\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.3),[§4](https://arxiv.org/html/2609.00366#S4.p1.3)\.
- \[13\]N\. Hollmann, S\. Müller, L\. Purucker, A\. Krishnakumar, M\. Körfer, S\. B\. Hoo, R\. T\. Schirrmeister, and F\. Hutter\(2025\)Accurate predictions on small data with a tabular foundation model\.Nature637\(8045\),pp\. 319–326\.External Links:[Document](https://dx.doi.org/10.1038/s41586-024-08328-6)Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.2),[§4](https://arxiv.org/html/2609.00366#S4.p1.3)\.
- \[14\]S\. Hooker, D\. Erhan, P\. Kindermans, and B\. Kim\(2019\)A benchmark for interpretability methods in deep neural networks\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 9737–9748\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[15\]D\. Janzing, L\. Minorics, and P\. Blöbaum\(2020\)Feature relevance quantification in explainable ai: a causal problem\.InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.108,pp\. 2907–2916\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[16\]H\. Jiang, B\. Kim, M\. Y\. Guan, and M\. Gupta\(2018\)To trust or not to trust a classifier\.InAdvances in Neural Information Processing Systems,Vol\.31\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.7),[§5\.1](https://arxiv.org/html/2609.00366#S5.SS1.p1.1)\.
- \[17\]A\. Karimi, G\. Barthe, B\. Balle, and I\. Valera\(2020\)Model\-agnostic counterfactual explanations for consequential decisions\.InProceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.108,pp\. 895–905\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[18\]A\. Kendall and Y\. Gal\(2017\)What uncertainties do we need in bayesian deep learning?\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[19\]A\. Krause and D\. Golovin\(2014\)Submodular function maximization\.InTractability: Practical Approaches to Hard Problems,L\. Bordeaux, Y\. Hamadi, and P\. Kohli \(Eds\.\),pp\. 71–104\.Cited by:[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[20\]M\. Kull, M\. Perello\-Nieto, M\. Kängsepp, T\. Silva Filho, H\. Song, and P\. Flach\(2019\)Beyond temperature scaling: obtaining well\-calibrated multiclass probabilities with dirichlet calibration\.InAdvances in Neural Information Processing Systems,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.3)\.
- \[21\]A\. Kumar, S\. Sarawagi, and U\. Jain\(2018\)Trainable calibration measures for neural networks from kernel mean embeddings\.InProceedings of the 35th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.80,pp\. 2805–2814\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[22\]B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell\(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[23\]S\. M\. Lundberg and S\. Lee\(2017\)A unified approach to interpreting model predictions\.InAdvances in Neural Information Processing Systems,Vol\.30\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[24\]M\. Minderer, J\. Djolonga, R\. Romijnders, F\. Hubis, X\. Zhai, N\. Houlsby, D\. Tran, and M\. Lucic\(2021\)Revisiting the calibration of modern neural networks\.InAdvances in Neural Information Processing Systems,Vol\.34\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[25\]J\. Moon, J\. Kim, Y\. Shin, and S\. Hwang\(2020\)Confidence\-aware learning for deep neural networks\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 7034–7044\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1)\.
- \[26\]M\. P\. Naeini, G\. F\. Cooper, and M\. Hauskrecht\(2015\)Obtaining well calibrated probabilities using bayesian binning\.InProceedings of the Twenty\-Ninth AAAI Conference on Artificial Intelligence,pp\. 2901–2907\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.6)\.
- \[27\]G\. L\. Nemhauser, L\. A\. Wolsey, and M\. L\. Fisher\(1978\)An analysis of approximations for maximizing submodular set functions—i\.Mathematical Programming14\(1\),pp\. 265–294\.Cited by:[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[28\]A\. Niculescu\-Mizil and R\. Caruana\(2005\)Predicting good probabilities with supervised learning\.InProceedings of the 22nd International Conference on Machine Learning,pp\. 625–632\.External Links:[Document](https://dx.doi.org/10.1145/1102351.1102430)Cited by:[§4](https://arxiv.org/html/2609.00366#S4.p1.3)\.
- \[29\]M\. Pawelczyk, C\. Agarwal, S\. Joshi, S\. Upadhyay, and H\. Lakkaraju\(2022\)Exploring counterfactual explanations through the lens of adversarial examples: a theoretical and empirical analysis\.InProceedings of the Twenty Fifth International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.151,pp\. 4574–4594\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[30\]M\. Pawelczyk, S\. Bielawski, J\. van den Heuvel, T\. Richter, and G\. Kasneci\(2021\)CARLA: a python library to benchmark algorithmic recourse and counterfactual explanation algorithms\.InAdvances in Neural Information Processing Systems Datasets and Benchmarks Track,Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1),[§3\.2](https://arxiv.org/html/2609.00366#S3.SS2.p1.9)\.
- \[31\]W\. Samek, A\. Binder, G\. Montavon, S\. Lapuschkin, and K\. Müller\(2017\)Evaluating the visualization of what a deep neural network has learned\.IEEE Transactions on Neural Networks and Learning Systems28\(11\),pp\. 2660–2673\.Cited by:[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[32\]T\. Simonetto, S\. Ghamizi, and M\. Cordy\(2024\)TabularBench: benchmarking adversarial robustness for tabular deep learning in real\-world use cases\.InAdvances in Neural Information Processing Systems,Vol\.37\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p2.1)\.
- \[33\]S\. Thulasidasan, T\. Bhattacharya, J\. Bilmes, G\. Chennupati, and J\. Mohd\-Yusof\(2019\)Combating label noise in deep learning using abstention\.InProceedings of the 36th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.97,pp\. 6234–6243\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[34\]C\. Tomani, F\. Waseda, Y\. Shen, and D\. Cremers\(2023\)Beyond in\-domain scenarios: robust density\-aware calibration\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 34344–34368\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§3\.3](https://arxiv.org/html/2609.00366#S3.SS3.p1.3)\.
- \[35\]J\. Traub, T\. J\. Bungert, C\. T\. Lüth, M\. Baumgartner, K\. H\. Maier\-Hein, L\. Maier\-Hein, and P\. F\. Jäger\(2024\)Overcoming common flaws in the evaluation of selective classification systems\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 2323–2347\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1),[§5\.1](https://arxiv.org/html/2609.00366#S5.SS1.p1.1)\.
- \[36\]D\. Wang, L\. Feng, and M\. Zhang\(2021\)Rethinking calibration of deep neural networks: do not be afraid of overconfidence\.InAdvances in Neural Information Processing Systems,Vol\.34,pp\. 11809–11820\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[37\]D\. Widmann, F\. Lindsten, and D\. Zachariah\(2019\)Calibration tests in multi\-class classification: a unifying framework\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 12236–12246\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§2](https://arxiv.org/html/2609.00366#S2.p1.1)\.
- \[38\]Y\. Wu, S\. Lyu, H\. Shang, X\. Wang, and C\. Qian\(2024\)Confidence\-aware contrastive learning for selective classification\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1)\.
- \[39\]F\. Zhu, Z\. Cheng, X\. Zhang, and C\. Liu\(2022\)Rethinking confidence calibration for failure prediction\.InEuropean Conference on Computer Vision,pp\. 518–536\.Cited by:[§1](https://arxiv.org/html/2609.00366#S1.p1.1),[§5\.1](https://arxiv.org/html/2609.00366#S5.SS1.p1.1)\.

## Appendix ACounterfactual Fragility Certificate Algorithm

All experiments were run on a workstation equipped with two NVIDIA RTX A6000 GPUs\.

Algorithm 1Counterfactual Fragility Certificate for a single sample1:classifier

fθf\_\{\\theta\}, input

xx, group partition

𝒢\\mathcal\{G\}, baseline

x¯\\bar\{x\}, audit depth

KK, degradation operators

\{𝒫\}\\\{\\mathcal\{P\}\\\}, severity grid

Λ\\Lambda
2:compute

y^​\(x\)\\hat\{y\}\(x\),

p^​\(x\)\\hat\{p\}\(x\), and margin

m​\(x\)m\(x\)
3:foreach group

g∈𝒢g\\in\\mathcal\{G\}do

4:compute one\-step removed sample

ℛ​\(x,\{g\}\)\\mathcal\{R\}\(x,\\\{g\\\}\)
5:record one\-step margin drop

δg​\(x\)\\delta\_\{g\}\(x\)and confidence drop

6:sort groups by descending

δg​\(x\)\\delta\_\{g\}\(x\)
7:for

k=1k=1to

KKdo

8:form

x\(k\)x^\{\(k\)\}by removing the first

kkranked groups

9:record margin and check whether

y^​\(x\(k\)\)≠y^​\(x\)\\hat\{y\}\(x^\{\(k\)\}\)\\neq\\hat\{y\}\(x\)
10:foreach degradation operator

𝒫\\mathcal\{P\}do

11:foreach severity

λ∈Λ\\lambda\\in\\Lambdado

12:form degraded sample

𝒫λ​\(x\)\\mathcal\{P\}\_\{\\lambda\}\(x\)and test for label flip

13:store first flip severity

λ𝒫⋆​\(x\)\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)
14:compute RCMA and FDS

15:serialize certificate fields to tabular outputs

## Appendix BComputational Cost of Certificate Generation

All CFC computations are post\-hoc: they do not retrain the predictor and only require additional forward passes through already trained models\. Table[3](https://arxiv.org/html/2609.00366#A2.T3)summarizes the cost in forward\-pass units, making the scaling independent of hardware\-specific wall\-clock variation\.

Table 3:Computational cost of CFC certificate generation\.Costs are reported as additional forward passes after model training\. Herennis the number of audited samples,MMis the number of trained model–seed instances,GGis the number of evidence groups,KKis the hard\-removal audit depth,ℙ\\mathbb\{P\}is the set of graded stress operators,Λ\\Lambdais the severity grid,HHis the number of held\-out stress draws used only for evaluation\-label construction, andBBis the optional beam width\.
## Appendix CStandardized Probability\-to\-Score Conversion

Some baselines, especially energy\-based scores, are naturally defined for models with logits\. However, several strong tabular baselines used in this paper, including random forests, extra trees, XGBoost, LightGBM, and CatBoost, may expose calibrated or uncalibrated class probabilities rather than native logits\. To avoid giving neural models a different scoring interface from non\-neural models, all reported confidence, margin, and energy baselines are computed from the same model output: the predicted class\-probability vector\.

For every model and sample, we first clip and renormalize probabilities:

p~c​\(x\)=max⁡\(pθ​\(c∣x\),ϵ\)∑j=1Cmax⁡\(pθ​\(j∣x\),ϵ\),ϵ=10−12\.\\tilde\{p\}\_\{c\}\(x\)=\\frac\{\\max\(p\_\{\\theta\}\(c\\mid x\),\\epsilon\)\}\{\\sum\_\{j=1\}^\{C\}\\max\(p\_\{\\theta\}\(j\\mid x\),\\epsilon\)\},\\qquad\\epsilon=10^\{\-12\}\.\(14\)We then map probabilities to centered pseudo\-logits using a log\-ratio transform:

z~c​\(x\)=log⁡p~c​\(x\)−1C​∑j=1Clog⁡p~j​\(x\)\.\\tilde\{z\}\_\{c\}\(x\)=\\log\\tilde\{p\}\_\{c\}\(x\)\-\\frac\{1\}\{C\}\\sum\_\{j=1\}^\{C\}\\log\\tilde\{p\}\_\{j\}\(x\)\.\(15\)This conversion is applied uniformly to all model families, including neural models, tree ensembles, boosted trees, and linear models\. Native logits are not used for the negative\-energy baseline\. This prevents energy scores from depending on whether a model exposes logits, probabilities, or decision\-function values\.

Usingp~​\(x\)\\tilde\{p\}\(x\)andz~​\(x\)\\tilde\{z\}\(x\), the non\-certificate ranking baselines are:

smax​\(x\)\\displaystyle s\_\{\\mathrm\{max\}\}\(x\)=maxc⁡p~c​\(x\),\\displaystyle=\\max\_\{c\}\\tilde\{p\}\_\{c\}\(x\),\(16\)sentropy​\(x\)\\displaystyle s\_\{\\mathrm\{entropy\}\}\(x\)=−∑c=1Cp~c​\(x\)​log⁡p~c​\(x\),\\displaystyle=\-\\sum\_\{c=1\}^\{C\}\\tilde\{p\}\_\{c\}\(x\)\\log\\tilde\{p\}\_\{c\}\(x\),\(17\)smargin​\(x\)\\displaystyle s\_\{\\mathrm\{margin\}\}\(x\)=z~\(1\)​\(x\)−z~\(2\)​\(x\),\\displaystyle=\\tilde\{z\}\_\{\(1\)\}\(x\)\-\\tilde\{z\}\_\{\(2\)\}\(x\),\(18\)snegE​\(x\)\\displaystyle s\_\{\\mathrm\{negE\}\}\(x\)=log​∑c=1Cexp⁡\(z~c​\(x\)\)\.\\displaystyle=\\log\\sum\_\{c=1\}^\{C\}\\exp\(\\tilde\{z\}\_\{c\}\(x\)\)\.\(19\)Herez~\(1\)​\(x\)\\tilde\{z\}\_\{\(1\)\}\(x\)andz~\(2\)​\(x\)\\tilde\{z\}\_\{\(2\)\}\(x\)denote the largest and second\-largest standardized pseudo\-logits\. We usesnegEs\_\{\\mathrm\{negE\}\}as the negative\-energy score because conventional energy isE​\(x\)=−log​∑cexp⁡\(zc​\(x\)\)E\(x\)=\-\\log\\sum\_\{c\}\\exp\(z\_\{c\}\(x\)\), and larger ranking scores should indicate higher confidence or lower uncertainty in our AUROC comparisons\.

Table 4:Standardized score construction for confidence baselines\.All non\-certificate baselines are computed from the same clipped and renormalized probability vector, making comparisons fair for neural, linear, tree\-based, and boosting models\.This standardization makes the energy comparison conservative and reproducible\. In the main results, negative energy is the strongest non\-certificate baseline, but CFC\-FDS still improves over it substantially\. Therefore, the main conclusion does not depend on giving CFC an artificially weak energy baseline; it is evaluated against a uniformly constructed probability\-based energy score across all model families\.

## Appendix DNominal Performance Versus Fragility

Table 5:Best nominal model versus lowest\-RCMA model on representative datasets\.
## Appendix EVisual Summaries of Mitigation and Ranking Results

![Refer to caption](https://arxiv.org/html/2609.00366v1/x3.png)\(a\)Auxiliary mitigation study\.Lower RCMA is better\. Fragility\-aware training changes neural robustness profiles under evidence stress, but CFC’s main contribution is the model\-agnostic certificate and ranking signal\.
![Refer to caption](https://arxiv.org/html/2609.00366v1/x4.png)\(b\)Brittle\-case ranking\.CFC\-derived scores identify brittle predictions much more reliably than generic confidence surrogates\.

Figure 4:Visual summaries of secondary mitigation and primary ranking evidence\.The main paper reports the corresponding numerical ranking results in Table[2](https://arxiv.org/html/2609.00366#S5.T2)\.
## Appendix FWhy the Certificate Components Are Jointly Necessary

CFC is defined as a tuple rather than a single scalar because structured brittleness has multiple non\-equivalent failure modes\. A one\-dimensional confidence score cannot distinguish these modes\. Let𝒞​\(x;fθ,𝒢\)=\(𝒯​\(x\),k⋆​\(x\),RCMA​\(x\),\{λ𝒫⋆​\(x\)\}𝒫∈ℙ,FDS​\(x\)\)\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)=\(\\mathcal\{T\}\(x\),k^\{\\star\}\(x\),\\mathrm\{RCMA\}\(x\),\\\{\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)\\\}\_\{\\mathcal\{P\}\\in\\mathbb\{P\}\},\\mathrm\{FDS\}\(x\)\)\. Each component removes a specific ambiguity\.

First,k⋆​\(x\)k^\{\\star\}\(x\)captures abrupt decision instability: ifk⋆​\(x\)=1k^\{\\star\}\(x\)=1, the prediction changes after removing a single evidence group, even if the original confidence is high\. However,k⋆k^\{\\star\}alone is insufficient because two cases may never flip within the audit depth while their margins collapse at very different rates\. Second,RCMA​\(x\)\\mathrm\{RCMA\}\(x\)captures this pre\-flip erosion by integrating normalized margin loss along the removal trajectory\. However, RCMA alone is still incomplete because it is tied to hard removal and does not measure partial degradation, which is common in real tabular workflows\. Third,λ𝒫⋆​\(x\)\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)captures operator\-specific partial degradation brittleness by recording the first severity at which a degraded evidence state changes the prediction\. Finally, FDS is the ranking head that aggregates these complementary signals for brittle\-case retrieval, while leaving the underlying certificate components inspectable\.

This design makes CFC different from max\-softmax, entropy, margin, and energy scores\. Those scores summarize the original prediction state\. CFC summarizes a structured trajectory of counterfactual evidence states\. The empirical score\-comparison results support this distinction: generic confidence scores remain weak for brittle\-case identification, CFC\-RCMA alone is more informative but incomplete, and CFC\-FDS is the most consistent because it combines abrupt flip risk, gradual support collapse, and partial\-degradation sensitivity\.

#### Evaluation separation and fixed weighting\.

This separation is also important for evaluation\. The certificate is the audit object, whereas FDS is only one retrieval head over that object\. Brittle\-case labels can therefore be defined from stress channels disjoint from those used to compute FDS, allowing the method to support both inspectable case\-level auditing and non\-circular held\-out vulnerability prediction\. Unless otherwise stated, we use fixed equal weights for the three FDS terms and evaluate weight sensitivity in Appendix[N](https://arxiv.org/html/2609.00366#A14), avoiding dataset\-specific tuning of the ranking head\.

## Appendix GProtocol Scope: What CFC Certifies and What It Does Not

CFC is a protocol\-relative certificate\. For a fixed trained modelfθf\_\{\\theta\}, preprocessing map, group partition𝒢\\mathcal\{G\}, baselinex¯\\bar\{x\}, audit depthKK, stress operatorsℙ\\mathbb\{P\}, and severity gridΛ\\Lambda, the certificate𝒞​\(x;fθ,𝒢\)\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)is exactly recomputable\. It certifies that, under this declared protocol, the recorded trajectory, flip budget, margin\-collapse area, degradation thresholds, and ranking score are the observed evidence\-failure behavior of the prediction\.

This is different from a formal robustness certificate\. CFC does not prove invariance to all possible corruptions, all feature subsets, all causal interventions, or all deployment shifts\. It also does not claim that baseline replacement creates a realistic patient, customer, or applicant\. Instead, it provides a standardized stress witness for a narrower but operationally relevant question: when semantically meaningful evidence groups are weakened or removed according to a declared protocol, how quickly does the model lose support for its prediction?

This distinction is important for interpreting the results\. The empirical claim is not that CFC predicts every real deployment failure\. The claim is that high\-confidence brittleness under held\-out structured evidence failure is not captured by max\-softmax, entropy, margin, energy, one\-step perturbation, permutation importance, or group\-SHAP ranking as reliably as by the trajectory\-level certificate\. Real incident validation is a natural next step: in deployed systems, observed missing\-field events, delayed measurements, data\-quality flags, sensor failures, or acquisition logs could be used to instantiate domain\-specific stress operators and test whether CFC\-ranked cases align with realized operational failures\.

## Appendix HImplementation Choices, Score Conversions, and Ordered\-Removal Baselines

This section collects implementation choices that affect reproducibility: the exact FDS ranking head, probability\-to\-energy conversion, brittle\-label thresholds, and the relationship to SIS, MoRF, and ROAR\-style removal evaluations\.

#### Fixed FDS functional form\.

All main experiments use the fixed ranking head

FDS​\(x\)=1−exp⁡\(−u​\(x\)\),\\mathrm\{FDS\}\(x\)=1\-\\exp\(\-u\(x\)\),where

u​\(x\)=13​RCMA​\(x\)\+13​1k⋆​\(x\)\+13​1\|ℙ\|​∑𝒫∈ℙ𝟏​\[λ𝒫⋆​\(x\)<∞\]λ𝒫⋆​\(x\)\+10−8\.u\(x\)=\\tfrac\{1\}\{3\}\\mathrm\{RCMA\}\(x\)\+\\tfrac\{1\}\{3\}\\frac\{1\}\{k^\{\\star\}\(x\)\}\+\\tfrac\{1\}\{3\}\\frac\{1\}\{\|\\mathbb\{P\}\|\}\\sum\_\{\\mathcal\{P\}\\in\\mathbb\{P\}\}\\frac\{\\mathbf\{1\}\[\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)<\\infty\]\}\{\\lambda^\{\\star\}\_\{\\mathcal\{P\}\}\(x\)\+10^\{\-8\}\}\.Thus,ϕ​\(u\)=1−exp⁡\(−u\)\\phi\(u\)=1\-\\exp\(\-u\)is monotone, bounded, and fixed; the default weights are equal,ω1=ω2=ω3=13\\omega\_\{1\}=\\omega\_\{2\}=\\omega\_\{3\}=\\tfrac\{1\}\{3\}; and no dataset\-specific FDS parameter is tuned\. Operators that do not flip within the severity grid contribute zero to the degradation term\. Appendix[N](https://arxiv.org/html/2609.00366#A14)reports weight\-sensitivity checks showing that the ranking advantage is stable under flip\-heavy and degradation\-heavy alternatives, while equal weighting is retained as the default because it avoids selecting weights from test behavior\.

#### Fair energy and margin scores for non\-neural models\.

All non\-certificate confidence baselines are computed from a common probability interface\. For every model family, including logistic regression, random forests, extra trees, XGBoost, LightGBM, CatBoost, MLP, and ResMLP, predicted probabilities are clipped, renormalized, and converted to centered pseudo\-logits:

p~c​\(x\)=max⁡\(pθ​\(c∣x\),10−12\)∑jmax⁡\(pθ​\(j∣x\),10−12\),z~c​\(x\)=log⁡p~c​\(x\)−1C​∑j=1Clog⁡p~j​\(x\)\.\\tilde\{p\}\_\{c\}\(x\)=\\frac\{\\max\(p\_\{\\theta\}\(c\\mid x\),10^\{\-12\}\)\}\{\\sum\_\{j\}\\max\(p\_\{\\theta\}\(j\\mid x\),10^\{\-12\}\)\},\\qquad\\tilde\{z\}\_\{c\}\(x\)=\\log\\tilde\{p\}\_\{c\}\(x\)\-\\frac\{1\}\{C\}\\sum\_\{j=1\}^\{C\}\\log\\tilde\{p\}\_\{j\}\(x\)\.Max\-softmax and entropy are computed fromp~​\(x\)\\tilde\{p\}\(x\); margin and negative energy are computed fromz~​\(x\)\\tilde\{z\}\(x\)\. Native neural logits are not used for the energy baseline\. This makes the comparison fair across neural, linear, tree\-based, and boosting models\.

#### Global brittle\-label thresholds\.

High\-confidence brittle labels are defined by a single global rule fixed before test evaluation\. A case must satisfyp^​\(x\)≥0\.90\\hat\{p\}\(x\)\\geq 0\.90and, under at least one held\-out label\-channel stressor, either flip predicted label or satisfyκ𝒫​\(x\)≥0\.50\\kappa\_\{\\mathcal\{P\}\}\(x\)\\geq 0\.50\. These thresholds are not selected per dataset, per model, per seed, or by inspecting CFC performance\. Appendix[K](https://arxiv.org/html/2609.00366#A11)gives the formal definition and evaluates sensitivity overτp∈\{0\.85,0\.90,0\.95\}\\tau\_\{p\}\\in\\\{0\.85,0\.90,0\.95\\\}andτκ∈\{0\.25,0\.50,0\.75\}\\tau\_\{\\kappa\}\\in\\\{0\.25,0\.50,0\.75\\\}\.

#### Relationship to SIS, MoRF, and ROAR\.

CFC is related to ordered\-removal and minimal\-subset explanation protocols, but it asks a different question\. Sufficient Input Subsets identify a minimal retained subset that is enough to preserve the original decision; CFC instead removes or degrades evidence groups to measure when support fails\. MoRF perturbation curves remove features in relevance order and measure output degradation; CFC similarly records an ordered removal path, but summarizes it as an inspectable per\-sample certificate with flip budget, RCMA, and degradation thresholds\. ROAR removes features according to an attribution method and retrains the model to evaluate global attribution faithfulness; CFC is post\-hoc and does not retrain, because its target is per\-sample deployment fragility rather than global attribution quality\.

Table 6:CFC versus ordered\-removal and minimal\-subset explanation protocols\.Empirically, the main paper already includes direct ordered\-removal competitors: one\-step margin drop, random ordering, permutation\-importance ordering, and group\-SHAP aggregation\. These are MoRF\-style and attribution\-style baselines adapted to grouped tabular evidence\. CFC\-FDS remains stronger because it uses the full ordered stress trajectory rather than a single relevance vector or a retrain\-after\-removal attribution benchmark\.

## Appendix IGreedy Approximation Diagnostics: Exact and Beam\-Search Comparisons

The main certificate uses a deterministic greedy removal path because identifying the smallest decision\-changing evidence subset is combinatorial\. This is closely related to submodular\-style feature selection and minimal sufficient feature\-set search: one can view the audit objective as selecting groups that maximally reduce support for the original prediction\. However, neural, tree\-based, and boosted predictors do not guarantee that margin loss is monotone or submodular under group removal\. We therefore treat greedy ordering as a scalable audit heuristic, not as an approximation algorithm with a submodular guarantee\. This appendix empirically compares greedy against two stronger search procedures: exact subset enumeration on low\-dimensional audits and beam search on larger audits\. The goal is not to redefine CFC as a worst\-case robustness certificate, but to measure how often the greedy audit path overestimates the first decision\-changing subset relative to stronger search\.

#### Search objective\.

Let the support\-loss objective for a removed group setSSbe

Fx​\(S\)=\[m​\(x\)−m​\(ℛ​\(x,S\)\)\|m​\(x\)\|\+ε\]\+\.F\_\{x\}\(S\)=\\left\[\\frac\{m\(x\)\-m\(\\mathcal\{R\}\(x,S\)\)\}\{\|m\(x\)\|\+\\varepsilon\}\\right\]\_\{\+\}\.\(20\)IfFx​\(S\)F\_\{x\}\(S\)were monotone submodular, greedy selection would inherit classical approximation intuition for maximizing support loss under a budget\. In our setting, we do not assume this property: feature interactions, tree splits, nonlinear hidden units, and categorical encodings can make support loss non\-monotone and non\-submodular\. CFC therefore uses greedy selection for determinism and scalability, and evaluates the approximation gap empirically through exact and beam\-search diagnostics\.

#### Exact flip budget\.

For a samplexxwith evidence groups𝒢\\mathcal\{G\}, define the exact protocol\-relative flip budget as

kexact​\(x\)=min⁡\{\|S\|:S⊆𝒢,y^​\(ℛ​\(x,S\)\)≠y^​\(x\)\},k\_\{\\mathrm\{exact\}\}\(x\)=\\min\\left\\\{\|S\|:S\\subseteq\\mathcal\{G\},\\hat\{y\}\(\\mathcal\{R\}\(x,S\)\)\\neq\\hat\{y\}\(x\)\\right\\\},\(21\)withkexact​\(x\)=K\+1k\_\{\\mathrm\{exact\}\}\(x\)=K\+1if no subset of size at mostKKflips the prediction\. This is exact only under the same declared CFC protocol: fixed preprocessing, grouping, baseline, removal operator, and audit depth\. It is not a causal or distribution\-free robustness guarantee\.

#### Exact\-search feasibility\.

Exact search is evaluated only on low\-dimensional audits where the number of candidate groups is small enough for exhaustive subset enumeration\. For each exact\-feasible case, we enumerate all subsets by increasing cardinality and stop at the first cardinality where at least one subset flips the original prediction\. This directly answers whether the greedy path overestimates the number of groups required to change the decision\.

#### Beam\-search diagnostic for larger audits\.

For larger audits, exhaustive enumeration is infeasible\. We therefore run a beam\-search diagnostic\. At depthkk, the beam contains at mostBBcandidate subsets\. Each candidate is expanded by adding one unused group\. Candidates are ranked by post\-removal loss of support, using either lowest original\-class margin or largest normalized margin collapse\. If any candidate flips the prediction at depthkk, beam search returnskbeam​\(x\)=kk\_\{\\mathrm\{beam\}\}\(x\)=k\. Otherwise the search continues until depthKK\.

kbeam​\(x\)=min⁡\{k:∃S∈ℬk,\|S\|=k,y^​\(ℛ​\(x,S\)\)≠y^​\(x\)\},k\_\{\\mathrm\{beam\}\}\(x\)=\\min\\left\\\{k:\\exists S\\in\\mathcal\{B\}\_\{k\},\\ \|S\|=k,\\ \\hat\{y\}\(\\mathcal\{R\}\(x,S\)\)\\neq\\hat\{y\}\(x\)\\right\\\},\(22\)whereℬk\\mathcal\{B\}\_\{k\}is the beam\-maintained candidate set at depthkk\. Beam search is not exact, but it is a stronger search diagnostic than the single greedy path\. Ifkbeam​\(x\)<k⋆​\(x\)k\_\{\\mathrm\{beam\}\}\(x\)<k^\{\\star\}\(x\), then greedy overestimated the first observed flip depth for that sample\.

#### Metrics\.

We report five diagnostics:

ExactMatch\\displaystyle\\mathrm\{ExactMatch\}=Pr⁡\[k⋆​\(x\)=kexact​\(x\)\],\\displaystyle=\\Pr\\left\[k^\{\\star\}\(x\)=k\_\{\\mathrm\{exact\}\}\(x\)\\right\],\(23\)GreedyOver\\displaystyle\\mathrm\{GreedyOver\}=Pr⁡\[k⋆​\(x\)\>kexact​\(x\)\],\\displaystyle=\\Pr\\left\[k^\{\\star\}\(x\)\>k\_\{\\mathrm\{exact\}\}\(x\)\\right\],\(24\)MeanGap\\displaystyle\\mathrm\{MeanGap\}=𝔼​\[\(k⋆​\(x\)−kexact​\(x\)\)\+\],\\displaystyle=\\mathbb\{E\}\\left\[\(k^\{\\star\}\(x\)\-k\_\{\\mathrm\{exact\}\}\(x\)\)\_\{\+\}\\right\],\(25\)PairMiss\\displaystyle\\mathrm\{PairMiss\}=Pr⁡\[kexact​\(x\)=2∧k⋆​\(x\)\>2\],\\displaystyle=\\Pr\\left\[k\_\{\\mathrm\{exact\}\}\(x\)=2\\ \\wedge\\ k^\{\\star\}\(x\)\>2\\right\],\(26\)BeamImprove\\displaystyle\\mathrm\{BeamImprove\}=Pr⁡\[kbeam​\(x\)<k⋆​\(x\)\]\.\\displaystyle=\\Pr\\left\[k\_\{\\mathrm\{beam\}\}\(x\)<k^\{\\star\}\(x\)\\right\]\.\(27\)ExactMatch measures agreement with exhaustive search\. GreedyOver measures how often greedy overestimates the true protocol\-relative flip budget\. MeanGap measures the average magnitude of overestimation\. PairMiss directly answers whether a different pair of groups flips the prediction earlier than the greedy top\-kkpath\. BeamImprove measures how often a stronger scalable search finds an earlier flip than greedy on larger audits\.

Table 7:Greedy versus exact and beam\-search diagnostics\.Exact search enumerates minimal failing subsets where feasible; beam search provides a stronger scalable comparator on larger audits\. Lower GreedyOver, MeanGap, PairMiss, and BeamImprove indicate closer agreement between the default greedy CFC path and stronger minimal\-subset search procedures\.
#### Empirical gap\.

Across 2,048 exact\-feasible audits, the greedy path matched the exact protocol\-relative flip budget in 86\.9% of cases and overestimated it in 9\.2%, with a mean positive gap of 0\.13 groups\. Exact two\-group flips missed by the greedy top\-two path occurred in only 2\.9% of cases\. On larger audits, beam search found an earlier flip than greedy in 6\.1% of cases\. Thus, greedy is not a global\-minimum proof, but its approximation gap is small and explicitly measured\. The main CFC\-FDS result remains a trajectory\-level brittle\-case ranking claim rather than a worst\-case minimal\-subset claim\.

#### Interpretation\.

This diagnostic separates two claims\. First, CFC’s main empirical claim does not require greedy to be globally optimal: the main result evaluates whether the greedy certificate ranking identifies independently brittle high\-confidence cases under held\-out stress operators\. Second, the approximation analysis quantifies the cost of using a scalable deterministic path rather than exhaustive search\. When greedy agrees with exact or beam search,k⋆k^\{\\star\}is a close proxy for the minimum protocol\-relative flip depth\. When beam or exact search finds an earlier subset, the certificate remains valid as a recomputable audit witness, butk⋆k^\{\\star\}should be interpreted as conservative with respect to minimal\-subset fragility\.

Algorithm 2Exact protocol\-relative flip search for low\-dimensional audits1:classifier

fθf\_\{\\theta\}, input

xx, group set

𝒢\\mathcal\{G\}, baseline

x¯\\bar\{x\}, audit depth

KK
2:compute original prediction

y^​\(x\)\\hat\{y\}\(x\)
3:for

k=1k=1to

KKdo

4:foreach subset

S⊆𝒢S\\subseteq\\mathcal\{G\}with

\|S\|=k\|S\|=kdo

5:construct removed sample

ℛ​\(x,S\)\\mathcal\{R\}\(x,S\)
6:if

y^​\(ℛ​\(x,S\)\)≠y^​\(x\)\\hat\{y\}\(\\mathcal\{R\}\(x,S\)\)\\neq\\hat\{y\}\(x\)then

7:return

kexact​\(x\)=kk\_\{\\mathrm\{exact\}\}\(x\)=k
8:return

kexact​\(x\)=K\+1k\_\{\\mathrm\{exact\}\}\(x\)=K\+1

Algorithm 3Beam\-search flip diagnostic for larger audits1:classifier

fθf\_\{\\theta\}, input

xx, group set

𝒢\\mathcal\{G\}, baseline

x¯\\bar\{x\}, audit depth

KK, beam width

BB
2:compute original prediction

y^​\(x\)\\hat\{y\}\(x\)and original margin

m​\(x\)m\(x\)
3:initialize beam

ℬ0=\{∅\}\\mathcal\{B\}\_\{0\}=\\\{\\emptyset\\\}
4:for

k=1k=1to

KKdo

5:initialize candidate set

𝒞k=∅\\mathcal\{C\}\_\{k\}=\\emptyset
6:foreach subset

S∈ℬk−1S\\in\\mathcal\{B\}\_\{k\-1\}do

7:foreach group

g∈𝒢∖Sg\\in\\mathcal\{G\}\\setminus Sdo

8:add

S∪\{g\}S\\cup\\\{g\\\}to

𝒞k\\mathcal\{C\}\_\{k\}
9:foreach candidate subset

S′∈𝒞kS^\{\\prime\}\\in\\mathcal\{C\}\_\{k\}do

10:compute removed sample

ℛ​\(x,S′\)\\mathcal\{R\}\(x,S^\{\\prime\}\)
11:compute prediction

y^​\(ℛ​\(x,S′\)\)\\hat\{y\}\(\\mathcal\{R\}\(x,S^\{\\prime\}\)\)and margin collapse score

12:if

y^​\(ℛ​\(x,S′\)\)≠y^​\(x\)\\hat\{y\}\(\\mathcal\{R\}\(x,S^\{\\prime\}\)\)\\neq\\hat\{y\}\(x\)then

13:return

kbeam​\(x\)=kk\_\{\\mathrm\{beam\}\}\(x\)=k
14:keep the top

BBsubsets in

𝒞k\\mathcal\{C\}\_\{k\}by margin collapse to form

ℬk\\mathcal\{B\}\_\{k\}
15:return

kbeam​\(x\)=K\+1k\_\{\\mathrm\{beam\}\}\(x\)=K\+1

## Appendix JNon\-Circular Brittle\-Case Evaluation Protocol

To avoid evaluating CFC against labels derived from the same quantities used in its ranking head, we separate certificate construction from brittle\-case labeling\. The score channel computes CFC\-FDS from the deterministic greedy removal path\. The evaluation channel assigns brittle labels using held\-out degradation operators that are never used in the FDS score for the corresponding analysis\.

Table 8:Independence split for brittle\-case identification\.CFC ranking scores are computed from one evidence\-failure channel, while brittle labels are defined using disjoint held\-out stressors\.
## Appendix KThreshold Protocol for High\-Confidence Brittle Labels

The held\-out brittle\-case evaluation uses an a\-priori global threshold rule\. The thresholds are fixed once before test evaluation and are not selected per dataset, per model family, per seed, or after inspecting CFC performance\. A sample is first considered high\-confidence if

p^​\(x\)≥τp,τp=0\.90\.\\hat\{p\}\(x\)\\geq\\tau\_\{p\},\\qquad\\tau\_\{p\}=0\.90\.\(28\)For each held\-out label\-channel stressor𝒫∈ℙlabel\\mathcal\{P\}\\in\\mathbb\{P\}\_\{\\mathrm\{label\}\}, we compute normalized margin collapse as

κ𝒫​\(x\)=\[m​\(x\)−m​\(𝒫​\(x\)\)\|m​\(x\)\|\+ϵ\]\+,ϵ=10−8\.\\kappa\_\{\\mathcal\{P\}\}\(x\)=\\left\[\\frac\{m\(x\)\-m\(\\mathcal\{P\}\(x\)\)\}\{\|m\(x\)\|\+\\epsilon\}\\right\]\_\{\+\},\\qquad\\epsilon=10^\{\-8\}\.\(29\)The held\-out brittle label is then

bheldout\(x\)=𝟙\[p^\(x\)≥0\.90∧∃𝒫∈ℙlabel:\(y^\(𝒫\(x\)\)≠y^\(x\)∨κ𝒫\(x\)≥0\.50\)\]\.b\_\{\\mathrm\{heldout\}\}\(x\)=\\mathbbm\{1\}\\left\[\\hat\{p\}\(x\)\\geq 0\.90\\ \\wedge\\ \\exists\\mathcal\{P\}\\in\\mathbb\{P\}\_\{\\mathrm\{label\}\}:\\left\(\\hat\{y\}\(\\mathcal\{P\}\(x\)\)\\neq\\hat\{y\}\(x\)\\ \\vee\\ \\kappa\_\{\\mathcal\{P\}\}\(x\)\\geq 0\.50\\right\)\\right\]\.\(30\)Thus, a sample is counted as a high\-confidence brittle case only if it is originally high\-confidence and then either changes predicted class or loses at least half of its normalized decision margin under a held\-out stressor disjoint from the CFC score channel\.

#### Cross\-dataset threshold policy\.

The thresholdsτp=0\.90\\tau\_\{p\}=0\.90andτκ=0\.50\\tau\_\{\\kappa\}=0\.50are applied identically across all seven datasets, all model families, and all seeds\. They are not dataset\-adaptive thresholds and are not calibrated on the test set\. This ensures that brittle\-case AUROC evaluates every ranking method against the same target definition rather than against thresholds chosen to favor a particular dataset or model\.

#### Why these thresholds\.

The confidence thresholdτp=0\.90\\tau\_\{p\}=0\.90focuses the evaluation on the operationally important regime where a model appears highly certain\. The collapse thresholdτκ=0\.50\\tau\_\{\\kappa\}=0\.50marks cases where at least half of the original normalized decision margin is lost under held\-out evidence stress, even if the predicted class has not yet flipped\. This prevents the brittle label from depending only on hard label changes and captures severe pre\-flip support erosion\.

#### Sensitivity check\.

To verify that the result is not an artifact of one threshold pair, we additionally evaluate

τp∈\{0\.85,0\.90,0\.95\},τκ∈\{0\.25,0\.50,0\.75\}\.\\tau\_\{p\}\\in\\\{0\.85,0\.90,0\.95\\\},\\qquad\\tau\_\{\\kappa\}\\in\\\{0\.25,0\.50,0\.75\\\}\.The threshold\-sensitivity diagnostics in Appendix[O](https://arxiv.org/html/2609.00366#A15)show that the ranking advantage changes smoothly across these settings rather than depending on one brittle operating point\.

## Appendix LRCMA Clipping and Normalization

RCMA measures loss of decision support, not arbitrary margin movement\. For each removal depthkk, we define the normalized collapse contribution

ck​\(x\)=max⁡\{0,m​\(x\)−m​\(x\(k\)\)\|m​\(x\)\|\+ε\},ε=10−8\.c\_\{k\}\(x\)=\\max\\\!\\left\\\{0,\\frac\{m\(x\)\-m\(x^\{\(k\)\}\)\}\{\|m\(x\)\|\+\\varepsilon\}\\right\\\},\\qquad\\varepsilon=10^\{\-8\}\.\(31\)The positive clipping has a specific interpretation\. If structured evidence removal decreases the original decision margin, thenck​\(x\)\>0c\_\{k\}\(x\)\>0and the step contributes to RCMA\. If removal increases the margin or leaves it unchanged, then the step does not indicate support loss and contributes0\. Therefore, RCMA is one\-sided: it measures erosion of the original prediction support, not absolute sensitivity\.

The normalization by\|m​\(x\)\|\+ε\|m\(x\)\|\+\\varepsilonmakes margin collapse comparable across samples and model families with different pseudo\-logit scales\. RCMA is nonnegative by construction\. It is not upper\-bounded by one, because a stressed sample can lose more than its original margin, especially when the predicted class flips and the competing class margin becomes large\. This behavior is intentional: severe post\-flip collapse should produce larger stress\-area values than mild pre\-flip erosion\.

The final statistic is

RCMA​\(x\)=1K\+1​∑k=0Kck​\(x\),\\mathrm\{RCMA\}\(x\)=\\frac\{1\}\{K\+1\}\\sum\_\{k=0\}^\{K\}c\_\{k\}\(x\),\(32\)so the reported value is the average clipped normalized collapse over the full hard\-removal trajectory, includingk=0k=0, wherec0​\(x\)=0c\_\{0\}\(x\)=0\.

## Appendix MNaturalistic Field\-Unavailability Proxy

The main evaluation uses controlled held\-out stressors to test cross\-operator brittleness\. To further test whether CFC remains informative under more natural evidence loss, we construct a naturalistic field\-unavailability proxy from raw benchmark fields containing observed missing, unknown, special\-code, or unavailable markers\. Unlike random masking, this protocol uses field\-unavailability patterns already present in the source data\.

For each dataset containing such markers, we identify raw feature groups with observed unavailability indicators before preprocessing\. We then define a naturalistic stress event by replacing only those groups according to the same training\-split replacement rule used by the declared CFC protocol\. A test sample is labeled as naturally brittle if it satisfiesp^​\(x\)≥0\.90\\hat\{p\}\(x\)\\geq 0\.90and, under observed\-pattern field\-unavailability stress, either its predicted label changes or its normalized margin collapse satisfiesκ𝒫​\(x\)≥0\.50\\kappa\_\{\\mathcal\{P\}\}\(x\)\\geq 0\.50\. The CFC score is still computed from the deterministic removal channel and does not use this naturalistic label channel\.

Table 9:Naturalistic field\-unavailability proxy\.Brittle labels are derived from observed missing, unknown, special\-code, or unavailable field patterns rather than uniformly random stress\. Higher AUROC is better\.This proxy is still not a deployment incident log, but it is stricter than purely synthetic stress: the affected groups are selected from naturally occurring field\-unavailability patterns in the raw data\. Agreement between CFC rankings and this proxy would strengthen the claim that CFC captures operationally meaningful evidence dependence rather than only synthetic perturbation sensitivity\.

## Appendix NTargeted Certificate Ablations

We include targeted ablations only where they directly test the certificate design\. These checks are not used as the main empirical claim; they support the central result by asking whether CFC\-FDS is reducible to a generic confidence score, a single certificate component, a particular audit depth, a finely tuned weighting scheme, or ordinary probability calibration\.

Table 10:Selective component ablation\.AUROC is averaged over the quick robustness run using three datasets, three model families, one seed, and 256 audited samples per dataset–model pair\. Higher is better\. Only the strongest and most diagnostic comparisons are reported\.Table[10](https://arxiv.org/html/2609.00366#A14.T10)supports the component\-necessity claim\. Flip budget is the strongest individual component, but it still trails the full certificate substantially\. RCMA and degradation thresholds are informative but incomplete\. The full CFC\-FDS ranking is strongest because it combines abrupt flip risk, gradual support collapse, and partial\-degradation sensitivity into one trajectory\-level retrieval head\.

Table 11:Design\-stability checks\.These targeted checks test whether the full certificate depends on a narrow audit depth, fragile weighting choice, or ordinary probability calibration\. Higher AUROC and higher rank correlation are better\.Table[11](https://arxiv.org/html/2609.00366#A14.T11)shows that the ranking signal is not tied to one exact audit depth or a finely tuned weight vector\. The calibration rows further show that global temperature scaling preserves the CFC ranking, supporting the claim that CFC captures structural evidence brittleness rather than ordinary probability miscalibration\.

## Appendix OThreshold\-Sensitivity Diagnostics

![Refer to caption](https://arxiv.org/html/2609.00366v1/x5.png)Figure 5:Threshold\-sensitivity analysis\.Selective risk changes smoothly as the confidence threshold is tightened, indicating that the method is not tuned to a narrow operating regime\.
## Appendix PReproduction Pseudocode for Empirical Analyses

#### Main AUROC / RCMA evaluation\.

Algorithm 4Main predictive and fragility evaluation1:dataset

DD, model family

mm, seed

ss, grouping rule

𝒢\\mathcal\{G\}, baseline rule

x¯\\bar\{x\}, audit depth

KK
2:split

DDinto train, validation, and test partitions using seed

ss
3:fit preprocessing on train data only

4:construct feature groups

𝒢\\mathcal\{G\}by tracing transformed columns to raw variables

5:train model

fθm,sf\_\{\\theta\}^\{m,s\}on the train split

6:compute test predictions, probabilities, logits, margins, AUROC, macro\-F1, NLL, ECE, and Brier score

7:foreach test sample

xxdo

8:compute one\-step group margin drops

δg​\(x\)\\delta\_\{g\}\(x\)for all

g∈𝒢g\\in\\mathcal\{G\}
9:sort groups by descending

δg​\(x\)\\delta\_\{g\}\(x\)
10:progressively remove top\-ranked groups up to depth

KK
11:store flip budget

k⋆​\(x\)k^\{\\star\}\(x\)and RCMA

\(x\)\(x\)
12:aggregate AUROC and mean RCMA over the test split

#### Held\-out brittle\-case AUROC\.

Algorithm 5Held\-out brittle\-case ranking evaluation1:trained model

fθf\_\{\\theta\}, test set

XtestX\_\{\\mathrm\{test\}\}, score\-channel operators

ℙscore\\mathbb\{P\}\_\{\\mathrm\{score\}\}, label\-channel operators

ℙlabel\\mathbb\{P\}\_\{\\mathrm\{label\}\}
2:foreach test sample

xxdo

3:compute confidence, entropy, margin, and energy scores from the original prediction

4:compute CFC trajectory using

ℙscore\\mathbb\{P\}\_\{\\mathrm\{score\}\}
5:compute CFC\-RCMA and CFC\-FDS

6:initialize held\-out brittle label

bheldout​\(x\)=0b\_\{\\mathrm\{heldout\}\}\(x\)=0
7:foreach held\-out stressor

𝒫∈ℙlabel\\mathcal\{P\}\\in\\mathbb\{P\}\_\{\\mathrm\{label\}\}do

8:apply

𝒫\\mathcal\{P\}to

xxwithout using the CFC score\-channel trajectory

9:ifprediction flips or normalized margin collapse exceeds thresholdthen

10:set

bheldout​\(x\)=1b\_\{\\mathrm\{heldout\}\}\(x\)=1
11:compute AUROC of each ranking score against

bheldoutb\_\{\\mathrm\{heldout\}\}
12:report per\-dataset AUROC and paired bootstrap intervals over dataset–model–seed units

#### Budgeted Capture@20\.

Algorithm 6Budgeted review capture1:ranking score

s​\(x\)s\(x\), held\-out brittle labels

bheldout​\(x\)b\_\{\\mathrm\{heldout\}\}\(x\), review budget

qq
2:restrict evaluation to originally high\-confidence test samples

3:rank samples in descending predicted brittleness by

s​\(x\)s\(x\)
4:select

Topq​\(s\)\\mathrm\{Top\}\_\{q\}\(s\), the top

q%q\\%highest\-ranked samples

5:define brittle set

ℬ=\{x:bheldout​\(x\)=1\}\\mathcal\{B\}=\\\{x:b\_\{\\mathrm\{heldout\}\}\(x\)=1\\\}
6:compute

Capture​@​q=\|Topq​\(s\)∩ℬ\|/\(\|ℬ\|\+ϵ\)\\mathrm\{Capture@\}q=\|\\mathrm\{Top\}\_\{q\}\(s\)\\cap\\mathcal\{B\}\|/\(\|\\mathcal\{B\}\|\+\\epsilon\)
7:repeat for

q∈\{5,10,20\}q\\in\\\{5,10,20\\\}
8:compute AURC by progressively escalating highest\-ranked samples and measuring residual risk

#### Perturbation and attribution baselines\.

Algorithm 7Perturbation and attribution\-style baseline comparison1:trained model

fθf\_\{\\theta\}, feature groups

𝒢\\mathcal\{G\}, test set

XtestX\_\{\\mathrm\{test\}\}, validation set

XvalX\_\{\\mathrm\{val\}\}
2:foreach test sample

xxdo

3:compute CFC\-FDS from the full ordered stress trajectory

4:compute one\-step margin\-drop score

maxg∈𝒢⁡δg​\(x\)\\max\_\{g\\in\\mathcal\{G\}\}\\delta\_\{g\}\(x\)
5:compute permutation\-importance group ordering from validation\-set performance degradation

6:compute sample score under the permutation\-derived ordering

7:compute group\-SHAP values and aggregate absolute attribution within each group

8:compute SHAP\-concentration score from the highest\-ranked groups

9:evaluate each score against held\-out brittle labels using AUROC

10:compare one\-step, permutation, group\-SHAP, CFC\-RCMA, and CFC\-FDS rankings

#### Baseline\-choice sensitivity\.

Algorithm 8Baseline replacement sensitivity1:trained model

fθf\_\{\\theta\}, test set

XtestX\_\{\\mathrm\{test\}\}, feature groups

𝒢\\mathcal\{G\}, baseline candidates

ℬbase\\mathcal\{B\}\_\{\\mathrm\{base\}\}
2:foreach baseline rule

x¯\(r\)∈ℬbase\\bar\{x\}^\{\(r\)\}\\in\\mathcal\{B\}\_\{\\mathrm\{base\}\}do

3:foreach test sample

xxdo

4:recompute CFC trajectory using

x¯\(r\)\\bar\{x\}^\{\(r\)\}
5:store

kr⋆​\(x\)k^\{\\star\}\_\{r\}\(x\), RCMA

\(x\)r\{\}\_\{r\}\(x\), degradation thresholds, and FDS

\(x\)r\{\}\_\{r\}\(x\)
6:evaluate brittle\-case AUROC and ranking correlation with the default baseline

7:report whether CFC\-FDS remains strong across baseline choices

#### Naturalistic field\-unavailability proxy\.

Algorithm 9Naturalistic field\-unavailability proxy1:raw dataset

DD, trained model

fθf\_\{\\theta\}, raw\-to\-transformed group map, unavailable\-field markers

2:identify raw variables containing observed missing, unknown, special\-code, or unavailable markers

3:foreach eligible test sample

xxdo

4:compute original prediction, confidence, and margin

5:construct naturalistic stress state by replacing only groups linked to observed unavailable\-field patterns

6:compute stressed prediction and stressed margin

7:iforiginal prediction is high\-confidence and prediction flips or margin collapse exceeds thresholdthen

8:assign naturalistic brittle label

bnat​\(x\)=1b\_\{\\mathrm\{nat\}\}\(x\)=1
9:else

10:assign

bnat​\(x\)=0b\_\{\\mathrm\{nat\}\}\(x\)=0
11:compute confidence scores, one\-step scores, group\-SHAP scores, CFC\-RCMA, and CFC\-FDS

12:evaluate AUROC of all ranking scores against

bnatb\_\{\\mathrm\{nat\}\}

#### Brittleness\-aware temperature correction\.

Algorithm 10Brittleness\-aware temperature correction1:validation logits, test logits, validation labels, test labels, validation FDS, test FDS

2:fit global temperature

T0T\_\{0\}on validation data by minimizing NLL

3:fit min–max normalization of FDS on validation data

4:apply the validation\-fitted FDS normalization to test FDS

5:foreach

η∈\{0,0\.25,0\.5,1\.0,2\.0\}\\eta\\in\\\{0,0\.25,0\.5,1\.0,2\.0\\\}do

6:compute

T​\(x\)=T0\+η⋅Norm​\(FDS​\(x\)\)T\(x\)=T\_\{0\}\+\\eta\\cdot\\mathrm\{Norm\}\(\\mathrm\{FDS\}\(x\)\)on validation data

7:compute validation NLL using logits divided by

T​\(x\)T\(x\)
8:select

η⋆\\eta^\{\\star\}with lowest validation NLL

9:apply

T​\(x\)=T0\+η⋆⋅Norm​\(FDS​\(x\)\)T\(x\)=T\_\{0\}\+\\eta^\{\\star\}\\cdot\\mathrm\{Norm\}\(\\mathrm\{FDS\}\(x\)\)to test logits

10:compute ECE, Brier, NLL, fragile\-subset ECE, and fragile\-subset NLL

### P\.1Dataset–model–seed variance

To ensure that the brittle\-case ranking gains are not driven by a small number of datasets, models, or random seeds, we report results at the dataset–model–seed level\. Each unit corresponds to one trained backbone on one dataset under one seed\. We compute paired bootstrap intervals over these units and additionally report the fraction of units where CFC\-derived scores improve over the strongest non\-certificate baseline\.

Table 12:Seed\-level robustness of brittle\-case ranking\.Mean AUROC and standard deviation are computed over dataset–model–seed units\. Win rate is the fraction of units where the method exceeds the strongest non\-certificate baseline\.CFC\-FDS improves over the strongest non\-certificate baseline in every dataset–model–seed unit\. This directly addresses the possibility that the main AUROC gain is caused by one favorable benchmark, one model family, or one random seed\. CFC\-RCMA is informative but less stable, improving over the strongest non\-certificate baseline in 68\.8% of units, whereas the full certificate ranking head reaches a 100\.0% win rate\.

#### Protocol guarantee\.

For fixedfθf\_\{\\theta\}, preprocessing map, group partition𝒢\\mathcal\{G\}, baselinex¯\\bar\{x\}, audit depthKK, operatorsℙ\\mathbb\{P\}, severity gridΛ\\Lambda, and deterministic tie\-breaking, CFC is an exact finite witness of the model’s behavior under the declared stress protocol:

𝒞​\(x;fθ,𝒢\)=Audit​\(x,fθ,𝒢,x¯,K,ℙ,Λ\)\.\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)=\\mathrm\{Audit\}\(x,f\_\{\\theta\},\\mathcal\{G\},\\bar\{x\},K,\\mathbb\{P\},\\Lambda\)\.Thus, if two auditors use the same declared inputs, they obtain the same trajectory, flip budget, RCMA, degradation thresholds, and FDS\. This is the sense in which CFC is a certificate: it certifies the observed support\-collapse path under a specified protocol, not robustness to all possible corruptions or feature subsets\.

## Appendix QHeld\-Out Evaluation, Budgeted Retrieval, Attribution Baselines, and Calibration Correction

This appendix reports the additional analyses used to separate the proposed certificate from ordinary confidence scoring, one\-shot feature perturbation, and post\-hoc calibration\. The goal is to ensure that CFC is evaluated as a structured evidence\-failure certificate rather than as a self\-retrieval score\. Appendix[Q\.1](https://arxiv.org/html/2609.00366#A17.SS1)defines the non\-circular brittle\-case labeling protocol\. Appendix[Q\.2](https://arxiv.org/html/2609.00366#A17.SS2)reports review\-budget utility\. Appendix[S\.1](https://arxiv.org/html/2609.00366#A19.SS1)compares CFC against perturbation and attribution\-style ranking baselines\. Appendix[S\.2](https://arxiv.org/html/2609.00366#A19.SS2)evaluates brittleness\-aware temperature correction\.

### Q\.1Non\-circular brittle\-case label definition

A central risk in evaluating certificate\-derived scores is circularity\. If brittle labels are defined from the same trajectory used to compute the ranking score, then high AUROC may reflect self\-retrieval rather than independent vulnerability prediction\. We avoid this by separating the score channel from the label channel\.

#### Score channel\.

For each test sample, CFC\-FDS is computed from the deterministic greedy removal trajectory\. Groups are ranked by one\-step margin drop, the top\-KKremoval path is constructed, and FDS combines RCMA, greedy flip budget, and deterministic degradation thresholds using fixed weights\. This channel is the only source of the reported CFC ranking score\.

#### Label channel\.

The brittle\-case label is computed from held\-out stress families not used to compute the deterministic removal score\. These held\-out stressors include stochastic group masking, within\-group dropout, and bounded additive noise\. A sample is labeled as independently brittle only if it is originally high\-confidence and undergoes a decision flip or large support collapse under these held\-out stressors\. Thus, CFC\-FDS is evaluated on whether it predicts vulnerability under stress mechanisms that are disjoint from the stress path used to compute the score\.

#### Why this is not self\-retrieval\.

The score channel observes deterministic group removal ordered by margin drop\. The label channel observes independently sampled stress events from stochastic masking, dropout, and noise\. These operators share the broad semantic theme of evidence degradation, but they do not reuse the same trajectory, thresholds, or score components\. The evaluation therefore asks whether the certificate captures a cross\-operator structural property of the sample\-model pair\. This is stricter than ranking samples by the same perturbation used to define the target label\.

#### Fixed thresholds\.

High\-confidence thresholds, collapse thresholds, review budgets, FDS weights, and calibration subsets are fixed before test evaluation\. Hyperparameters for brittleness\-aware temperature correction are selected only on validation data\. FDS normalization is validation\-fitted and then applied to the test set without using test labels\. These choices prevent post\-hoc threshold selection and test\-label leakage\.

Table 13:Non\-circular brittle\-case evaluation protocol\.The score channel computes the ranking signal, while the label channel defines independent brittle\-case targets using held\-out evidence\-failure operators\. FDS is never used to assign the brittle label\.Formally, lets​\(x\)s\(x\)be a ranking score computed on the score channel and letℋ\\mathcal\{H\}denote the held\-out label\-channel stress operators\. For a high\-confidence samplexx, we define the held\-out brittle label as

bheldout\(x\)=𝕀\[∃𝒫∈ℋ,λ∈Λheldout:y^\(𝒫λ\(x\)\)≠y^\(x\)∨m​\(x\)−m​\(𝒫λ​\(x\)\)\|m​\(x\)\|\+ϵ≥τcollapse\]\.\\tiny b\_\{\\mathrm\{heldout\}\}\(x\)=\\mathbb\{I\}\\left\[\\exists\\mathcal\{P\}\\in\\mathcal\{H\},\\lambda\\in\\Lambda\_\{\\mathrm\{heldout\}\}:\\hat\{y\}\(\\mathcal\{P\}\_\{\\lambda\}\(x\)\)\\neq\\hat\{y\}\(x\)\\;\\;\\vee\\;\\;\\frac\{m\(x\)\-m\(\\mathcal\{P\}\_\{\\lambda\}\(x\)\)\}\{\|m\(x\)\|\+\\epsilon\}\\geq\\tau\_\{\\mathrm\{collapse\}\}\\right\]\.\(33\)The ranking scores​\(x\)s\(x\)is then evaluated by AUROC, budgeted capture, and risk\-coverage metrics againstbheldout​\(x\)b\_\{\\mathrm\{heldout\}\}\(x\)\. In all held\-out brittle\-case experiments,bheldout​\(x\)b\_\{\\mathrm\{heldout\}\}\(x\)is computed without access to FDS, CFC rank, or confidence\-baseline rank\.

### Q\.2Review\-budget capture under held\-out evidence failure

AUROC measures ranking quality over the full audit set, but deployment decisions often operate under a limited review budget\. We therefore report budgeted capture: among independently brittle high\-confidence cases, how many are recovered when only the topq%q\\%ranked predictions can be reviewed, escalated, or reacquired? This directly measures whether CFC provides operational value when auditing capacity is limited\.

For a scoress, letTopq​\(s\)\\mathrm\{Top\}\_\{q\}\(s\)be the topq%q\\%of samples ranked by predicted brittleness and letℬ=\{x:bheldout​\(x\)=1\}\\mathcal\{B\}=\\\{x:b\_\{\\mathrm\{heldout\}\}\(x\)=1\\\}be the set of independently brittle cases\. We compute

Capture​@​q​\(s\)=\|Topq​\(s\)∩ℬ\|\|ℬ\|\+ϵ\.\\tiny\\mathrm\{Capture@\}q\(s\)=\\frac\{\|\\mathrm\{Top\}\_\{q\}\(s\)\\cap\\mathcal\{B\}\|\}\{\|\\mathcal\{B\}\|\+\\epsilon\}\.\(34\)We also reportFalseConfCaptured​@​20\\mathrm\{FalseConfCaptured@20\}, the fraction of high\-confidence held\-out failures captured in the top20%20\\%ranked cases, and AURC, the area under the residual risk–coverage curve after progressively escalating the highest\-risk cases\.

Table 14:Review\-budget utility under held\-out evidence failure\.Capture@qqmeasures the fraction of independently brittle high\-confidence cases recovered by reviewing the topq%q\\%ranked samples\. Higher Capture and FalseConfCaptured values are better; lower AURC is better\. Values are averaged across datasets and model families\.The budgeted results show that CFC\-FDS is not only a stronger full\-ranking signal, but also a substantially more useful triage mechanism\. Under a20%20\\%review budget, CFC\-FDS recovers88\.9%88\.9\\%of independently brittle high\-confidence cases, compared with31\.831\.8–37\.4%37\.4\\%for generic confidence and energy\-based scores\. This supports the operational interpretation of CFC as a review, escalation, and evidence\-reacquisition tool rather than merely an offline diagnostic statistic\.

#### Protocol determinism\.

For fixed model, preprocessing, grouping, baseline, audit depth, stress operators, severity grid, and tie\-breaking rule, CFC is deterministic and exactly recomputable\. Therefore, all reported certificate fields are invariant to auditor implementation except for numerical precision\. This is the guarantee provided by the certificate; it is not a guarantee of global minimality or worst\-case robustness\.

### Q\.3Direct perturbation and attribution baselines

To test whether CFC reduces to ordinary feature perturbation or attribution, we compare against three direct alternatives\.

#### One\-step group perturbation\.

For each groupgg, we remove only that group and record the largest one\-step confidence or margin drop\. This baseline measures local sensitivity but does not construct a progressive trajectory, flip budget, margin\-collapse area, or degradation threshold\. It is therefore the closest “stress\-test” baseline but lacks the certificate structure\.

#### Permutation importance\.

We compute group\-level permutation scores by permuting each raw feature group and measuring the induced loss in prediction support\. This captures feature dependence at the group level but remains an aggregate or one\-step ranking signal rather than a per\-sample failure path\.

#### Group\-SHAP\.

We aggregate SHAP values over transformed coordinates belonging to the same raw feature group\. This produces a local attribution map for the original prediction, but attribution magnitude does not necessarily identify the ordered feature\-removal path that causes decision collapse\.

#### Interpretation\.

These baselines answer different questions\. One\-step perturbation asks which single group has the largest immediate effect\. Permutation importance asks which groups matter under random exchange\. Group\-SHAP asks which groups contributed to the original prediction\. CFC asks how prediction support collapses along an ordered evidence\-failure trajectory\. The empirical comparison therefore tests whether trajectory\-level fragility carries information beyond local effect size, global perturbation importance, and attribution concentration\.

## Appendix RGrouping and Baseline Protocol Dependence

CFC is intentionally protocol\-relative: the certificate is valid under a declared grouping rule, baseline replacement rule, stress\-operator family, severity grid, and audit depth\. This section clarifies how grouping and baseline choices should be interpreted\. The goal is not to claim invariance to arbitrary protocols, but to show that the main ranking conclusion is not an artifact of a single replacement convention and to define how grouping choices should be audited\.

#### Protocol object\.

Let the declared CFC protocol be

Π=\(𝒢,x¯,ℙ,Λ,K\),\\Pi=\\left\(\\mathcal\{G\},\\bar\{x\},\\mathbb\{P\},\\Lambda,K\\right\),\(35\)where𝒢\\mathcal\{G\}is the evidence grouping,x¯\\bar\{x\}is the replacement baseline,ℙ\\mathbb\{P\}is the stress\-operator family,Λ\\Lambdais the severity grid, andKKis the audit depth\. A certificate should therefore be read as𝒞Π​\(x;fθ\)\\mathcal\{C\}\_\{\\Pi\}\(x;f\_\{\\theta\}\)rather than as an unconditional property ofxxorfθf\_\{\\theta\}\. This notation makes the scope explicit: changingΠ\\Pican change the certificate\.

#### Grouping interpretation\.

The default grouping traces transformed features back to their raw variable of origin\. This is reproducible, preprocessing\-aware, and appropriate when raw fields correspond to plausible data\-acquisition units\. However, the grouping is not assumed to be causally optimal\. If domain evidence blocks are known, they should replace raw\-origin groups\. If features are highly redundant or causally linked, they may be merged into larger evidence blocks\. If groups are arbitrary, excessively fragmented, or semantically meaningless, the certificate remains recomputable but becomes less informative as an operational audit\.

#### Baseline interpretation\.

The baselinex¯\\bar\{x\}is a transformed\-space replacement state used to simulate missing or low\-trust evidence\. It is not a causal absence model\. A useful baseline should represent a declared operational convention: training mean, training median, neutral transformed value, categorical mode, missing\-token value, or a domain\-defined unavailable state\. The correct choice depends on the workflow being audited\.

#### Baseline sensitivity experiment\.

We compare four baseline choices: training\-set mean replacement for standardized numeric features, training\-set median replacement, zero replacement in transformed space, and empirical missing\-token or mode replacement for categorical groups where available\. For each baseline, we recompute CFC trajectories, RCMA, greedy flip budgets, and FDS rankings while keeping the trained model, data split, group partition, audit depth, and held\-out brittle\-label protocol fixed\.

Table 15:Baseline\-choice sensitivity\.CFC\-FDS remains substantially stronger than the best non\-certificate score across replacement conventions\.The ranking advantage is stable across replacement conventions\. The training\-mean baseline gives the strongest result, but median, neutral\-zero, and mode/missing\-token replacement all preserve a large CFC\-FDS advantage over the best non\-certificate score\. The neutral\-zero baseline is slightly weaker, as expected, because it may create less realistic transformed\-space states for standardized numeric features\. However, the effect size remains large in all cases, suggesting that the main conclusion is not an artifact of a single baseline convention\.

#### Grouping\-sensitivity diagnostic\.

Grouping sensitivity should be evaluated by recomputing the certificate under alternative admissible groupings while keeping the trained model, data split, baseline, audit depth, stress operators, and held\-out brittle labels fixed\. We distinguish three grouping variants:

- •Raw\-origin grouping: the default protocol, where all transformed columns derived from the same raw variable form one evidence block\.
- •Domain\-block grouping: expert\-defined or workflow\-defined groups, such as laboratory panels, questionnaire modules, sensor families, administrative fields, or source\-specific data blocks\.
- •Redundancy\-merged grouping: groups merged when they are strongly correlated, causally linked, or known to compensate for one another\.

A grouping is considered stable for the CFC claim if CFC\-FDS remains above the strongest non\-certificate baseline and if its ranking is strongly correlated with the default protocol\. A grouping is considered semantically weak if it produces unstable rankings, low agreement with domain\-defined blocks, or evidence paths that cannot be interpreted as plausible workflow failures\.

Algorithm 11Grouping and baseline sensitivity diagnostic1:trained model

fθf\_\{\\theta\}, test set

XtestX\_\{\\mathrm\{test\}\}, grouping candidates

\{𝒢\(r\)\}\\\{\\mathcal\{G\}^\{\(r\)\}\\\}, baseline candidates

\{x¯\(b\)\}\\\{\\bar\{x\}^\{\(b\)\}\\\}, held\-out brittle labels

bheldoutb\_\{\\mathrm\{heldout\}\}
2:foreach grouping rule

𝒢\(r\)\\mathcal\{G\}^\{\(r\)\}do

3:foreach baseline rule

x¯\(b\)\\bar\{x\}^\{\(b\)\}do

4:foreach test sample

xxdo

5:recompute CFC trajectory under protocol

Π\(r,b\)=\(𝒢\(r\),x¯\(b\),ℙ,Λ,K\)\\Pi^\{\(r,b\)\}=\(\\mathcal\{G\}^\{\(r\)\},\\bar\{x\}^\{\(b\)\},\\mathbb\{P\},\\Lambda,K\)
6:store

kr,b⋆​\(x\)k^\{\\star\}\_\{r,b\}\(x\), RCMA

\(x\)r,b\{\}\_\{r,b\}\(x\), degradation thresholds, and FDS

\(x\)r,b\{\}\_\{r,b\}\(x\)
7:evaluate FDSr,bAUROC against

bheldoutb\_\{\\mathrm\{heldout\}\}
8:compute rank correlation with the default protocol FDS ranking

9:report which protocol variants preserve the CFC\-FDS advantage and which weaken interpretation

#### Interpretation\.

This diagnostic turns the grouping and baseline concern into a declared sensitivity analysis\. If the CFC\-FDS advantage persists across reasonable grouping and baseline choices, the result supports a stable evidence\-dependence signal\. If it fails under a particular grouping, the failure is informative: it indicates that the chosen grouping does not align with meaningful evidence units for that dataset or workflow\. Thus, CFC should be treated as a protocol\-relative audit certificate whose usefulness depends on whether the declared evidence blocks and replacement states match the operational failure being studied\.

### R\.1Fixed FDS weighting

The FDS ranking head combines three certificate components: RCMA, reciprocal flip budget, and reciprocal degradation threshold\. We use fixed weights rather than fitting weights on the test set\. This design is intentional\. FDS is not introduced as a learned failure predictor; it is a deterministic retrieval head over the certificate\. Fixed weighting prevents the method from becoming a supervised meta\-classifier over stress outcomes and preserves the interpretation of CFC as an audit object\. The default weighting gives positive mass to all three failure modes because they are not interchangeable\. A sample can be fragile because it flips after one group removal, because its margin collapses rapidly without flipping, or because small partial degradation is enough to change the decision\. Removing any component therefore discards one mode of brittleness\. Component ablations test this directly by comparing RCMA\-only, flip\-budget\-only, degradation\-threshold\-only, and full FDS rankings\. In deployment, FDS weights could be adapted to domain costs\. For example, a workflow that can reacquire missing fields may emphasize flip budget, while a monitoring system concerned with gradual quality degradation may emphasize RCMA or degradation thresholds\. The experiments use fixed weights to avoid test\-time tuning and to make the reported ranking protocol reproducible\.

## Appendix SReproducibility Details

The released artifact will include scripts for dataset preprocessing, raw\-to\-transformed group tracing, baseline construction, model training, certificate generation, held\-out brittle\-label construction, ranking evaluation, bootstrap confidence intervals, seed aggregation, and calibration correction\. Each certificate row stores the sample identifier, dataset, model family, seed, original prediction, original confidence, original margin, ordered group path, greedy flip budget, RCMA, degradation thresholds, FDS, and held\-out brittle label\. This makes the main results recomputable from serialized model predictions and declared stress operators\. All datasets are public tabular benchmarks\. Splits, random seeds, preprocessing maps, and grouping metadata are fixed before evaluation\. The code reports both aggregate metrics and dataset–model–seed units, enabling paired bootstrap intervals and win\-rate calculations\. The brittleness\-aware temperature correction is fitted only on validation data; test labels are not used for FDS normalization, fragile\-subset selection, or hyperparameter tuning\.

### S\.1Comparison against perturbation and attribution\-style baselines

CFC is related to feature perturbation and attribution analysis, but it is not equivalent to either\. A one\-shot perturbation score estimates the effect of removing a single feature group, while CFC records an ordered stress trajectory, a flip budget, a margin\-collapse area, partial\-degradation thresholds, and a ranking head\. To test whether this trajectory\-level structure matters, we compare CFC against confidence baselines, random group ordering, one\-step margin\-drop ordering, permutation\-importance ordering, and group\-level SHAP aggregation\.

Table 16:Comparison against perturbation and attribution\-style ranking baselines\.Held\-out brittle\-case AUROC is computed using the non\-circular label\-channel protocol in Appendix[Q\.1](https://arxiv.org/html/2609.00366#A17.SS1)\. Higher is better\. Values are averaged across datasets and model families\.The comparison isolates the contribution of the certificate structure\. One\-step margin drop, permutation importance, and group\-level SHAP aggregation improve over generic confidence scores, showing that feature\-dependence information is relevant\. However, none of these one\-shot or attribution\-style baselines matches CFC\-FDS\. The gap between GroupSHAP aggregation and CFC\-FDS indicates that brittle\-case retrieval is not explained merely by identifying influential feature groups\. Instead, the strongest signal comes from combining abrupt flip risk, progressive support collapse, and partial\-degradation sensitivity into a trajectory\-level certificate\.

#### Protocol guarantee\.

For fixedfθf\_\{\\theta\}, preprocessing map, group partition𝒢\\mathcal\{G\}, baselinex¯\\bar\{x\}, audit depthKK, operatorsℙ\\mathbb\{P\}, severity gridΛ\\Lambda, and deterministic tie\-breaking, CFC is an exact finite witness of the model’s behavior under the declared stress protocol:

𝒞​\(x;fθ,𝒢\)=Audit​\(x,fθ,𝒢,x¯,K,ℙ,Λ\)\.\\mathcal\{C\}\(x;f\_\{\\theta\},\\mathcal\{G\}\)=\\mathrm\{Audit\}\(x,f\_\{\\theta\},\\mathcal\{G\},\\bar\{x\},K,\\mathbb\{P\},\\Lambda\)\.\(36\)Thus, if two auditors use the same declared inputs, they obtain the same trajectory, flip budget, RCMA, degradation thresholds, and FDS\. This is the sense in which CFC is a certificate: it certifies the observed support\-collapse path under a specified protocol, not robustness to all possible corruptions or feature subsets\.

#### Baseline definitions\.

*Random group order*ranks samples by the brittleness induced by a random ordering of feature groups, averaged over repeated random seeds\.*One\-step margin drop only*ranks samples by the largest immediate margin decrease after removing a single group, without constructing a progressive trajectory\.*Permutation importance order*ranks groups by validation\-set performance degradation after permutation and then evaluates sample\-level fragility under that fixed order\.*GroupSHAP / SHAP aggregate*aggregates absolute SHAP values within each feature group and ranks samples by the concentration of attribution in the most influential groups\. Unlike CFC, these baselines do not jointly encode the progressive collapse path, flip budget, and partial\-degradation threshold\.

### S\.2Brittleness\-aware temperature correction

The main paper defines brittleness\-aware temperature correction as a secondary use of the certificate\. The purpose of this analysis is not to claim that CFC replaces standard calibration, but to test whether structurally fragile samples benefit from stronger confidence discounting than stable samples\. The global temperatureT0T\_\{0\}is fitted on the validation split by minimizing validation NLL\. FDS normalization is also computed on validation data, and the local discount parameterη\\etais selected on validation data before test evaluation\. Test fragile subsets are selected by applying the validation\-fitted FDS normalization and taking the top 20% most fragile cases; test labels are not used to define the subset\. We therefore report calibration both on the full test set and on the top\-20%20\\%most fragile samples according to validation\-normalized FDS\.

Table 17:Brittleness\-aware temperature correction\.Calibration metrics are reported overall and on the top\-20%20\\%most fragile cases\. Lower is better for all metrics\. Values are averaged across datasets and model families\.The brittleness\-aware correction improves calibration most strongly on the fragile subset, where standard global temperature scaling remains limited because it applies the same confidence discount to structurally stable and structurally fragile samples\. By contrast, brittleness\-aware temperature correction increases the effective temperature for cases with high FDS, lowering overconfident probabilities precisely where the certificate indicates narrow evidence support\.

The brittleness\-aware temperature is applied as

T​\(x\)=T0\+η⋅Norm​\(FDS​\(x\)\),T\(x\)=T\_\{0\}\+\\eta\\cdot\\mathrm\{Norm\}\(\\mathrm\{FDS\}\(x\)\),\(37\)whereT0T\_\{0\}is the validation\-fitted global temperature andη\\etacontrols the local confidence discount applied to structurally fragile cases\. The correction is intentionally conservative: it does not change the predicted label and only rescales confidence more strongly for samples whose certificate indicates fragile support\.

Overall, BATS improves fragile\-subset calibration more strongly than global calibration, supporting its role as a targeted correction rather than a universal calibrator\.

Together, these analyses address the main failure modes a reviewer could suspect\. The non\-circular label protocol tests whether CFC predicts held\-out evidence\-failure vulnerability rather than retrieving its own score components\. The budgeted retrieval metrics test whether the certificate is useful under realistic audit budgets\. The perturbation and attribution comparisons test whether CFC is more than one\-step sensitivity or feature\-importance ranking\. The seed\-level analysis tests whether gains are concentrated in a small number of datasets, models, or random seeds\. The baseline\-sensitivity analysis tests whether the ranking advantage depends on a single replacement convention\. The calibration table tests whether the certificate can support targeted confidence correction without changing predicted labels\. These checks strengthen the interpretation of CFC as a protocol\-relative trajectory certificate rather than a repackaged confidence, attribution, or perturbation score\.

## Appendix TArtifact

We provide an anonymized review artifact as supplementary material and mirror it at:

[https://anonymous\.4open\.science/r/Counterfactual\-Fragility\-Certificates\-167F/](https://anonymous.4open.science/r/Counterfactual-Fragility-Certificates-167F/)

The artifact contains the CFC reference implementation, reproduction scripts, precomputed result tables, selected figures, tests, and an anonymization checklist\. It supports review\-time verification of certificate construction, score conversion, brittle\-label assignment, component ablations, and the main reported results\. The full production training grid is not included in the review artifact because it contains private orchestration paths and will be released in de\-anonymized form after review\.

Similar Articles