Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
Summary
A paper introducing a generator-based evaluation framework to study how Concept Bottleneck Models and prototype-based models respond to geometric-latent versus semantic-concept perturbations, arguing that interpretability does not inherently confer robustness but instead redistributes sensitivity across perturbation spaces.
View Cached Full Text
Cached at: 10/02/26, 09:51 AM
# Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
Source: [https://arxiv.org/html/2609.38625](https://arxiv.org/html/2609.38625)
Hanwei ZhangAffiliation:Department of Computer ScienceAffiliation:Saarland UniversityAffiliation:Saarbrücken, GermanyEmail:[zhang@depend\.uni\-saarland\.de](mailto:)Tianma HuAffiliation:School of Computer Science and EngineeringAffiliation:Tianjin University of TechnologyAffiliation:Tianjin, ChinaEmail:[hutianma15299@stud\.tjut\.edu\.cn](mailto:)Gaojie JinAffiliation:Department of Artificial Intelligence &Affiliation:Institute of Artificial Intelligence and Brain SciencesAffiliation:University of MacauAffiliation:Macau SAR, ChinaEmail:[gaojiejin@um\.edu\.mo](mailto:)Xu ChengAffiliation:School of Computer Science and EngineeringAffiliation:Tianjin University of TechnologyAffiliation:Tianjin, ChinaEmail:[xu\.cheng@ieee\.org](mailto:)Ronghui Mu††thanks:corresponding authorAffiliation:Department of Computer ScienceAffiliation:University of ExeterAffiliation:Exeter,UKEmail:[r\.mu2@exeter\.ac\.uk](mailto:)
###### Abstract
Concept Bottleneck Models \(CBMs\) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings\. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves\. To disentangle these factors, we introduce a generator\-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space\. We further evaluate a prototype\-based interpretable model, PixPNet, showing that the same perturbation\-to\-prediction evaluation and certification pipeline extends beyond concept bottlenecks without requiring alignment between heterogeneous internal representations\. Within this framework, we evaluate robustness both empirically, via prediction and concept\-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces\. Across experiments on CUB and RIVAL10 variants, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness\. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness\. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability robustness trade off that depends critically on the perturbation regime and task structure\. Together, our results show that interpretability and robustness are distinct objectives: interpretable intermediate representations do not uniformly improve robustness, but instead redistribute sensitivity across perturbation spaces and model families\.
## 1Introduction
Concept Bottleneck Models \(CBMs\)\([Koh et al\., 2020](https://arxiv.org/html/2609.38625#bib.bib23)\)are a prominent paradigm of interpretable learning, explicitly structuring predictions around human\-interpretable concepts\. By exposing intermediate concept representations, CBMs provide semantically grounded explanations, support targeted interventions, and offer a modular design that extends naturally in vision and multimodal settings\. These properties make CBMs attractive, particularly in high\-stakes and regulated domains such as those governed by the EU AI Act\([Hermanns et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib13);[Parliament and of the EU, 2024](https://arxiv.org/html/2609.38625#bib.bib33)\), where interpretability is often believed to promote robustness\. However, the robustness implications of introducing a concept bottleneck remain underexplored, and existing evidence is mixed\.
Prior work reaches conflicting conclusions regarding the robustness of CBMs\.[Rasheed et al\. \(2024\)](https://arxiv.org/html/2609.38625#bib.bib34)report robustness gains from concept bottlenecks, attributing them to the removal of non\-essential variation, with robustness evaluated at the level of final class predictions\. In contrast,[Sinha et al\. \(2023\)](https://arxiv.org/html/2609.38625#bib.bib38)argues that concept bottlenecks can reduce robustness: concept representations themselves are often unstable under perturbations, and exposing an explicit concept interface introduces additional attack surfaces\. As illustrated in[Figure 1](https://arxiv.org/html/2609.38625#S1.F1)\(c\), we observe this tension directly: under different task settings of the same dataset, evaluating the same standard classifier and its concept bottleneck counterpart can lead to opposite conclusions about whether concept bottlenecks improve robustness\.
\(a\)Evaluation framework
\(b\)Robustness gap summary
\(c\)Concrete contrast
Figure 1:Evaluation framework and representative results\.\(a\) Evaluation robustness across geometric*vs\.*semantic perturbations and empirical*vs\.*certified criteria\. \(b\) The heatmap reports representativefcbmf\_\{cbm\}–fstdf\_\{std\}robustness gaps usingϵ=0\.3\\epsilon=0\.3,τ=5\\tau=5,σ=0\.10\\sigma=0\.10, andρ=0\.10\\rho=0\.10for the four columns, respectively\. All gaps are oriented so that positive values indicate greater CBM robustness\. Detailed computations in Appendix[C\.1](https://arxiv.org/html/2609.38625#A3.SS1)\(c\) Certified concept\-radius examples atρ=0\.10\\rho=0\.10show opposite outcomes across task settings\.At first glance, these findings appear irreconcilable\. We argue that they instead stem from different notions of robustness and perturbation regimes\. In practice, robustness is still most often assessed using small pixel\-level perturbations, which lack semantic meaning and fail to reflect changes in underlying concepts\. Whether a concept bottleneck improves or degrades robustness depends on*where robustness is measured*\(class*vs\.*concept level\),*how perturbations are applied*, and*the task conditions*under which the model operates\. Making these distinctions explicit is key to understanding the role of concept bottlenecks in robust learning\. This leads to the central question of this work:
> How does introducing an interpretable intermediate representation reshape robustness relative to a standard classifier, and how does this depend on the perturbation space and robustness criterion?
To study this systematically, we adopt a concept bottleneck generator\([Kulkarni et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib26)\)that maps latent representations to an explicit concept space, enabling controlled concept\-level interventions before image generation\. The generator is frozen and shared across classifiers, serving only as an experimental instrument\. This design allows us to disentangle two fundamentally different types of perturbations: \(i\)*geometric perturbations*, which induce continuous variations in the latent embedding space, and \(ii\)*semantic perturbations*, which directly modify individual concepts\. Within this framework, we evaluate robustness*empirically*by measuring changes in class predictions and intermediate representations, and*certifiably*by deriving guarantees via randomized smoothing\. Standard models and their corresponding CBMs share the same backbone, ensuring a fair comparison\.
Our analysis focuses on generator\-based perturbations to enable controlled evaluation of the robustness induced by concept bottlenecks, without claiming universal robustness guarantees\. Our results show that concept bottlenecks do not exhibit a uniform robustness behavior\. Instead, they can either improve or degrade robustness, depending on the setting\. In particular, we find that the effect of the bottleneck is strongly influenced by the semantic similarity between classes, and the number and relevance of concepts used by the model\. These findings provide a unifying explanation for previously conflicting results and clarify the conditions under which concept bottlenecks act as beneficial regularizers versus when they introduce fragility\.
Overall, our contributions are threefold:
\(i\)A unified robustness perspective for concept bottlenecks\.We reconcile previously conflicting findings on CBM robustness by showing that they often arise from conflating different perturbation spaces and robustness criteria\. By distinguishing prediction\-level and concept\-level sensitivity under geometric and semantic perturbations, as well as empirical, certified, and adversarial robustness, we clarify when concept bottlenecks suppress sensitivity and when they instead introduce additional vulnerabilities\. Our results emphasize that interpretability and robustness are distinct objectives, and that concept bottlenecks reshape robustness rather than uniformly improving it\.
\(ii\)A controlled evaluation and certification framework for geometric and semantic robustness\.We introduce a generator\-based framework \(Figure[1](https://arxiv.org/html/2609.38625#S1.F1)\(a\)\) that separates continuous latent\-space geometric perturbations from discrete concept\-level semantic interventions\. The framework enables matched comparisons under a common perturbation\-to\-prediction interface, together with empirical robustness metrics and perturbation\-specific randomized\-smoothing certificates in latent and concept spaces\. Importantly, the evaluation interface does not require different classifier architectures to share or align their internal representations\.
\(iii\)Empirical and certified analysis of when robustness gains emerge\.Through extensive experiments on CUB\-200\-2011 and RIVAL\-10 variants, we show that concept bottlenecks do not provide an inherent robustness advantage \(*e\.g\.*, Figure[1](https://arxiv.org/html/2609.38625#S1.F1)\(b\)\)\. Instead, robustness depends strongly on the perturbation regime, robustness criterion, class semantic similarity, and concept vocabulary design\. These findings provide practical guidance for designing CBMs that better balance interpretability and robustness\. We further validate the broader applicability of the evaluation pipeline on PixPNet, a prototype\-based interpretable model, showing that the same generator\-defined evaluation and certification procedure can be applied beyond CBMs without requiring alignment between generator concepts and model\-specific prototypes\.
## 2Related Work
##### Robustness of Concept Bottlenecks\.
Early work on interpretable models highlighted connections between interpretability and robustness, with Self\-Explaining Neural Networks \(SENN\) enforcing robustness as the stability of explanations under small perturbations\([Sawada and Nakamura, 2022](https://arxiv.org/html/2609.38625#bib.bib35)\)\. More recent studies directly examine robustness in Concept Bottleneck Models \(CBMs\)\.[Sinha et al\. \(2023\)](https://arxiv.org/html/2609.38625#bib.bib38)show that interpretability alone does not guarantee robustness, demonstrating that concept representations can be fragile under adversarial perturbations and that explicit concept interfaces introduce new attack surfaces, revealing a tension between concept\-level stability and adversarial robustness\. In contrast,[Rasheed et al\. \(2024\)](https://arxiv.org/html/2609.38625#bib.bib34)find that CBMs, particularly sequentially trained variants, can achieve higher adversarial accuracy than end\-to\-end models under weak to moderate attacks, suggesting that conceptual bottlenecks may act as an information\-filtering regularizer at the prediction level\. Related observations beyond CBMs further indicate that highly interpretable representations can remain non\-robust\([Li et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib28)\), while recent extensions such as language\-guided and flexible CBMs improve scalability and adaptability without addressing robustness explicitly\([Yu et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib45);[Du et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib9)\)\. Overall, prior work reports seemingly conflicting conclusions, which we argue arise from differences in robustness definitions, perturbation models, and task settings rather than fundamental inconsistencies\.
##### Randomized Smoothing for Certified Robustness\.
Randomized smoothing provides a general framework for certifying robustness by constructing a smoothed classifier through noise injection and aggregation, yielding probabilistic guarantees against bounded perturbations\([Cohen et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib6)\)\. While most work applies smoothing in pixel space, recent studies have explored extensions to learned representations or structured perturbations\([Hao et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib10);[Wang et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib44);[Mu et al\., 2023](https://arxiv.org/html/2609.38625#bib.bib31);[Mu et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib32)\)\. In this work, randomized smoothing is used primarily as a model\-agnostic certification and comparison tool\. The resulting smoothed predictor can also improve empirical robust accuracy in some adversarial settings, but defense is not the primary objective of our analysis\.
##### Interpretable intermediate architectures beyond CBMs\.
Interpretable intermediate representations are also exposed by post\-hoc concept methods\([Yuksekgonul et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib46)\)and prototype\-based architectures\([Zhang et al\., 2021](https://arxiv.org/html/2609.38625#bib.bib48);[Carmichael et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib3);[Sicre et al\., 2023](https://arxiv.org/html/2609.38625#bib.bib37)\)\. PCBM\([Yuksekgonul et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib46)\)constructs concept representations using CAV\-style directions, while prototype\-based models such as PixPNet\([Carmichael et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib3)\)ground predictions in localized parts\. Although these models expose different internal representations, our evaluation does not require these representations to be aligned: the generator defines a common external perturbation interface, and robustness is measured at the resulting prediction level\.
## 3Methodology
### 3\.1Perturbation Models
Letg:d→𝒳g:\\real^\{d\}\\rightarrow\\mathcal\{X\}denote a pretrained generator that maps a latent noise vector𝐳∈d\\mathbf\{z\}\\in\\real^\{d\}to an image𝐱=g\(𝐳\)\\mathbf\{x\}=g\(\\mathbf\{z\}\)\. We assume thatggadmits a decompositiong=g2∘g1g=g\_\{2\}\\circ g\_\{1\}, whereg1g\_\{1\}maps the latent code to an intermediate representation andg2g\_\{2\}maps this representation to the image space\. Following concept bottleneck generators, an encoder–decoder pair\(E,D\)\(E,D\)is intermediate layer, including an explicit concept space representation𝐜=E\(g1\(𝐳\)\)∈ℝK\\mathbf\{c\}=E\(g\_\{1\}\(\\mathbf\{z\}\)\)\\in\\mathbb\{R\}^\{K\}such that image generation can be written as𝐱=g\(𝐳\)=g2\(D\(E\(g1\(𝐳\)\)\)\)\\mathbf\{x\}=g\(\\mathbf\{z\}\)=g\_\{2\}\(D\(E\(g\_\{1\}\(\\mathbf\{z\}\)\)\)\)\. Each conceptcic\_\{i\}corresponds to a semantically meaningful attribute and is associated with a binary state indicating the presence or absence of the concept\. This representation allows us to distinguish two fundamentally different perturbation regimes:
##### Geometric perturbations
are applied in latent space by adding noise to the latent code,𝐳~=𝐳\+δ,‖δ‖2≤ϵ\\tilde\{\\mathbf\{z\}\}=\\mathbf\{z\}\+\\mathbf\{\\delta\},\\\|\\mathbf\{\\delta\}\\\|\_\{2\}\\leq\\epsilon, which induces a continuous variation along the generator manifold and results in a geometrically perturbed image𝐱~=g\(𝐳~\)\\tilde\{\\mathbf\{x\}\}=g\(\\tilde\{\\mathbf\{z\}\}\)\.
##### Semantic perturbations
are applied directly in concept space by modifying the concept vector𝐜\\mathbf\{c\}\. We uniformly sample a subsetS⊆\{1,⋯,K\}S\\subseteq\\\{1,\\cdots,K\\\}of size\|S\|=τ\|S\|=\\tauand, for eachi∈Si\\in S, swap the binary representation of corresponding concept from\[ci\+,ci−\]\[c\_\{i\}^\{\+\},c\_\{i\}^\{\-\}\]to\[ci−,ci\+\]\[c\_\{i\}^\{\-\},c\_\{i\}^\{\+\}\], simulating changes in high\-level semantic attributes while preserving the underlying geometric structure\. To detect concept changes under perturbation, we train a set of binary concept classifiers𝒞\(𝐜\)∈\{0,1\}K\\mathcal\{C\}\(\\mathbf\{c\}\)\\in\\\{0,1\\\}^\{K\}that map a concept representation to a binary concept state\. We then define a pseudo\-labeling functionϕi\(𝐜,𝐜~\):=𝟙\[𝒞\(ci\)≠𝒞\(ci~\)\]\\phi\_\{i\}\(\\mathbf\{c\},\\tilde\{\\mathbf\{c\}\}\)\\mathrel\{:=\}\\mathbbm\{1\}\[\\mathcal\{C\}\(c\_\{i\}\)\\neq\\mathcal\{C\}\(\\tilde\{c\_\{i\}\}\)\]to determine whether a given conceptcic\_\{i\}has flipped under perturbation\.
### 3\.2Statistical Robustness Analysis
Letf:𝒳→𝒴f:\\mathcal\{X\}\\to\\mathcal\{Y\}denote a classifier operating on generated images\. We consider two classes of models:*standard classifier*fstdf\_\{std\}directly maps the input image to a label;*concept bottleneck model*is defined as a tuplefcbm:=\(r,h\)f\_\{cbm\}\\mathrel\{:=\}\(r,h\), wherer:𝒳→Kr:\\mathcal\{X\}\\to\\real^\{K\}predicts a concept representation from image andh:K→𝒴h:\\real^\{K\}\\to\\mathcal\{Y\}maps concepts to the final prediction\. This setup enables a controlled comparison of how standard and concept\-based classifiers respond to geometric perturbations in latent space and semantic perturbations in concept space\. Further methodological details are provided in Appendix[B\.1](https://arxiv.org/html/2609.38625#A2.SS1)\.
##### Empirical Robustness\.
To quantify empirical robustness under geometric and semantic perturbations, we measure changes at the concept and prediction levels using the following metrics:
*1\. Concept Flip Rate \(CFR\)*measures the average fraction of concepts whose predicted states change under perturbation:
CFR:=𝔼\{𝐜=E\(𝐳\)\}\[1K∑i=1Kϕi\(𝐜,𝐜~\)\];\\mathrm\{CFR\}\\;\\mathrel\{:=\}\\;\\mathbbm\{E\}\_\{\\\{\\mathbf\{c\}=E\(\\mathbf\{z\}\)\\\}\}\\left\[\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\\phi\_\{i\}\(\\mathbf\{c\},\\tilde\{\\mathbf\{c\}\}\)\\right\];*2\. Prediction Flip Rate \(PFR\)*measures the probability that the classifier’s predicted label changes under perturbation:
PFR\(f\):=𝔼𝐱\[𝟙\[f\(𝐱\)≠f\(𝐱~\)\]\];\\mathrm\{PFR\}\(f\)\\;\\mathrel\{:=\}\\;\\mathbbm\{E\}\_\{\\mathbf\{x\}\}\\left\[\\mathbbm\{1\}\\bigl\[f\(\\mathbf\{x\}\)\\neq f\(\\tilde\{\\mathbf\{x\}\}\)\\bigr\]\\right\];*3\. Decision Function Sensitivity \(DFS\)*quantifies the sensitivity of a concept bottleneck classifierfcbmf\_\{\\mathrm\{cbm\}\}by measuring the change in its internal concept representation:
DFS:=𝔼𝐳\[‖r\(g\(𝐳\)\)−r\(g\(𝐳~\)\)‖2\]\.\\mathrm\{DFS\}\\;\\mathrel\{:=\}\\;\\mathbbm\{E\}\_\{\\mathbf\{z\}\}\\left\[\\left\\\|r\\bigl\(g\(\\mathbf\{z\}\)\\bigr\)\-r\\bigl\(g\(\\tilde\{\\mathbf\{z\}\}\)\\bigr\)\\right\\\|\_\{2\}\\right\]\.
CFR and DFS characterize concept\-level robustness, capturing both true concept flips induced by the generator and the sensitivity of the concept bottleneck classifier, while PFR measures robustness at the prediction level\. Together, these metrics allow us to study how concept instability propagates to model predictions\.
##### Certified Robustness\.
We use randomized smoothing to certify robustness in the two perturbation spaces induced by the generator\.
##### Latent\-space smoothing\.
For geometric robustness, we apply Gaussian smoothing in latent space:
f^G\(𝐳\)=argmaxy∈𝒴ℙδ∼𝒩\(0,σ2I\)\[f\(g\(𝐳\+δ\)\)=y\]\.\\hat\{f\}\_\{\\mathrm\{G\}\}\(\\mathbf\{z\}\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\}\\mathbb\{P\}\_\{\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\}\\\!\\left\[f\(g\(\\mathbf\{z\}\+\\delta\)\)=y\\right\]\.\(1\)
###### Theorem 3\.1\(Latent\-space Gaussian certificate\([Cohen et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib6)\)\)\.
LetAAandBBbe the top and runner\-up classes off^G\\hat\{f\}\_\{\\mathrm\{G\}\}with probabilitiespA\>pBp\_\{A\}\>p\_\{B\}\. Thenf^G\\hat\{f\}\_\{\\mathrm\{G\}\}is constant for all latent perturbations‖ϵ‖2≤RG\\\|\\epsilon\\\|\_\{2\}\\leq R\_\{\\mathrm\{G\}\}, where
RG=σ2\(Φ−1\(pA\)−Φ−1\(pB\)\)\.R\_\{\\mathrm\{G\}\}=\\frac\{\\sigma\}\{2\}\\left\(\\Phi^\{\-1\}\(p\_\{A\}\)\-\\Phi^\{\-1\}\(p\_\{B\}\)\\right\)\.\(2\)
##### Concept\-space smoothing\.
For semantic robustness, we sample independent concept flipsηi∼Bernoulli\(ρ\)\\eta\_\{i\}\\sim\\mathrm\{Bernoulli\}\(\\rho\)and writeℱ\(𝐜,η\)\\mathcal\{F\}\(\\mathbf\{c\},\\eta\)for the flipped concept representation\. The smoothed classifier is
f^C\(𝐜\)=argmaxy∈𝒴ℙη∼Bernoulli\(ρ\)K\[f\(g2\(D\(ℱ\(𝐜,η\)\)\)\)=y\]\.\\hat\{f\}\_\{\\mathrm\{C\}\}\(\\mathbf\{c\}\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\}\\mathbb\{P\}\_\{\\eta\\sim\\mathrm\{Bernoulli\}\(\\rho\)^\{K\}\}\\\!\\left\[f\\\!\\left\(g\_\{2\}\(D\(\\mathcal\{F\}\(\\mathbf\{c\},\\eta\)\)\)\\right\)=y\\right\]\.\(3\)
###### Theorem 3\.2\(Concept\-spaceℓ0\\ell\_\{0\}certificate\([Lee et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib27)\)\)\.
Lety⋆y^\{\\star\}be the top class off^C\\hat\{f\}\_\{\\mathrm\{C\}\}, and letp¯y⋆\\underline\{p\}\_\{y^\{\\star\}\}andp¯y\\overline\{p\}\_\{y\}be conservative confidence bounds on the top and competing class probabilities\. For any concept\-flip budgetΔ\\Delta, if
ℒΔ\(p¯y⋆\)\>𝒰Δ\(p¯y\),∀y≠y⋆,\\mathcal\{L\}\_\{\\Delta\}\(\\underline\{p\}\_\{y^\{\\star\}\}\)\>\\mathcal\{U\}\_\{\\Delta\}\(\\overline\{p\}\_\{y\}\),\\qquad\\forall y\\neq y^\{\\star\},\(4\)then, with probability at least1−α1\-\\alphaover Monte Carlo estimation,f^C\(𝐜~\)=y⋆\\hat\{f\}\_\{\\mathrm\{C\}\}\(\\tilde\{\\mathbf\{c\}\}\)=y^\{\\star\}for all‖𝒞\(𝐜~\)−𝒞\(𝐜\)‖0≤Δ\\\|\\mathcal\{C\}\(\\tilde\{\\mathbf\{c\}\}\)\-\\mathcal\{C\}\(\\mathbf\{c\}\)\\\|\_\{0\}\\leq\\Delta\.
We report the largest certified semantic radius
Δ⋆\(𝐜\)=max\{Δ:ℒΔ\(p¯y⋆\)\>𝒰Δ\(p¯y\),∀y≠y⋆\}\.\\Delta^\{\\star\}\(\\mathbf\{c\}\)=\\max\\left\\\{\\Delta:\\mathcal\{L\}\_\{\\Delta\}\(\\underline\{p\}\_\{y^\{\\star\}\}\)\>\\mathcal\{U\}\_\{\\Delta\}\(\\overline\{p\}\_\{y\}\),\\ \\forall y\\neq y^\{\\star\}\\right\\\}\.\(5\)All confidence bounds use Clopper–Pearson intervals; full transfer\-function definitions, proofs, and algorithms are in Appendix[B\.2](https://arxiv.org/html/2609.38625#A2.SS2)\.
## 4Experiments
In this section, we empirically and certifiably evaluate the robustness offstdf\_\{std\}andfcbmf\_\{cbm\}under controlled geometric and semantic perturbations\. Due to page limitations, we report a subset of the experimental results in the main paper; the complete results, including additional statistical analyses and backbone comparisons, are provided in[Appendix C](https://arxiv.org/html/2609.38625#A3)\. Across perturbation regimes, task conditions, and evaluation metrics, our experiments reveal a consistent pattern:*Concept bottlenecks do not induce monotonic robustness improvements; instead, they shift where and how sensitivity manifests\.*
### 4\.1Experimental Setup
We conduct experiments on two datasets with rich concept annotations: CUB\-200\-2011 and RIVAL\-10\. CUB\-200\-2011\([Wah et al\., 2011](https://arxiv.org/html/2609.38625#bib.bib42)\)is a fine\-grained bird classification benchmark consisting of 11,788 images from 200 species, each annotated with 112 expert\-defined visual concepts\.111[https://www\.vision\.caltech\.edu/datasets/cub\_200\_2011/](https://www.vision.caltech.edu/datasets/cub_200_2011/)RIVAL\-10\([Moayeri et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib30)\)aligns CIFAR\-10\([Krizhevsky et al\., 2009](https://arxiv.org/html/2609.38625#bib.bib24)\)classes with ImageNet\([Krizhevsky et al\., 2012](https://arxiv.org/html/2609.38625#bib.bib25)\)categories and contains approximately 26,000 images annotated with 18 visual attributes and corresponding segmentations\.
For CUB\-200\-2011, we selectK=15K=15concepts by default based on empirical frequency\. We adopt a post\-hoc concept bottleneck autoencoder \(CB\-AE\)\([Kulkarni et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib26)\)built on StyleGAN\([Karras et al\., 2020](https://arxiv.org/html/2609.38625#bib.bib21)\)as the generator, and train it on CUB\-200\-2011 to induce an explicit concept space aligned with these concepts\. For RIVAL\-10, models are trained on the official training split of 21,098 images and evaluated on the test split of 5,286 images, with all reported results computed on the test set\. We adopt BigGAN\([Brock et al\., 2018](https://arxiv.org/html/2609.38625#bib.bib1)\)as the base image generator and treat all 18 annotated attributes as concepts by default\. For classification, all compared classifiers share the same pretrained backbone to ensure a fair and controlled comparison\. For both datasets, we use a ResNet\-50\([He et al\., 2016](https://arxiv.org/html/2609.38625#bib.bib12)\)for the standard classifierfstdf\_\{std\}and a ResNet\-50 version of the post\-hoc concept bottleneck model \(PCBM\)\([Yuksekgonul et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib46)\)for the concept\-based classifierfcbmf\_\{cbm\}\.
We restrict evaluation to generated inputs on whichfstdf\_\{std\}andfcbmf\_\{cbm\}agree, denoted by𝒮agree\\mathcal\{S\}\_\{\\mathrm\{agree\}\}, and use the shared prediction as a reference label\. We use 100 images per run to balance computational cost with repeated\-subset statistical testing\.
#### 4\.1\.1Empirical robustness results
0\.050\.10\.30\.50\.8002020404060608080100100ϵ\\epsilonCUB PFR \(%\)↓\\downarrowGeometric perturbations001122⋅10−2\\cdot 10^\{\-2\}RIVAL PFR \(%\)↓\\downarrow1122551010002020404060608080100100τ\\tauCUB PFR \(%\)↓\\downarrowSemantic perturbations00224466RIVAL PFR \(%\)↓\\downarrowColor code:CUB\-200\-2011RIVAL\-10Line code:fstdf\_\{\\mathrm\{std\}\}fcbmf\_\{\\mathrm\{cbm\}\}Figure 2:Empirical robustness under geometric and semantic perturbations\.Orange \(left axis\): CUB\-200\-2011; blue \(right axis\): RIVAL\-10\. Dataset\-specific y\-axis scales are used for clarity\.11223344552020404060608080100100Decision Function Sensitivity \(DFS\)PFR %↓\\downarrow005510102020404060608080100100Concept Flip Rate \(CFR\)PFR %↓\\downarrowColor code:GeometricSemantic
Figure 3:Prediction sensitivity in concept\-level measures on CUB\-200\-2011\.To quantify variability due to subset selection, we repeat the evaluation on55independent random subsets of size 200 drawn uniformly without replacement from the pool \(subsets may overlap\)\. For each subset, bothfstdf\_\{\{std\}\}andfcbmf\_\{\{cbm\}\}are evaluated on the*same*inputs \(paired design\), yielding five paired measurements per metric\. We then testH0:Δ=0H\_\{0\}:\\Delta=0using a paired two\-sided test across the55subset\-level measurements, whereΔ\\Deltadenotes the per\-subset difference between methods\. When multiple budgets and metrics are tested, we additionally report Holm–Bonferroni correctedpp\-values in Appendix[C\.3](https://arxiv.org/html/2609.38625#A3.SS3)\.
As shown in[Figure 2](https://arxiv.org/html/2609.38625#S4.F2), the prediction flip rate \(PFR\) increases with perturbation strength\. On CUB\-200\-2011,fcbmf\_\{cbm\}exhibits slightly higher PFR thanfstdf\_\{std\}across most geometric and semantic budgets, indicating greater sensitivity\. In contrast, on RIVAL\-10, both models remain nearly perfectly robust over the same perturbation ranges\. As shown in[Figure 3](https://arxiv.org/html/2609.38625#S4.F3), raw concept flip rates \(CFR\) alone do not fully account for prediction instability\. In contrast, decision function sensitivity \(DFS\) shows a more consistent relationship with PFR, suggesting that instability in the CBM decision representation better explains prediction flips than the number of concept changes\. This arises because CBMs primarily rely on task\-relevant concepts and are relatively insensitive to perturbations of irrelevant ones, aligning with the intuition that concept bottlenecks can act as a form of regularization under structured semantic noise\. Additional analysis of DFS, CFR, and PFR, along with the complete numerical results, is provided in Appendix[C\.3](https://arxiv.org/html/2609.38625#A3.SS3.SSS0.Px2)\.
### 4\.2Randomised Smoothing Experiments
Empirical perturbations characterize sensitivity under sampled inputs but do not provide guarantees over all perturbations within a neighborhood\. To assess whether these trends persist under worst\-case perturbations, we evaluate certified robustness using randomized smoothing in latent \(ℓ2\\ell\_\{2\}\) and concept \(ℓ0\\ell\_\{0\}\) spaces\. Unless otherwise stated, we usen=200n=200inputs sampled from𝒮agree\\mathcal\{S\}\_\{\\mathrm\{agree\}\}and estimate smoothed class probabilities withN=1000N=1000Monte Carlo samples\. Confidence bounds are computed using one\-sided Clopper–Pearson intervals at level1−α=95%1\-\\alpha=95\\%\. All smoothed and certified accuracies are measured with respect to the reference labels defined by the agreement set\. To quantify variability, we repeat the evaluation over five independently sampled subsets of size200200\.
For latent\-space smoothing, we use Gaussian perturbationsδ∼𝒩\(0,σ2I\)\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)withσ∈\{0\.01,0\.05,0\.10,0\.30\}\\sigma\\in\\\{0\.01,0\.05,0\.10,0\.30\\\}and report the average certifiedℓ2\\ell\_\{2\}radiusRRover correctly classified, non abstained samples\. For concept space smoothing, each concept is independently flipped with probabilityρ\\rho, and we report the average certifiedℓ0\\ell\_\{0\}concept flip radiusΔ\\Deltaover correctly classified, non\-abstained samples\. In both settings, smoothed accuracy \(SmAcc\) counts abstentions as incorrect\.
[Figure 4](https://arxiv.org/html/2609.38625#S4.F4)summarizes the results\. Under latent\-space smoothing, thefstdf\_\{std\}generally achieves higher SmAcc and larger certified radii than thefcbmf\_\{cbm\}on CUB\-200\-2011, especially asσ\\sigmaincreases\. This indicates that, on generator native inputs,fcbmf\_\{cbm\}does not improve certifiable robustness to continuous latent perturbations\. On the original RIVAL\-10 split, both models are saturated: SmAcc remains at 100% and certified radii are nearly identical across allσ\\sigma, reflecting the strong class separability of this benchmark and indicating that samples remain far from decision boundaries in latent space\.
Under concept\-space smoothing, increasing the flip probabilityρ\\rhoreduces both SmAcc and the certifiedℓ0\\ell\_\{0\}radius, indicating thatstronger semantic randomization reduces the probability margin between the top class and competing classes\. On CUB\-200\-2011, thefcbmf\_\{cbm\}does not consistently improve certified concept robustness: although it slightly improves SmAcc for some noise levels, its certified radius is comparable to or smaller than that of the ResNet baseline\. On RIVAL\-10, the certificates are again saturated, motivating the harder RIVAL\-10 variants studied in[section 5](https://arxiv.org/html/2609.38625#S5)\.
0\.010\.050\.100\.3060608080100100σ\\sigmaSmAcc \(%\)↑\\uparrow\(a\) Latent0\.010\.050\.100\.30000\.20\.20\.40\.40\.60\.60\.80\.8σ\\sigmaRR↑\\uparrow\(b\) Latent5⋅10−25\\cdot 10^\{\-2\}0\.10\.10\.20\.2404060608080100100ρ\\rhoSmAcc \(%\)↑\\uparrow\(c\) Concept5⋅10−25\\cdot 10^\{\-2\}0\.10\.10\.20\.200551010ρ\\rhoΔ\\Delta↑\\uparrow\(d\) ConceptColor code:CUB\-200\-2011RIVAL\-10Line code:fstdf\_\{\\mathrm\{std\}\}fcbmf\_\{\\mathrm\{cbm\}\}
Figure 4:Certified robustness under latent\-space and concept\-space randomized smoothing\.
### 4\.3Adversarial Evaluation
We apply the APGD\-CE attack\([Croce and Hein, 2020](https://arxiv.org/html/2609.38625#bib.bib7)\)to the entire pipeline, including the generator, for a practical stress test of adversarial robustness rather than a head\-to\-head robustness comparison\. As shown in[Figure 5](https://arxiv.org/html/2609.38625#S4.F5), randomized smoothing consistently improves robust accuracy under both latent\-space and concept\-space attacks, with larger gains at higher attack strengths\. While[subsection 4\.2](https://arxiv.org/html/2609.38625#S4.SS2)indicates thatfcbmf\_\{cbm\}does not outperformfstdf\_\{std\}under randomized smoothing in either latent or concept space, adversarial evaluations reveal a complementary pattern:fcbmf\_\{cbm\}is more robust to latent\-space attacks but less robust to concept\-space attacks\. This reflects the dual role of concept bottlenecks as both a new attack interface and a regularizing mechanism against geometric perturbations, helping resolve prior conflicting conclusions about whether concept bottlenecks introduce robustness\.
latent spacezzconcept spacecc002020404060608080Robust Accuracy \(%\)↑\\uparrow\(a\)ϵ=0\.03\\epsilon=0\.03latent spacezzconcept spacecc002020404060608080Robust Accuracy \(%\)↑\\uparrow\(b\)ϵ=0\.01\\epsilon=0\.01Color code:VanillaSmoothedFill code:fstdf\_\{\\mathrm\{std\}\}fcbmf\_\{\\mathrm\{cbm\}\}Figure 5:Adversarial robustness under APGD\-CE\([Croce and Hein, 2020](https://arxiv.org/html/2609.38625#bib.bib7)\)attacks on the CUB generated dataset\.*Takeaway\.\(1\) Concept bottlenecks do not uniformly improve prediction\-level robustness, but they can offer increased robustness to large semantic perturbations by down\-weighting irrelevant concepts, acting as a form of regularization\. \(2\) Concept bottlenecks introduce a robustness trade\-off: they improve robustness against geometric adversarial perturbations while increasing vulnerability to concept\-space attacks\. Our unified framework explains previously contradictory observations and clarifies how future designs can better exploit the robustness trade\-off\.*
## 5Robustness under Task Conditions
The preceding sections show that concept bottleneck robustness varies substantially across datasets, despite using identical model architectures\. This suggests that robustness is shaped not only by the presence of a bottleneck, but also by task\-level factors\. To isolate these effects, we conduct controlled studies on RIVAL\-10 by varying \(i\) class structure and semantic similarity, and \(ii\) concept vocabulary size, while keeping the backbone and training protocol fixed\.
### 5\.1Effect of Class Structure and Semantic Similarity
0\.050\.050\.10\.10\.30\.30\.50\.50\.80\.8000\.50\.5111\.51\.522ϵ\\epsilonPFR \(%\)↓\\downarrow\(a\) Dissimilar0\.050\.050\.10\.10\.30\.30\.50\.50\.80\.8005510101515ϵ\\epsilonPFR \(%\)↓\\downarrow\(b\) Similar\+10 dissim\+5 dissim\+5 sim\+10 sim−4\-4−2\-20022Δ\\DeltaPFR↓\\downarrow\(c\)fcbm−fstdf\_\{cbm\}\-f\_\{std\}GeometricSemanticColor code:\+5 classes\+10 classesLine code:fstdf\_\{std\}fcbmf\_\{cbm\}Figure 6:Effect of class structure and semantic similarity on task\-conditional robustness under geometric perturbations\.\(a,b\) PFR*v\.s\.*ϵ\\epsilonfor dissimilar and similar class settings\. \(c\) Difference in PFR \(denoted asΔ\\DeltaPFR\) betweenfcbmf\_\{cbm\}andfstdf\_\{std\}whenϵ=0\.8\\epsilon=0\.8\. LowerΔ\\DeltaPFR favorsfcbmf\_\{cbm\}\.To study the effect of class similarity on robustness, we augment the original RIVAL\-10 classification task with additional classes designed to vary semantic proximity to the original categories\. Specifically, we construct two groups of new classes: similar classes, which are semantically close to the original RIVAL\-10 categories, and dissimilar classes, which are semantically distant\. Using these groups, we define four task configurations by adding 5 \(or 10\) similar \(or dissimilar\) classes, to the original 10\-class task, yielding classification problems of increasing size and semantic difficulty\. Across all configurations, we retain the original 18 concept attributes provided by RIVAL\-10, ensuring that only the class structure changes while the concept vocabulary remains fixed\. We then evaluate how robustness varies across these controlled settings\.
Across all settings, class semantic similarity emerges as the dominant factor shaping robustness differences betweenfstdf\_\{std\}andfcbmf\_\{cbm\}\.[Figure 6](https://arxiv.org/html/2609.38625#S5.F6)shows that semantic class dissimilarity generally improves robustness, with classifiers exhibiting lower prediction instability when classes are well separated\. Across configurations,fcbmf\_\{cbm\}tends to match or slightly outperformfstdf\_\{std\}under geometric perturbations and mild semantic changes, but becomes increasingly sensitive under large semantic perturbations, reflecting greater vulnerability at the concept level\. We observe consistent trends across both empirical robustness evaluations under geometric and semantic perturbations and certified robustness analyses in latent and concept spaces\. In addition, randomized smoothing yields larger robustness gains forfcbmf\_\{cbm\}in several settings, suggesting that smoothing can interact favorably with structured concept representations, albeit without providing a uniform robustness advantage\.
### 5\.2Effect of Concept Vocabulary Size
We study the effect of concept set size on robustness using the original RIVAL\-10 task\. Starting from the original set of 18 annotated concepts, we use a large language model to generate an additional 18 concepts that are semantically relevant to these classes\. This yields three concept configurations with increasing granularity: 27 and 36 concepts\. For each configuration, we evaluate robustness while keeping the classification task fixed, allowing us to isolate the impact of concept vocabulary size on model robustness\.
[Figure 7](https://arxiv.org/html/2609.38625#S5.F7)shows that changing the concept vocabulary affects the robustness behavior of bothfstdf\_\{\\mathrm\{std\}\}andfcbmf\_\{\\mathrm\{cbm\}\}, but the direction of the effect depends on the perturbation setting\. Under geometric perturbations, the original RIVAL\-10 setting is nearly saturated\. Thus, geometric perturbations provide little evidence that either model is more robust\. Semantic concept flips are more diagnostic\. With 27 concepts,fcbmf\_\{cbm\}is slightly more stable at the largest flip budget \(9\.299\.29vs\.9\.429\.42PFR\), whereas with 36 concepts the trend reverses \(3\.113\.11vs\.2\.652\.65PFR\)\. This suggests that the effect of the concept vocabulary is not monotonic\. One possible explanation is that, in the 27 concept setting, some flipped concepts are redundant or weakly used by thefcbmf\_\{cbm\}decision head, allowing the model to absorb large semantic flips slightly better\. In the 36 concept setting, however, the additional concepts may introduce more correlated or unevenly useful attributes; perturbing them can destabilize thefcbmf\_\{cbm\}concept representation without yielding a consistent semantic robustness benefit\. The randomized smoothing certified results support this interpretation\. In latent space, thefstdf\_\{std\}usually obtains larger certified radii thanfcbmf\_\{cbm\}across both 27 and 36 concept settings\. This is consistent with the fact that latent smoothing perturbs the generated image manifold continuously;fcbmf\_\{cbm\}must first recover a stable concept representation from each perturbed image, so instability in this intermediate representation can reduce the smoothed class margin\. In concept space, where perturbations are applied directly at the concept interface, the two models are much closer:fcbmf\_\{cbm\}is sometimes marginally better and sometimes marginally worse\. Overall, concept vocabulary size changes the measured robustness profile, but it does not induce a uniform robustness trend forfcbmf\_\{cbm\}relative to thefstdf\_\{std\}\.
Color code:K=27K=36Line code:fstdf\_\{std\}fcbmf\_\{cbm\}Color:K=27K=36Fill:fstdf\_\{\\mathrm\{std\}\}fcbmf\_\{\\mathrm\{cbm\}\}
0\.050\.100\.300\.500\.8000224466⋅10−2\\cdot 10^\{\-2\}ϵ\\epsilonPFR %↓\\downarrow\(a\) Geometric perturbation112255101000224466881010τ\\tauPFR %↓\\downarrow\(b\) Semantic perturbation0\.010\.050\.100\.30000\.50\.511σ\\sigmaRR↑\\uparrow\(c\) Certified in Latent space0\.050\.100\.20005510101515ρ\\rhoΔ⋆\\Delta^\{\\star\}↑\\uparrow\(d\) Certified in Concept space
Figure 7:Impact of concept vocabulary size on robustness\.Concept\-vocabulary ablation on RIVAL\-10\. \(a,b\) Empirical PFR under geometric latent perturbations and semantic concept flips\. \(c,d\) Certified robustness under Gaussian latent smoothing and Bernoulli concept smoothing\.*Takeaway\.\(1\) Lower inter\-class similarity is generally associated with higher classifier robustness\. \(2\) As class similarity increases, concept bottlenecks tend to remain more stable and exhibit a slower degradation in robustness\. \(3\) Increasing the number of concepts does not monotonically improve robustness: additional concepts can help when perturbations target redundant attributes, but may also introduce instability when concepts are correlated or weakly grounded\.*
##### Design Implications for Robust Concept Bottlenecks\.
These results indicate that robustness differences betweenfstdf\_\{std\}andfcbmf\_\{cbm\}are primarily governed by task structure\. This suggests that concept bottlenecks are particularly beneficial for classification tasks with high inter\-class similarity, provided that the concept vocabulary is chosen carefully\. Our findings further highlight that robustness should be treated as a key criterion when selecting and designing concept sets\. While concept bottlenecks reduce sensitivity to semantic perturbations and improve robustness to adversarial geometric perturbations, they are less robust to adversarial semantic attacks\. Consequently, improving the robustness of concept bottlenecks requires focusing on enhancing robustness within the concept space itself\.
## 6Limitation and Conclusion
##### Limitation\.
Our analysis relies on generator\-native inputs to disentangle geometric and semantic perturbations\. While this enables controlled robustness evaluation, results may differ for real\-world inputs outside the generator manifold\. We focus on post\-hoc CBMs with frozen backbones; end\-to\-end CBMs or alternative concept learning paradigms may exhibit different behaviors\. Finally, our conclusions are limited to robustness under structured geometric and semantic perturbations and do not replace standard pixel\-space adversarial robustness evaluations\.
##### Conclusion\.
We introduced a generator\-based framework that disentangles continuous latent/geometric perturbations from discrete concept/semantic interventions and used it to systematically evaluate robustness under empirical, certified, and adversarial criteria\. Across CUB\-200\-2011 and RIVAL\-10 variants, we find that concept bottlenecks do not confer uniform robustness improvements\. Instead, they reshape where and how sensitivity emerges, with robustness depending strongly on the perturbation space, robustness criterion, task structure, and concept vocabulary design\.
Within our controlled evaluation setting, these results clarify when concept bottlenecks suppress irrelevant variation and when they introduce additional sensitivity through the semantic interface\. In particular, class semantic similarity and concept granularity emerge as important design factors governing this trade\-off\. Our findings reinforce that interpretability and robustness are distinct objectives: concept bottlenecks should not be treated as an implicit robustness mechanism, but as a modeling choice whose robustness consequences depend on how concepts are defined and used\.
Finally, although CBMs provide our primary analytical setting, the additional PixPNet experiments show that the same generator\-defined perturbation\-to\-prediction evaluation and randomized\-smoothing certification pipeline can be applied to a prototype\-based interpretable classifier without requiring alignment between heterogeneous internal representations\. This broader applicability suggests a path toward evaluating robustness across a wider family of interpretable architectures while preserving the controlled perturbation framework developed here\.
##### Acknowledgment
This work received support from DFG under grant No\. 389792660 \(TRR 248\), and No\. 547583482\. This work was also supported by The Royal Society Grant \(Ensuring Trustworthy AI: Robustness Certification for Large Language Models\)\[Reference RGS\\R2\\252444\]\. GJ is supported by the University of Macau under grants MYRG\-SRG2026\-00033\-FIC\. This work was supported by the National Natural Science Foundation of China \(NSF\) \(T2422015, 62306212\) and the Beijing\-Tianjin\-Hebei Natural Science Foundation Cooperation Project \(25JJJJC0009\)\.
## References
- Brock et al\. \(2018\)Andrew Brock, Jeff Donahue, and Karen Simonyan\.Large scale gan training for high fidelity natural image synthesis\.*arXiv preprint arXiv:1809\.11096*, 2018\.
- Carlini et al\. \(2019\)Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin\.On evaluating adversarial robustness\.*arXiv preprint arXiv:1902\.06705*, 2019\.
- Carmichael et al\. \(2024\)Zachariah Carmichael, Suhas Lohit, Anoop Cherian, Michael J Jones, and Walter J Scheirer\.Pixel\-grounded prototypical part networks\.In*2024 IEEE/CVF Winter Conference on Applications of Computer Vision \(WACV\)*, pages 4756–4767\. IEEE, 2024\.
- Chauhan et al\. \(2023\)Kushal Chauhan, Rishabh Tiwari, Jan Freyberg, Pradeep Shenoy, and Krishnamurthy Dvijotham\.Interactive concept bottleneck models\.In*Proceedings of the aaai conference on artificial intelligence*, volume 37, pages 5948–5955, 2023\.
- Chen et al\. \(2018\)Ricky TQ Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud\.Isolating sources of disentanglement in variational autoencoders\.*Advances in neural information processing systems*, 31, 2018\.
- Cohen et al\. \(2019\)Jeremy Cohen, Elan Rosenfeld, and Zico Kolter\.Certified adversarial robustness via randomized smoothing\.In*international conference on machine learning*, pages 1310–1320\. PMLR, 2019\.
- Croce and Hein \(2020\)Francesco Croce and Matthias Hein\.Reliable evaluation of adversarial robustness with an ensemble of diverse parameter\-free attacks\.In*International conference on machine learning*, pages 2206–2216\. PMLR, 2020\.
- Ding et al\. \(2020\)Zheng Ding, Yifan Xu, Weijian Xu, Gaurav Parmar, Yang Yang, Max Welling, and Zhuowen Tu\.Guided variational autoencoder for disentanglement learning\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pages 7920–7929, 2020\.
- Du et al\. \(2026\)Xingbo Du, Qiantong Dou, Lei Fan, and Rui Zhang\.Flexible concept bottleneck model\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 40, pages 20932–20940, 2026\.
- Hao et al\. \(2022\)Zhongkai Hao, Chengyang Ying, Yinpeng Dong, Hang Su, Jian Song, and Jun Zhu\.Gsmooth: Certified robustness against semantic transformations via generalized randomized smoothing\.In*International Conference on Machine Learning*, pages 8465–8483\. PMLR, 2022\.
- Havasi et al\. \(2022\)Marton Havasi, Sonali Parbhoo, and Finale Doshi\-Velez\.Addressing leakage in concept bottleneck models\.*Advances in Neural Information Processing Systems*, 35:23386–23397, 2022\.
- He et al\. \(2016\)Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun\.Deep residual learning for image recognition\.In*Proceedings of the IEEE conference on computer vision and pattern recognition*, pages 770–778, 2016\.
- Hermanns et al\. \(2024\)Holger Hermanns, Anne Lauber\-Rönsberg, Philip Meinel, Sarah Sterz, and Hanwei Zhang\.Ai act for the working programmer\.In*International Conference on Bridging the Gap between AI and Reality*, pages 74–98\. Springer, 2024\.
- Higgins et al\. \(2017\)Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner\.beta\-vae: Learning basic visual concepts with a constrained variational framework\.In*International conference on learning representations*, 2017\.
- Hu et al\. \(2024\)Lijie Hu, Chenyang Ren, Zhengyu Hu, Hongbin Lin, Cheng\-Long Wang, Hui Xiong, Jingfeng Zhang, and Di Wang\.Editable concept bottleneck models\.*arXiv preprint arXiv:2405\.15476*, 2024\.
- Huang et al\. \(2017\)Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger\.Densely connected convolutional networks\.In*Proceedings of the IEEE conference on computer vision and pattern recognition*, pages 4700–4708, 2017\.
- Huang et al\. \(2020\)Xiaowei Huang, Daniel Kroening, Wenjie Ruan, James Sharp, Youcheng Sun, Emese Thamo, Min Wu, and Xinping Yi\.A survey of safety and trustworthiness of deep neural networks: Verification, testing, adversarial attack and defence, and interpretability\.*Computer Science Review*, 37:100270, 2020\.
- Ismail et al\. \(2023\)Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho\.Concept bottleneck generative models\.In*The Twelfth International Conference on Learning Representations*, 2023\.
- Jin et al\. \(2022\)Gaojie Jin, Xinping Yi, Wei Huang, Sven Schewe, and Xiaowei Huang\.Enhancing adversarial training with second\-order statistics of weights\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pages 15273–15283, 2022\.
- Jin et al\. \(2025\)Gaojie Jin, Sihao Wu, Jiaxu Liu, Tianjin Huang, and Ronghui Mu\.Enhancing robust fairness via confusional spectral regularization\.In*International Conference on Learning Representations*, volume 2025, pages 52945–52966, 2025\.
- Karras et al\. \(2020\)Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila\.Analyzing and improving the image quality of stylegan\.In*Proceedings of the IEEE/CVF conference on computer vision and pattern recognition*, pages 8110–8119, 2020\.
- Kim et al\. \(2023\)Eunji Kim, Dahuin Jung, Sangha Park, Siwon Kim, and Sungroh Yoon\.Probabilistic concept bottleneck models\.*arXiv preprint arXiv:2306\.01574*, 2023\.
- Koh et al\. \(2020\)Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang\.Concept bottleneck models\.In*International conference on machine learning*, pages 5338–5348\. PMLR, 2020\.
- Krizhevsky et al\. \(2009\)Alex Krizhevsky, Geoffrey Hinton, et al\.Learning multiple layers of features from tiny images\.2009\.
- Krizhevsky et al\. \(2012\)Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton\.Imagenet classification with deep convolutional neural networks\.*Advances in neural information processing systems*, 25, 2012\.
- Kulkarni et al\. \(2025\)Akshay Kulkarni, Ge Yan, Chung\-En Sun, Tuomas Oikarinen, and Tsui\-Wei Weng\.Interpretable generative models through post\-hoc concept bottlenecks\.In*Proceedings of the Computer Vision and Pattern Recognition Conference*, pages 8162–8171, 2025\.
- Lee et al\. \(2019\)Guang\-He Lee, Yang Yuan, Shiyu Chang, and Tommi Jaakkola\.Tight certificates of adversarial robustness for randomly smoothed classifiers\.*Advances in Neural Information Processing Systems*, 32, 2019\.
- Li et al\. \(2026\)Aaron J Li, Suraj Srinivas, Usha Bhalla, and Himabindu Lakkaraju\.Evaluating adversarial robustness of concept representations in sparse autoencoders\.In*Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\)*, pages 5940–5957, 2026\.
- Meng et al\. \(2022\)Mark Huasong Meng, Guangdong Bai, Sin Gee Teo, Zhe Hou, Yan Xiao, Yun Lin, and Jin Song Dong\.Adversarial robustness of deep neural networks: A survey from a formal verification perspective\.*IEEE Transactions on Dependable and Secure Computing*, 2022\.
- Moayeri et al\. \(2022\)Mazda Moayeri, Phillip Pope, Yogesh Balaji, and Soheil Feizi\.A comprehensive study of image classification model sensitivity to foregrounds, backgrounds, and visual attributes\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pages 19087–19097, 2022\.
- Mu et al\. \(2023\)Ronghui Mu, Wenjie Ruan, Leandro Soriano Marcolino, Gaojie Jin, and Qiang Ni\.Certified policy smoothing for cooperative multi\-agent reinforcement learning\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 37, pages 15046–15054, 2023\.
- Mu et al\. \(2024\)Ronghui Mu, Leandro Soriano Marcolino, Yanghao Zhang, Tianle Zhang, Xiaowei Huang, and Wenjie Ruan\.Reward certification for policy smoothed reinforcement learning\.In*Proceedings of the AAAI conference on artificial intelligence*, volume 38, pages 21429–21437, 2024\.
- Parliament and of the EU \(2024\)European Parliament and Council of the EU\.Regulation \(eu\) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence and amending regulations \(ec\) no 300/2008, \(eu\) no 167/2013, \(eu\) no 168/2013, \(eu\) 2018/858, \(eu\) 2018/1139 and \(eu\) 2019/2144 and directives 2014/90/eu, \(eu\) 2016/797 and \(eu\) 2020/1828 \(artificial intelligence act\), 2024\.URL[https://eur\-lex\.europa\.eu/legal\-content/EN/TXT/?uri=OJ:L\_202401689](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=OJ:L_202401689)\.
- Rasheed et al\. \(2024\)Bader Rasheed, Mohamed Abdelhamid, Adil Khan, Igor Menezes, and Asad Masood Khatak\.Exploring the impact of conceptual bottlenecks on adversarial robustness of deep neural networks\.*IEEE Access*, 12:131323–131335, 2024\.
- Sawada and Nakamura \(2022\)Yoshihide Sawada and Keigo Nakamura\.C\-senn: Contrastive self\-explaining neural network\.*arXiv preprint arXiv:2206\.09575*, 2022\.
- Shang et al\. \(2024\)Chenming Shang, Shiji Zhou, Hengyuan Zhang, Xinzhe Ni, Yujiu Yang, and Yuwang Wang\.Incremental residual concept bottleneck models\.In*Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition*, pages 11030–11040, 2024\.
- Sicre et al\. \(2023\)Ronan Sicre, Hanwei Zhang, Julien Dejasmin, Chiheb Daaloul, Stéphane Ayache, and Thierry Artières\.Dp\-net: learning discriminative parts for image recognition\.In*2023 IEEE international conference on image processing \(ICIP\)*, pages 1230–1234\. IEEE, 2023\.
- Sinha et al\. \(2023\)Sanchit Sinha, Mengdi Huai, Jianhui Sun, and Aidong Zhang\.Understanding and enhancing robustness of concept\-based models\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 37, pages 15127–15135, 2023\.
- Szegedy et al\. \(2013\)Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus\.Intriguing properties of neural networks\.*arXiv preprint arXiv:1312\.6199*, 2013\.
- Vandenhirtz et al\. \(2024\)Moritz Vandenhirtz, Sonia Laguna, Ričards Marcinkevičs, and Julia Vogt\.Stochastic concept bottleneck models\.*Advances in Neural Information Processing Systems*, 37:51787–51810, 2024\.
- Wagh et al\. \(2022\)Neeraj Wagh, Jionghao Wei, Samarth Rawal, Brent M Berry, and Yogatheesan Varatharajah\.Evaluating latent space robustness and uncertainty of eeg\-ml models under realistic distribution shifts\.*Advances in Neural Information Processing Systems*, 35:21142–21156, 2022\.
- Wah et al\. \(2011\)C\. Wah, S\. Branson, P\. Welinder, P\. Perona, and S\. Belongie\.The caltech\-ucsd birds\-200\-2011 dataset\.Technical Report CNS\-TR\-2011\-001, California Institute of Technology, 2011\.
- Wang et al\. \(2021\)Binghui Wang, Jinyuan Jia, Xiaoyu Cao, and Neil Zhenqiang Gong\.Certified robustness of graph neural networks against adversarial structural perturbation\.In*Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining*, pages 1645–1653, 2021\.
- Wang et al\. \(2026\)Zixia Wang, Gaojie Jin, Jia Hu, and Ronghui Mu\.Clucert: Certifying llm robustness via clustering\-guided denoising smoothing\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 40, pages 37998–38006, 2026\.
- Yu et al\. \(2025\)Lu Yu, Haoyu Han, Zhe Tao, Hantao Yao, and Changsheng Xu\.Language guided concept bottleneck models for interpretable continual learning\.In*Proceedings of the Computer Vision and Pattern Recognition Conference*, pages 14976–14986, 2025\.
- Yuksekgonul et al\. \(2022\)Mert Yuksekgonul, Maggie Wang, and James Zou\.Post\-hoc concept bottleneck models\.*arXiv preprint arXiv:2205\.15480*, 2022\.
- Zhang et al\. \(2026\)Hanwei Zhang, Luo Cheng, Rui Wen, Yang Zhang, Lijun Zhang, and Holger Hermanns\.Sl\-cbm: Enhancing concept bottleneck models with semantic locality for better interpretability, 2026\.URL[https://arxiv\.org/abs/2601\.12804](https://arxiv.org/abs/2601.12804)\.
- Zhang et al\. \(2021\)Xing\-Xing Zhang, Zhen\-Feng Zhu, Ya\-Wei Zhao, and Yao Zhao\.Prototype learning in machine learning: A literature review\.*Journal of Software*, 33\(10\):3732–3753, 2021\.
To support reproducibility and completeness, the supplementary material provides extended background and related work \([Appendix A](https://arxiv.org/html/2609.38625#A1)\), additional methodological details \([Appendix B](https://arxiv.org/html/2609.38625#A2)\), including full randomized smoothing proofs and algorithms, comprehensive experimental results \([Appendix C](https://arxiv.org/html/2609.38625#A3)\), and cross\-family applicability experiments \([Appendix D](https://arxiv.org/html/2609.38625#A4)\)\.
## Appendix ABackground and Related Work
##### Concept Bottleneck Models\.
Concept Bottleneck Models \(CBMs\) were introduced by[Koh et al\. \[2020\]](https://arxiv.org/html/2609.38625#bib.bib23)as a framework that decomposes prediction into concept extraction followed by concept\-based decision making\. Building on this idea, subsequent work has explored CBMs along several directions\. Post\-hoc CBMs \(PCBM\) leverage frozen vision or multimodal backbones and add lightweight projection layers to obtain concept representations\[[Yuksekgonul et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib46)\]\. Other approaches focus on improving the quality, uncertainty, or robustness of learned concepts\[[Vandenhirtz et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib40),[Kim et al\., 2023](https://arxiv.org/html/2609.38625#bib.bib22),[Havasi et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib11),[Zhang et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib47)\], while additional work aims to increase the flexibility and adaptability of concept representations through interactive or editable mechanisms\[[Chauhan et al\., 2023](https://arxiv.org/html/2609.38625#bib.bib4),[Hu et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib15),[Shang et al\., 2024](https://arxiv.org/html/2609.38625#bib.bib36)\]\.
##### Robustness of Concept Bottlenecks\.
Early work on interpretable models highlighted connections between interpretability and robustness, with Self\-Explaining Neural Networks \(SENN\) enforcing robustness as the stability of explanations under small perturbations\[[Sawada and Nakamura, 2022](https://arxiv.org/html/2609.38625#bib.bib35)\]\. More recent studies directly examine robustness in Concept Bottleneck Models \(CBMs\)\.[Sinha et al\. \[2023\]](https://arxiv.org/html/2609.38625#bib.bib38)show that interpretability alone does not guarantee robustness, demonstrating that concept representations can be fragile under adversarial perturbations and that explicit concept interfaces introduce new attack surfaces, revealing a tension between concept\-level stability and adversarial robustness\. In contrast,[Rasheed et al\. \[2024\]](https://arxiv.org/html/2609.38625#bib.bib34)find that CBMs, particularly sequentially trained variants, can achieve higher adversarial accuracy than end\-to\-end models under weak to moderate attacks, suggesting that conceptual bottlenecks may act as an information\-filtering regularizer at the prediction level\. Related observations beyond CBMs further indicate that highly interpretable representations can remain non\-robust\[[Li et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib28)\], while recent extensions such as language\-guided and flexible CBMs improve scalability and adaptability without addressing robustness explicitly\[[Yu et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib45),[Du et al\., 2026](https://arxiv.org/html/2609.38625#bib.bib9)\]\. Overall, prior work reports seemingly conflicting conclusions, which we argue arise from differences in robustness definitions, perturbation models, and task settings rather than fundamental inconsistencies\.
##### Concept\-Based Generative Models\.
Prior work explores disentangled or concept\-based generative models for controllable generation\[[Chen et al\., 2018](https://arxiv.org/html/2609.38625#bib.bib5),[Ding et al\., 2020](https://arxiv.org/html/2609.38625#bib.bib8),[Higgins et al\., 2017](https://arxiv.org/html/2609.38625#bib.bib14),[Ismail et al\., 2023](https://arxiv.org/html/2609.38625#bib.bib18)\], while CB\-AE\[[Kulkarni et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib26)\]enables concept\-level control by augmenting frozen pretrained generators with minimal supervision\. In this work, we use CB\-AE\[[Kulkarni et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib26)\]solely as an experimental instrument to enable controlled geometric and semantic perturbations\.
##### Robustness Beyond Pixel Space\.
Deep neural networks are known to be highly sensitive to small input perturbations, motivating extensive work on adversarial robustness and robustness evaluation\[[Szegedy et al\., 2013](https://arxiv.org/html/2609.38625#bib.bib39),[Carlini et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib2),[Meng et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib29),[Huang et al\., 2020](https://arxiv.org/html/2609.38625#bib.bib17)\]\. However, most studies focus on pixel\-level perturbations, which often lack semantic meaning and do not reflect changes in underlying factors of variation\. Recent work has therefore explored robustness beyond the pixel space, including robustness to perturbations in learned representations or structured transformations\[[Hao et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib10),[Wang et al\., 2021](https://arxiv.org/html/2609.38625#bib.bib43),[Wagh et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib41),[Jin et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib19),[Jin et al\., 2025](https://arxiv.org/html/2609.38625#bib.bib20)\]\. Despite this progress, how architectural choices such as concept\-aligned or interpretable representations affect robustness under semantic perturbations remains underexplored, motivating our focus on concept\-level robustness\.
##### Randomized Smoothing for Certified Robustness\.
Randomized smoothing\[[Cohen et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib6)\]was developed to evaluate probabilistic certified robustness for classification tasks\. By constructing a smoothed classifierf^\(x\)\\hat\{f\}\(x\)through aggregation of predictions over randomized perturbations, it can produce the most probable prediction of the base classifierf\(x\)f\(x\)over perturbed inputs from Gaussian noise in a test instance\. The smoothed classifierf^\(x\)\\hat\{f\}\(x\)is supposed to be provably robust tol2l\_\{2\}\-norm bounded perturbations within a certain radiusRR:
###### Theorem A\.1\.
\[[Cohen et al\., 2019](https://arxiv.org/html/2609.38625#bib.bib6)\]For a classifierf:ℝ→𝒴f:\\mathbb\{R\}\\to\\mathcal\{Y\}, supposey∈𝒴y\\in\\mathcal\{Y\}, letδ∼𝒩\(0,σ2I\)\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\. Suppose the top classAAis predicted with probabilitypAp\_\{A\}, and the runner\-up prediceted classBB, with probabilitypBp\_\{B\}\. The smoothed classifier bef^\(𝐱\):=argmaxyℙ\(f\(𝐱\+δ\)=y\)\\hat\{f\}\(\\mathbf\{x\}\):=\\mathop\{\\arg\\max\}\\limits\_\{y\}\\mathbb\{P\}\(f\(\\mathbf\{x\}\+\\delta\)=y\), supposepA¯,pB¯∈\[0,1\]\\underline\{p\_\{A\}\},\\overline\{p\_\{B\}\}\\in\[0,1\], if
ℙ\(f\(𝐱\+δ\)=yA\)≥pA¯≥pB¯≥maxy≠yAℙ\(f\(𝐱\+δ\)=y\),\\mathbb\{P\}\(f\(\\mathbf\{x\}\+\\delta\)=y\_\{A\}\)\\geq\\underline\{p\_\{A\}\}\\geq\\overline\{p\_\{B\}\}\\geq\\mathop\{\\max\}\\limits\_\{y\\neq y\_\{A\}\}\\mathbb\{P\}\(f\(\\mathbf\{x\}\+\\delta\)=y\),\(6\)thenf^\(𝐱\+ϵ\)=yA\\hat\{f\}\(\\mathbf\{x\}\+\\epsilon\)=y\_\{A\}for all‖ϵ‖2≤R\|\|\\epsilon\|\|\_\{2\}\\leq R, where
R=σ2\(Φ−1\(pA¯\)−Φ−1\(pB¯\)\)\.R=\\frac\{\\sigma\}\{2\}\(\\Phi^\{\-1\}\(\\underline\{p\_\{A\}\}\)\-\\Phi^\{\-1\}\(\\overline\{p\_\{B\}\}\)\)\.\(7\)
HereΦ−1\\Phi^\{\-1\}is the inverse cumulative distribution function \(CDF\) of the normal distribution\.
This approach has been applied primarily in pixel space to certify robustness againstℓp\\ell\_\{p\}\-bounded adversarial perturbations, and has been extended to different noise distributions and norm settings\. While most work applies smoothing in pixel space, recent studies have explored extensions to learned representations or structured perturbations\[[Hao et al\., 2022](https://arxiv.org/html/2609.38625#bib.bib10)\]\. In this work, we use randomized smoothing not as a defense mechanism, but as a comparative evaluation tool to assess certified robustness of concept\-based and standard classifiers under controlled geometric and semantic perturbations\.
## Appendix BMethodology
### B\.1Empirical Robustness Characterization
As illustrated in[Figure 8](https://arxiv.org/html/2609.38625#A2.F8), we introduce a generator\-based evaluation framework that disentangles continuous latent\-space perturbations from discrete concept\-level interventions\. In the latent space, geometric perturbations are applied by modifying the corresponding latent vector𝐳\\mathbf\{z\}, while in the concept space, semantic interventions are performed by flipping entries in the concept representation𝐜\\mathbf\{c\}\. These operations yield controlled sets of generated images reflecting geometric and semantic perturbations, which we use to systematically analyze the robustness of classifiers with and without concept bottlenecks\.
Figure 8:Generator\-based evaluation framework for geometric vs\. semantic perturbations\.\(i\) geometric perturbations by adding latent\-space noise within anℓ2\\ell\_\{2\}budgetϵ\\epsilon, and \(ii\) semantic perturbations by flipping up toτ\\taubinary concepts in concept space before image generation\.
### B\.2Certified Robustness via Random Smoothing
### B\.3Latent\-Space Gaussian Smoothing
Given a latent representation𝐳\\mathbf\{z\}, the latent\-smoothed classifier is
f^G\(𝐳\)=argmaxy∈𝒴ℙδ∼𝒩\(0,σ2I\)\[f\(g\(𝐳\+δ\)\)=y\]\.\\hat\{f\}\_\{\\mathrm\{G\}\}\(\\mathbf\{z\}\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\}\\mathbb\{P\}\_\{\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\}\\left\[f\(g\(\\mathbf\{z\}\+\\delta\)\)=y\\right\]\.\(8\)LetAAandBBdenote the top and runner\-up classes with probabilities
pA=ℙδ\[f\(g\(𝐳\+δ\)\)=A\],pB=ℙδ\[f\(g\(𝐳\+δ\)\)=B\]\.p\_\{A\}=\\mathbb\{P\}\_\{\\delta\}\\left\[f\(g\(\\mathbf\{z\}\+\\delta\)\)=A\\right\],\\qquad p\_\{B\}=\\mathbb\{P\}\_\{\\delta\}\\left\[f\(g\(\\mathbf\{z\}\+\\delta\)\)=B\\right\]\.\(9\)By the Gaussian randomized smoothing certificate of[Cohen et al\. \[2019\]](https://arxiv.org/html/2609.38625#bib.bib6), ifpA\>pBp\_\{A\}\>p\_\{B\}, thenf^G\(𝐳\+ϵ\)=A\\hat\{f\}\_\{\\mathrm\{G\}\}\(\\mathbf\{z\}\+\\epsilon\)=Afor all
‖ϵ‖2≤σ2\(Φ−1\(pA\)−Φ−1\(pB\)\)\.\\\|\\epsilon\\\|\_\{2\}\\leq\\frac\{\\sigma\}\{2\}\\left\(\\Phi^\{\-1\}\(p\_\{A\}\)\-\\Phi^\{\-1\}\(p\_\{B\}\)\\right\)\.\(10\)
In practice,pAp\_\{A\}andpBp\_\{B\}are estimated by Monte Carlo sampling\. GivenNNi\.i\.d\. samplesδ1,…,δN∼𝒩\(0,σ2I\)\\delta\_\{1\},\\ldots,\\delta\_\{N\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\), we estimate
p^y\(𝐳\)=1N∑i=1N𝕀\[f\(g\(𝐳\+δi\)\)=y\]\.\\widehat\{p\}\_\{y\}\(\\mathbf\{z\}\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbb\{I\}\\left\[f\(g\(\\mathbf\{z\}\+\\delta\_\{i\}\)\)=y\\right\]\.\(11\)We compute one\-sided Clopper–Pearson confidence bounds at level1−α1\-\\alpha: a lower boundp¯A\\underline\{p\}\_\{A\}for the top class and an upper boundp¯B\\overline\{p\}\_\{B\}for the runner\-up class\. The empirical certified radius is
R^G=σ2\(Φ−1\(p¯A\)−Φ−1\(p¯B\)\)\.\\widehat\{R\}\_\{\\mathrm\{G\}\}=\\frac\{\\sigma\}\{2\}\\left\(\\Phi^\{\-1\}\(\\underline\{p\}\_\{A\}\)\-\\Phi^\{\-1\}\(\\overline\{p\}\_\{B\}\)\\right\)\.\(12\)The certificate holds with probability at least1−α1\-\\alphaover the sampling procedure\.
### B\.4Concept\-Space Bernoulli Flip Smoothing
Let𝐜∈ℝK\\mathbf\{c\}\\in\\mathbb\{R\}^\{K\}denote a concept representation and let𝒞\(𝐜\)∈\{0,1\}K\\mathcal\{C\}\(\\mathbf\{c\}\)\\in\\\{0,1\\\}^\{K\}be its binarized concept vector\. For paired concept coordinates\[ci\+,ci−\]\[c\_\{i\}^\{\+\},c\_\{i\}^\{\-\}\], flipping theii\-th concept swaps\[ci\+,ci−\]\[c\_\{i\}^\{\+\},c\_\{i\}^\{\-\}\]into\[ci−,ci\+\]\[c\_\{i\}^\{\-\},c\_\{i\}^\{\+\}\]\.
We define Bernoulli concept\-flip noise by sampling
ηi∼Bernoulli\(ρ\),i=1,…,K,\\eta\_\{i\}\\sim\\mathrm\{Bernoulli\}\(\\rho\),\\qquad i=1,\\ldots,K,\(13\)independently across concepts\. Letℱ\(𝐜,η\)\\mathcal\{F\}\(\\mathbf\{c\},\\eta\)denote the concept representation obtained after flipping all coordinates withηi=1\\eta\_\{i\}=1\. The concept\-smoothed class probabilities are
py\(𝐜\)=ℙη∼Bernoulli\(ρ\)K\[f\(g2\(D\(ℱ\(𝐜,η\)\)\)\)=y\],p\_\{y\}\(\\mathbf\{c\}\)=\\mathbb\{P\}\_\{\\eta\\sim\\mathrm\{Bernoulli\}\(\\rho\)^\{K\}\}\\left\[f\\\!\\left\(g\_\{2\}\(D\(\\mathcal\{F\}\(\\mathbf\{c\},\\eta\)\)\)\\right\)=y\\right\],\(14\)and the concept\-smoothed classifier is
f^C\(𝐜\)=argmaxy∈𝒴py\(𝐜\)\.\\hat\{f\}\_\{\\mathrm\{C\}\}\(\\mathbf\{c\}\)=\\arg\\max\_\{y\\in\\mathcal\{Y\}\}p\_\{y\}\(\\mathbf\{c\}\)\.\(15\)
We certify robustness to adversarial concept flips satisfying
‖𝒞\(𝐜~\)−𝒞\(𝐜\)‖0≤Δ\.\\\|\\mathcal\{C\}\(\\tilde\{\\mathbf\{c\}\}\)\-\\mathcal\{C\}\(\\mathbf\{c\}\)\\\|\_\{0\}\\leq\\Delta\.\(16\)Because randomized concept flipping modifies semantic values rather than masking them, sparse\-ablation certificates do not directly apply\. We instead use the Neyman–Pearson smoothing framework of[Lee et al\. \[2019\]](https://arxiv.org/html/2609.38625#bib.bib27)on the Hamming cube\.
For a candidate adversarial budgetΔ\\Delta, letℒΔ\\mathcal\{L\}\_\{\\Delta\}and𝒰Δ\\mathcal\{U\}\_\{\\Delta\}denote the Neyman–Pearson transfer functions induced by the Bernoulli smoothing distribution\. Intuitively,ℒΔ\(p\)\\mathcal\{L\}\_\{\\Delta\}\(p\)gives the smallest possible probability of an event after anyΔ\\Delta\-bit adversarial shift, assuming that the event has clean smoothing probability at leastpp\. Similarly,𝒰Δ\(p\)\\mathcal\{U\}\_\{\\Delta\}\(p\)gives the largest possible perturbed probability of an event whose clean smoothing probability is at mostpp\.
Lety⋆=argmaxypy\(𝐜\)y^\{\\star\}=\\arg\\max\_\{y\}p\_\{y\}\(\\mathbf\{c\}\)be the top class\. Suppose conservative confidence bounds satisfy
py⋆\(𝐜\)≥p¯y⋆\(𝐜\),py\(𝐜\)≤p¯y\(𝐜\),∀y≠y⋆\.p\_\{y^\{\\star\}\}\(\\mathbf\{c\}\)\\geq\\underline\{p\}\_\{y^\{\\star\}\}\(\\mathbf\{c\}\),\\qquad p\_\{y\}\(\\mathbf\{c\}\)\\leq\\overline\{p\}\_\{y\}\(\\mathbf\{c\}\),\\qquad\\forall y\\neq y^\{\\star\}\.\(17\)If
ℒΔ\(p¯y⋆\(𝐜\)\)\>𝒰Δ\(p¯y\(𝐜\)\),∀y≠y⋆,\\mathcal\{L\}\_\{\\Delta\}\\left\(\\underline\{p\}\_\{y^\{\\star\}\}\(\\mathbf\{c\}\)\\right\)\>\\mathcal\{U\}\_\{\\Delta\}\\left\(\\overline\{p\}\_\{y\}\(\\mathbf\{c\}\)\\right\),\\qquad\\forall y\\neq y^\{\\star\},\(18\)then with probability at least1−α1\-\\alphaover Monte Carlo estimation,
f^C\(𝐜~\)=y⋆,∀𝐜~:‖𝒞\(𝐜~\)−𝒞\(𝐜\)‖0≤Δ\.\\hat\{f\}\_\{\\mathrm\{C\}\}\(\\tilde\{\\mathbf\{c\}\}\)=y^\{\\star\},\\qquad\\forall\\tilde\{\\mathbf\{c\}\}:\\\|\\mathcal\{C\}\(\\tilde\{\\mathbf\{c\}\}\)\-\\mathcal\{C\}\(\\mathbf\{c\}\)\\\|\_\{0\}\\leq\\Delta\.\(19\)
##### Certified concept\-flip radius\.
The certified adversarial concept\-flip radius is
Δ⋆\(𝐜\)=max\{Δ∈\{0,1,…,K\}:\\displaystyle\\Delta^\{\\star\}\(\\mathbf\{c\}\)=\\max\\Big\\\{\\Delta\\in\\\{0,1,\\ldots,K\\\}:ℒΔ\(p¯y⋆\(𝐜\)\)\>𝒰Δ\(p¯y\(𝐜\)\),\\displaystyle\\mathcal\{L\}\_\{\\Delta\}\\left\(\\underline\{p\}\_\{y^\{\\star\}\}\(\\mathbf\{c\}\)\\right\)\>\\mathcal\{U\}\_\{\\Delta\}\\left\(\\overline\{p\}\_\{y\}\(\\mathbf\{c\}\)\\right\),∀y∈𝒴∖\{y⋆\}\}\.\\displaystyle\\forall y\\in\\mathcal\{Y\}\\setminus\\\{y^\{\\star\}\\\}\\Big\\\}\.\(20\)
### B\.5Monte Carlo Estimation and Certification Algorithm
For concept\-space smoothing, we estimate the smoothed probabilities usingNNi\.i\.d\. Bernoulli flip samples\. Abstentions occur when the top\-class lower confidence bound does not exceed the runner\-up upper confidence bound\. The procedure for estimating the smoothed class probabilities is presented in Algorithm[1](https://arxiv.org/html/2609.38625#alg1)\.
Algorithm 1Concept\-spaceℓ0\\ell\_\{0\}certification via Bernoulli concept flips0:Concept representation
𝐜\\mathbf\{c\}, base classifier
ff, generator mapping
g2\(D\(⋅\)\)g\_\{2\}\(D\(\\cdot\)\), flip probability
ρ\\rho, samples
NN, confidence level
1−α1\-\\alpha\.
1:for
i=1i=1to
NNdo
2:Sample
ηi∼Bernoulli\(ρ\)K\\eta\_\{i\}\\sim\\mathrm\{Bernoulli\}\(\\rho\)^\{K\}\.
3:Compute
𝐜i′=ℱ\(𝐜,ηi\)\\mathbf\{c\}\_\{i\}^\{\\prime\}=\\mathcal\{F\}\(\\mathbf\{c\},\\eta\_\{i\}\)\.
4:Compute
y^i=f\(g2\(D\(𝐜i′\)\)\)\\hat\{y\}\_\{i\}=f\(g\_\{2\}\(D\(\\mathbf\{c\}\_\{i\}^\{\\prime\}\)\)\)\.
5:endfor
6:Compute class counts
count\[y\]=\|\{i:y^i=y\}\|\\mathrm\{count\}\[y\]=\|\\\{i:\\hat\{y\}\_\{i\}=y\\\}\|for all
y∈𝒴y\\in\\mathcal\{Y\}\.
7:Compute Clopper–Pearson confidence bounds
p¯y\(𝐜\)\\underline\{p\}\_\{y\}\(\\mathbf\{c\}\)and
p¯y\(𝐜\)\\overline\{p\}\_\{y\}\(\\mathbf\{c\}\)from
count\[y\]\\mathrm\{count\}\[y\]\.
8:Let
y⋆=argmaxyp¯y\(𝐜\)y^\{\\star\}=\\arg\\max\_\{y\}\\underline\{p\}\_\{y\}\(\\mathbf\{c\}\)\.
9:Return
Δ⋆\(𝐜\)\\Delta^\{\\star\}\(\\mathbf\{c\}\)using Eq\. \([20](https://arxiv.org/html/2609.38625#A2.E20)\)\.
## Appendix CExperiments
### C\.1Computation of the robustness\-gap heatmap\.
Each entry in Figure 1\(b\) is computed from the corresponding representative experimental setting in the main results\. For empirical geometric robustness we useϵ=0\.3\\epsilon=0\.3; for empirical semantic robustness we useτ=5\\tau=5; for latent\-space randomized smoothing we useσ=0\.10\\sigma=0\.10; and for concept\-space randomized smoothing we useρ=0\.10\\rho=0\.10\.
Since lower PFR indicates greater robustness, empirical gaps are computed as
GapPFR=PFR\(fstd\)−PFR\(fcbm\)\.\\mathrm\{Gap\}\_\{\\mathrm\{PFR\}\}=\\mathrm\{PFR\}\(f\_\{\\mathrm\{std\}\}\)\-\\mathrm\{PFR\}\(f\_\{\\mathrm\{cbm\}\}\)\.Thus, a positive PFR gap means that the CBM has fewer prediction flips\.
For certified robustness, larger radii indicate greater robustness\. Therefore, latent\-space certified gaps are computed as
GapR=R\(fcbm\)−R\(fstd\),\\mathrm\{Gap\}\_\{R\}=R\(f\_\{\\mathrm\{cbm\}\}\)\-R\(f\_\{\\mathrm\{std\}\}\),and concept\-space certified gaps are computed as
GapΔ⋆=Δ⋆\(fcbm\)−Δ⋆\(fstd\)\.\\mathrm\{Gap\}\_\{\\Delta^\{\\star\}\}=\\Delta^\{\\star\}\(f\_\{\\mathrm\{cbm\}\}\)\-\\Delta^\{\\star\}\(f\_\{\\mathrm\{std\}\}\)\.Thus, positive values consistently indicate that the CBM is more robust\. PFR gaps are reported in percentage points, while certified\-radius gaps are reported in absolute radius units\.
### C\.2More Details on Experiments Setups
##### Hareware\.
We profiled the GPU memory usage of the two main experimental components under the default setting of batch size 8\. All experiments were run on a single NVIDIA RTX 4090 GPU with 24GB memory\. For the latent\-space perturbation experiment, each epsilon setting involves approximately400×1000=400,000400\\times 1000=400\{,\}000latent perturbations, with a peak allocated GPU memory of about 5\.3GB and a PyTorch reserved memory of about 6\.6GB\. For the concept\-space flipping experiment, the default setting involves approximately4×1000×100=1,600,0004\\times 1000\\times 100=1\{,\}600\{,\}000concept perturbations, with a peak allocated GPU memory of about 5\.0GB and a PyTorch reserved memory of about 6\.3GB\. These measurements indicate that, under the reported experimental scale and batch\-size setting, a single RTX 4090 24GB GPU is sufficient for the memory requirements of these experiments\. To ensure stable execution, we recommend keeping at least 8GB of free GPU memory, and preferably 10–12GB\. Each experiment typically uses a single GPU; multiple experiments are run in parallel via separate CUDA\_VISIBLE\_DEVICES configurations
### C\.3Empirical Robustness Results
Table 1:Robustness comparison under geometric \(latentℓ2\\ell\_\{2\}noise\) and semantic \(concept flip\) perturbations\.[Table 1](https://arxiv.org/html/2609.38625#A3.T1)reports the complete empirical robustness results on both CUB\-200\-2011 and RIVAL\-10\. Here, CFR reflects the surrogate ground\-truth rate of concept flips induced by the generator\. We then measure prediction instability via the classifier’s prediction flip rate\. For concept bottleneck models, we additionally evaluate robustness in the intermediate concept space\. The results show that, althoughfcbmf\_\{cbm\}is less robust thanfstdf\_\{std\}at the prediction level, its concept representation is comparatively stable: even when up to 10 concepts are semantically flipped, changes in the concept space remain limited\.
##### Statistical Analysis\.
We test for differences in robustness betweenfstdf\_\{\\text\{std\}\}andfcbmf\_\{\\text\{cbm\}\}using a paired design: for each repeatkk, both classifiers are evaluated on the same subset, yielding paired measurements\. For each perturbation budget, we defineΔk=PFRk\(fstd\)−PFRk\(fcbm\)\\Delta\_\{k\}=\\mathrm\{PFR\}\_\{k\}\(f\_\{\\text\{std\}\}\)\-\\mathrm\{PFR\}\_\{k\}\(f\_\{\\text\{cbm\}\}\)and test the null hypothesisH0:𝔼\[Δ\]=0H\_\{0\}:\\mathbb\{E\}\[\\Delta\]=0with a two\-sided paired test across theK=5K=5repeats\. Since we test multiple budgets, we additionally report Holm–Bonferroni correctedpp\-values\. The per\-budgetpp\-values and corrected values are reported in[Table 2](https://arxiv.org/html/2609.38625#A3.T2)\.
Table 2:Paired significance tests for PFR differences betweenfstdf\_\{\\text\{std\}\}andfcbmf\_\{\\text\{cbm\}\}acrossK=5K=5subset repeats \(each repeat usesn=200n=200images\)\.Δ\\Deltais measured in percentage points \(pp\)\. We report raw two\-sided pairedtt\-testpp\-values and Holm–Bonferroni correctedpHolmp\_\{\\mathrm\{Holm\}\}across all 9 budgets in the table\.Across perturbation budgets, stronger perturbations increase PFR for both classifiers and increase DFS, indicating that the perturbations induce larger changes in the CBM concept representation\. The repeated\-subset statistics show that the reported trends are stable across different random selections of evaluation images, and the paired tests quantify when the observed differences in PFR betweenfstdf\_\{\\text\{std\}\}andfcbmf\_\{\\text\{cbm\}\}are statistically significant\.
##### Additional Analysis of CFR, PFR, and DFS
22442020404060608080DFSPFR \(%\)005510102244CFRDFSFigure 9:mean PFR*vs*mean DFS \(left\) and mean DFS*vs*mean CFR \(right\)\.Red: Geometric,Blue: Semantic\. Solid:fstdf\_\{std\}; Dashed:fcbmf\_\{cbm\}\.As shown in[Figure 9](https://arxiv.org/html/2609.38625#A3.F9), PFR generally increases with CFR, but the relationship is not strictly monotonic\. For example, when the semantic budget isτ=5\\tau=5\(CFR = 5\) and the geometric budget isϵ=0\.5\\epsilon=0\.5\(CFR = 3\.15\), the corresponding PFR values are similar, indicating that higher CFR does not always imply higher prediction instability\.
DFS exhibits a more consistent relationship with PFR\. Geometric perturbations tend to induce larger DFS values than semantic perturbations at comparable CFR levels, suggesting that latent\-space perturbations can distort the model decision representation more strongly than direct concept flips\. This may be due to the nonlinear CB\-AE mapping between latent, concept, and image spaces\.
### C\.4Randomised Smoothing Experiments
Table 3:Unified randomized smoothing results in latent and concept space\. Latent RM uses Gaussian noiseδ∼𝒩\(0,σ2I\)\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)and reports certifiedℓ2\\ell\_\{2\}radiusRR\. Concept RM uses Bernoulli concept flips with rateρ\\rhoand reports certifiedℓ0\\ell\_\{0\}radiusΔ\\Delta\. SmAcc denotes smoothed accuracy\. Values are mean±\\pmstd over five random subsets\. Bold indicates the better value betweenfstdf\_\{\\mathrm\{std\}\}andfcbmf\_\{\\mathrm\{cbm\}\}for CUB\-200\-2011 only; we omit highlighting on RIVAL\-10 due to saturation\.[Table 3](https://arxiv.org/html/2609.38625#A3.T3)reports the complete set of randomized smoothing experimental results\. Under latent\-space smoothing, theResNet baseline generally achieves higher SmAcc and larger certified radii than the PCBMon CUB\-200\-2011, especially asσ\\sigmaincreases\. This indicates that, on generator native inputs, the concept bottleneck does not improve certifiable robustness to continuous latent perturbations\. On the original RIVAL\-10 split, both models are saturated: SmAcc remains at 100% and certified radii are nearly identical across allσ\\sigma, reflecting the strong class separability of this benchmark and indicating that samples remain far from decision boundaries in latent space\.
#### C\.4\.1Statistical Analysis for Latent\-space Randomized Smoothing
We evaluate latent\-space randomized smoothing on a pool of 1000 generator\-native inputs\. To quantify sampling variability from subset selection, we formK=5K=5independent random subsets of sizen=200n=200drawn uniformly without replacement from the 400\-sample pool \(subsets may overlap\)\. We use the same subsets for both methods and all noise scales\. For each input, we estimate smoothed class probabilities usingN=1000N=1000i\.i\.d\. Gaussian perturbations, and compute one\-sided Clopper–Pearson confidence bounds atα=0\.05\\alpha=0\.05\. Following multi\-class confidence separation, the smoothed classifier*abstains*unless the top\-class lower bound exceeds the runner\-up upper bound \(pA\>pBp\_\{A\}\>p\_\{B\}\)\.
##### Metrics\.
Smoothed accuracy \(SmAcc\) counts abstentions as incorrect\. The average certified radiusRRis computed over inputs that are both non\-abstained and correct \(with respect to the reference/no\-noise label\)\.
##### Significance testing\.
For eachσ\\sigmaand each metric, we compare methods using a paired two\-sided test across theK=5K=5subset\-level measurements \(paired because both methods are evaluated on the same subset\)\. We additionally report Holm–Bonferroni corrected p\-values across all hypothesis tests\.
Table 4:Significance analysis overK=5K=5random 200\-sample subsets \(mean±\\pmstd across subsets\)\. Differences are reported asΔ=fstd−fcbm\\Delta=f\_\{std\}\-f\_\{cbm\}\(pp = percentage points\)\. We report paired two\-sided p\-values across the five subset\-level measurements; Holm denotes Holm–Bonferroni correction over all tests\.
#### C\.4\.2Statistical Analysis for Concept\-space Randomized Smoothing
Table 5:Significance analysis for concept\-space \(ℓ0\\ell\_\{0\}\) randomized smoothing overK=5K=5random 200\-sample subsets\. Differences areΔ=fstd−fcbm\\Delta=\\text\{fstd\}\-\\text\{fcbm\}\(pp = percentage points\), reported as mean±\\pmstd across subsets\. We report paired two\-sided p\-values across the five subset\-level measurements; Holm denotes Holm–Bonferroni correction over all66tests \(3 flip budgets×\\times2 metrics\)\.We begin from a pool of 1000 generator\-native inputs\. To quantify variability due to subset selection, we generateK=5K=5independent random subsets of sizen=200n=200by uniform sampling without replacement from the pool \(subsets may overlap\)\. For each subset, we evaluate both methods on the*same*200 inputs \(paired design\), and repeat this for all concept flip budgetsm∈\{1,2,4\}m\\in\\\{1,2,4\\\}\.
##### Smoothed prediction and abstention\.
For each input we estimate class probabilities under fixed\-count concept flips usingN=1000N=1000Monte Carlo samples\. We compute one\-sided Clopper–Pearson confidence bounds atα=0\.05\\alpha=0\.05for the top classAAand runner\-upBB\. The smoothed classifier outputsy^=A\\hat\{y\}=Aonly ifp¯A\>p¯B\\underline\{p\}\_\{A\}\>\\overline\{p\}\_\{B\}; otherwise it abstains\.
##### Metrics\.
Smoothed accuracy \(SmAcc\) counts abstentions as incorrect: a sample is correct iff the smoothed prediction matches the reference label and does not abstain\. The certified flip radiusτ\\tauis computed using theℓ0\\ell\_\{0\}smoothing certificate \(Def\. 5\.7 / Cor\. 5\.6\) and averaged over samples that are both correct and non\-abstained\.
##### Significance testing\.
For each flip budgetmmand each metric \(SmAcc andτ\\tau\), we obtain one measurement per subset for each method, yielding five paired measurements\. We compare methods using a paired two\-sided test across the five subset\-level measurements atα=0\.05\\alpha=0\.05\. Since we test 3 flip budgets and 2 metrics \(6 hypotheses\), we additionally report Holm–Bonferroni corrected p\-values across all 6 tests\.
### C\.5Robustness under Task Conditions
#### C\.5\.1Effect of Class Structure and Semantic Similarity
To construct additional classes for RIVAL\-10, we use a large language model222[https://deploymentsafety\.openai\.com/gpt\-5\-4\-thinking](https://deploymentsafety.openai.com/gpt-5-4-thinking)to identify two sets of ImageNet classes: ten classes that are semantically similar to the original RIVAL\-10 categories and ten classes that are semantically dissimilar\. The selected classes are shown in[Figure 10](https://arxiv.org/html/2609.38625#A3.F10)\. Class similarity is estimated using feature distances extracted from a ResNet\-50 backbone: for each class, we compute a centroid feature via K\-means clustering and measure pairwise similarity using Euclidean distances between these class centroids\.
Figure 10:Distance heatmap among 20 classes under similar and dissimilar settings\.Class distances are computed as Euclidean distances between class\-level centroid features extracted using a ResNet\-50 backbone\.[Table 6](https://arxiv.org/html/2609.38625#A3.T6)and[Table 7](https://arxiv.org/html/2609.38625#A3.T7)present the complete set of empirical and certified robustness experiments\.
Table 6:Robustness on RIVAL\-10 \(18 concepts, varying class count\)\. Backbone: ResNet50\.Geo\.: latentℓ2\\ell\_\{2\}perturbation \(values mean±\\pmSE\)\.Sem\.: concept number flip,τ\\tauconcepts flipped \(global average\)\. CFR \(%\): concept flip rate /τ\\tau; PFR \(%\): prediction flip rate; DFS: mean PCBM concept distance\.Table 7:Randomized smoothing ablation on RIVAL\-10 variants after removing the saturated 10\-class setting\. Latent RM uses Gaussian latent noiseη∼𝒩\(0,σ2I\)\\eta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)and reports certifiedℓ2\\ell\_\{2\}radiusRR\. Concept RM uses Bernoulli concept flips with probabilityρ\\rhoand reports certifiedℓ0\\ell\_\{0\}concept\-flip radiusΔτ\\Delta\_\{\\tau\}\. SmAcc denotes smoothed accuracy\. All variants use 18 concepts\.
#### C\.5\.2Effect of Concept Vocabulary Size
[Table 8](https://arxiv.org/html/2609.38625#A3.T8)and[Table 9](https://arxiv.org/html/2609.38625#A3.T9)presents the complete set of empirical and certified robustness experiments\.
Table 8:Empirical robustness under geometric \(latentℓ2\\ell\_\{2\}\) and semantic \(concept flip\) perturbations on RIVAL\-10 with varying concept vocabulary size\. CFR: concept flip rate; PFR \(%\): prediction flip rate; DFS: PCBM concept distance\.Table 9:RIVAL\-10 10\-class results with 27 vs 36 concepts \(ResNet50\), recomputed with image\-level 5×\\times200 subset aggregation \(mean±\\pmstd\)\. Better value is bolded only when strictly better \(ties are not bolded\)\.
#### C\.5\.3Experiments on DenseNet\-161\.
We test similar experiments on DenseNet\-161\[[Huang et al\., 2017](https://arxiv.org/html/2609.38625#bib.bib16)\]and the results are shown in[Table 10](https://arxiv.org/html/2609.38625#A3.T10)and[Table 11](https://arxiv.org/html/2609.38625#A3.T11)\.
Table 10:DenseNet\-161 direct robustness on RIVAL\-10 variants\. PFR \(%\): prediction flip rate forfdnf\_\{\\text\{dn\}\}\(DenseNet\-161\) andfcbmf\_\{\\text\{cbm\}\}; DFS: average PCBM concept distance\. Mean±\\pmstd over 5 random subsets of 200 images\.Table 11:DenseNet\-161 latent Gaussian randomized smoothing\.z~=z\+η\\tilde\{z\}=z\+\\eta,η∼𝒩\(0,σ2I\)\\eta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\. SmAcc: smoothed accuracy;RR: average certified radius over correctly classified and certified samples\. Mean±\\pmstd over five subsets of 200 images\.
#### C\.5\.4Experiments on Attacks\.
We evaluate adversarial robustness by attacking the entire pipeline using APGD\-CE\[[Croce and Hein, 2020](https://arxiv.org/html/2609.38625#bib.bib7)\]; the corresponding results are reported in[Table 12](https://arxiv.org/html/2609.38625#A3.T12)\.
Table 12:APGD\-CE attack results on the CUB generated dataset based on backbone ResNet\-50\.
## Appendix DCross\-Family Applicability Beyond CBMs
Our primary analysis focuses on Concept Bottleneck Models \(CBMs\), since their explicit concept interface provides a particularly clean setting for separating geometric and semantic robustness\. To examine whether the proposed evaluation procedure is restricted to CBMs, we additionally evaluate a prototype\-based interpretable architecture, PixPNet, together with ResNet and PCBM under the same generator\-defined perturbation protocol\.
Importantly, we do not assume that different model families share a common internal representation\. In particular, the concepts exposed by the generator are not assumed to align one\-to\-one with the localized prototypes learned by PixPNet\. Instead, what is shared across models is the*external perturbation\-to\-prediction interface*: the generator defines a controlled perturbation space, and each downstream classifier is evaluated on the resulting perturbed inputs using its own internal representation\.
This distinction allows us to compare standard, concept\-based, and prototype\-based classifiers under matched perturbations without requiring their heterogeneous internal representations to be aligned\.
### D\.1Models and Evaluation Protocol
We consider three representative model families:
- •Standard classifier:ResNet;
- •Concept\-based interpretable classifier:PCBM, using CAV\-style concept representations;
- •Prototype\-based interpretable classifier:PixPNet, using localized part prototypes\.
For each classifierfmf\_\{m\}, we compose it with the same fixed generatorggand define
Fm\(z\)=fm\(g\(z\)\)\.F\_\{m\}\(z\)=f\_\{m\}\(g\(z\)\)\.Therefore, the generator and perturbation process are shared across model families, while only the downstream classifier changes\.
We evaluate two perturbation geometries:
1. 1\.Gaussian latent perturbations, which induce continuous geometric variation through z′=z\+δ,δ∼𝒩\(0,σ2I\);z^\{\\prime\}=z\+\\delta,\\qquad\\delta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\);
2. 2\.Bernoulli concept interventions, which independently flip binary generator concepts and therefore induce discrete semantic variation\.
For empirical evaluation, we measure prediction flip rate \(PFR\)\. For certified evaluation, we use a perturbation\-specific randomized\-smoothing certificate: anℓ2\\ell\_\{2\}latent\-space certificate for Gaussian perturbations and a Hamming\-space certificate for discrete concept perturbations\. All cross\-family experiments use the same CUB\-200\-2011 generator\-defined perturbation spaces and the same evaluation and certification procedures as the corresponding experiments in the main paper\. The purpose of this section is therefore to vary the downstream classifier family while keeping the perturbation protocol fixed\.
### D\.2Empirical Robustness of a Prototype\-Based Classifier
We first evaluate whether the same generator\-defined perturbations can be used to characterize the prediction stability of a prototype\-based classifier\. Table[13](https://arxiv.org/html/2609.38625#A4.T13)reports PFR for ResNet and PixPNet under Gaussian latent perturbations and Bernoulli semantic interventions\.
Table 13:Prediction flip rate \(PFR\) for ResNet and PixPNet under the same generator\-defined perturbations\. The semantic perturbations are defined in the generator concept space; they are not prototype\-space interventions\.The purpose of this experiment is not to establish a robustness ranking between ResNet and PixPNet\. Rather, it demonstrates that the same generator\-defined perturbation protocol produces measurable and systematic robustness trends for a prototype\-based interpretable classifier\.
In particular, increasing either the latent perturbation strength or the semantic intervention probability increases prediction instability for both architectures\. This suggests that the proposed perturbation\-to\-prediction evaluation interface can be applied to a non\-CBM interpretable architecture without modifying its internal prototype representation\.
### D\.3Cross\-Family Gaussian Randomized Smoothing
We next evaluate whether the randomized\-smoothing component of the framework transfers across classifier architectures\.
For each modelfmf\_\{m\}, we smooth the composite classifier
Fm\(z\)=fm\(g\(z\)\)F\_\{m\}\(z\)=f\_\{m\}\(g\(z\)\)under Gaussian latent perturbations\. Because the perturbation distribution is defined before the downstream classifier, the same smoothing and certification procedure can be applied to ResNet, PCBM, and PixPNet without requiring their internal representations to be aligned\.
We report two quantities:
- •Smoothed accuracy \(SmAcc\):classification accuracy of the smoothed classifier, with abstentions counted as incorrect;
- •Average certified radius \(ACR\):the average certified latent\-space radius over correctly classified, non\-abstaining samples\.
Table 14:Gaussian randomized smoothing across standard, concept\-based, and prototype\-based classifiers\. SmAcc denotes smoothed accuracy \(%\), with abstentions counted as incorrect\. ACR denotes the average certified latent\-space radius over correctly classified, non\-abstaining samples\.All three model families achieve non\-trivial smoothed accuracy and certified latent\-space radii across the tested noise levels\. The purpose of Table[14](https://arxiv.org/html/2609.38625#A4.T14)is therefore not to identify a uniformly most robust architecture\. Instead, it demonstrates that the same randomized\-smoothing procedure and certificate can be applied to standard, concept\-based, and prototype\-based classifiers through the common composite interfaceFm\(z\)=fm\(g\(z\)\)F\_\{m\}\(z\)=f\_\{m\}\(g\(z\)\)\.
The results also reinforce the main conclusion of the paper: robustness differences are regime\-dependent rather than determined solely by whether a model exposes an interpretable intermediate representation\.
### D\.4Discrete Semantic Perturbations Across Model Families
We further examine whether the same evaluation pipeline extends from continuous latent perturbations to discrete semantic interventions\. We use the same CUB\-200\-2011 generator, concept representation, and Bernoulli concept\-space perturbation protocol as in the main experiments\.
Specifically, each binary concept in the*generator concept representation*is independently flipped with probabilityρ\\rho\. The perturbed concept representation is then decoded into an image, which is subsequently evaluated by the downstream classifier\. Thus, the perturbation process is identical across classifier families; only the downstream classifier is changed\.
Importantly, for PixPNet this should not be interpreted as directly perturbing or removing PixPNet prototypes\. The perturbation is defined entirely in the generator concept space\. PixPNet is evaluated as a prototype\-based classifier under the resulting generator\-defined semantic changes, without assuming any one\-to\-one correspondence between generator concepts and PixPNet prototypes\.
Table[15](https://arxiv.org/html/2609.38625#A4.T15)reports the smoothed accuracy under the same Bernoulli smoothing procedure used in the main experiments\.
Table 15:Smoothed accuracy under Bernoulli generator\-concept interventions on CUB\-200\-2011\. The same generator\-defined perturbation protocol is applied to ResNet and PixPNet\.Asρ\\rhoincreases, smoothed accuracy decreases for both classifiers, reflecting the increasing severity of the generator\-defined semantic perturbations\. The purpose of this experiment is not to compare concept representations across architectures, but to test whether the same semantic perturbation and evaluation procedure can be applied to a prototype\-based classifier without modifying its internal prototype representation\.
### D\.5Certified Robustness under Discrete Concept Interventions
For discrete semantic perturbations, we use the same Bernoulli randomized\-smoothing certification procedure described in Sec\.[B\.4](https://arxiv.org/html/2609.38625#A2.SS4)\. Since the perturbation space is discrete, robustness is certified in Hamming space rather than using the Gaussianℓ2\\ell\_\{2\}radius employed for latent perturbations\.
For each downstream classifier, the certificate is defined with respect to the same generator concept space\. Therefore, the certification procedure does not require the downstream model itself to expose concepts or prototypes that are aligned with the generator concepts\.
We report two complementary quantities:
- •ACR\-H: the average certified Hamming radius over correctly classified, non\-abstaining samples;
- •CA@1: the percentage of all evaluated images that are correctly classified and certified against at least one additional generator\-concept flip\.
Table 16:Certified robustness under Bernoulli generator\-concept interventions on CUB\-200\-2011\. ACR\-H denotes the average certified Hamming radius over correctly classified, non\-abstaining samples\. CA@1 denotes the percentage of all evaluated images certified against at least one additional generator\-concept flip\.The certified radii are modest in this setting, and we therefore do not interpret these results as evidence that either architecture is uniformly robust to semantic concept changes\. Instead, the experiment demonstrates that the same generator\-defined semantic perturbation can be paired with a perturbation\-specific certificate for different downstream classifier architectures\.
Together with the Gaussian results, this shows that the framework supports both continuous and discrete perturbation geometries: continuous latent perturbations are paired with a Gaussianℓ2\\ell\_\{2\}certificate, whereas discrete generator\-concept interventions are paired with a Hamming\-space certificate\.
### D\.6Discussion and Scope
The cross\-family experiments support two main conclusions\.
First, the proposed evaluation procedure is not tied to the internal concept representation of a CBM\. The same generator\-defined perturbations can be propagated through standard, concept\-based, and prototype\-based classifiers, enabling matched comparisons of prediction stability under a common perturbation process\.
Second, randomized smoothing can be attached to this common external interface using a certificate matched to the perturbation geometry: Gaussian latent perturbations admit continuous latent\-space certificates, whereas Bernoulli concept interventions admit discrete Hamming\-space certificates\.
Importantly, these results do*not*imply that CBMs, PCBMs, and prototype\-based models share a common interpretable representation\. Nor do we assume that generator concepts correspond directly to PixPNet prototypes\. What is shared across architectures is the external
perturbation→generated input→prediction\\text\{perturbation\}\\rightarrow\\text\{generated input\}\\rightarrow\\text\{prediction\}evaluation interface\.
CBMs remain our primary analytical setting because their explicit semantic bottleneck enables a more detailed analysis of concept\-level sensitivity, concept vocabulary design, and task\-dependent robustness\. The PixPNet experiments provide complementary evidence that the same evaluation and certification pipeline can be applied to classifier families with different internal representations\.Similar Articles
Towards Fine-Grained and Verifiable Concept Bottleneck Models
This paper proposes a fine-grained concept bottleneck model framework that grounds each concept in localized visual evidence, enabling direct verification of concept correctness and improving transparency in medical imaging tasks.
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
This paper evaluates the robustness of AI code agents when codebases are perturbed with semantics-preserving transformations, revealing a jagged frontier where model performance varies unpredictably across different scaffolds and benchmarks.
ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models
ReCBM proposes an uncertainty-gated relational reasoning framework for Concept Bottleneck Models, introducing concept relations like co-occurrence, implication, and exclusion to recover unreliable or missing concept states and improve interpretability and downstream predictions.
How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models
This paper conducts a multi-level analysis of how input perturbations propagate through decoder-only language models, assessing robustness via output behavior, hidden-state geometry, and attention-head function across models like GPT-2 and Qwen2.5.
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
This paper presents a comprehensive empirical evaluation of how large language models handle corruptions in chain-of-thought reasoning steps, testing 13 models across 5 perturbation types (MathError, UnitConversion, Sycophancy, SkippedSteps, ExtraSteps) on mathematical reasoning tasks. The findings reveal heterogeneous vulnerability patterns with implications for deploying LLMs in multi-stage reasoning pipelines.