ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models

arXiv cs.AI Papers

Summary

ReCBM proposes an uncertainty-gated relational reasoning framework for Concept Bottleneck Models, introducing concept relations like co-occurrence, implication, and exclusion to recover unreliable or missing concept states and improve interpretability and downstream predictions.

arXiv:2608.10004v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling semantic inspection and test-time intervention. Recent variants have improved CBMs through richer concept representations, uncertainty estimation, and dependency modeling. However, robust reasoning under unreliable concept states remains underexplored. Without such reasoning, misleading semantic evidence can propagate through the bottleneck, compromising both explanations and downstream predictions. To address this issue, we propose ReCBM, an uncertainty-gated relational reasoning framework for CBMs. ReCBM introduces semantically defined concept relations into the bottleneck and uses uncertainty to guide their refinement. By modeling co-occurrence, implication, and exclusion, ReCBM specifies how evidence is exchanged across concepts, while uncertainty modulates the contribution of each concept during this process. Experiments across diverse datasets showed that ReCBM improved concept and task recovery under missing and flipped concepts, supported uncertainty-aware intervention, and extracted compact task-relevant concept subsets without degrading downstream performance.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:20 AM

# ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models
Source: [https://arxiv.org/html/2608.10004](https://arxiv.org/html/2608.10004)
###### Abstract

Concept Bottleneck Models \(CBMs\) provide an interpretable framework by grounding predictions in human\-understandable concepts, enabling semantic inspection and test\-time intervention\. Recent variants have improved CBMs through richer concept representations, uncertainty estimation, and dependency modeling\. However, robust reasoning under unreliable concept states remains underexplored\. Without such reasoning, misleading semantic evidence can propagate through the bottleneck, compromising both explanations and downstream predictions\. To address this issue, we propose ReCBM, an uncertainty\-gated relational reasoning framework for CBMs\. ReCBM introduces semantically defined concept relations into the bottleneck and uses uncertainty to guide their refinement\. By modeling co\-occurrence, implication, and exclusion, ReCBM specifies how evidence is exchanged across concepts, while uncertainty modulates the contribution of each concept during this process\. Experiments across diverse datasets showed that ReCBM improved concept and task recovery under missing and flipped concepts, supported uncertainty\-aware intervention, and extracted compact task\-relevant concept subsets without degrading downstream performance\.

## 1Introduction

Deep neural networks have substantially advanced a wide range of applications, yet their predictions are rarely grounded in explicit, human\-verifiable evidence\(LeCunet al\.[2015](https://arxiv.org/html/2608.10004#bib.bib1); Rudin[2019](https://arxiv.org/html/2608.10004#bib.bib3); Lipton[2018](https://arxiv.org/html/2608.10004#bib.bib2)\)\. Concept Bottleneck Models \(CBMs\) address this limitation by routing predictions through human\-interpretable concepts, making their decisions more transparent and editable\(Kohet al\.[2020](https://arxiv.org/html/2608.10004#bib.bib9)\)\. However, this interface is useful only when the exposed concept states are reliable\. In practice, concept predictions can be noisy, missing, or inconsistent, while the human feedback used to correct them may be incomplete or uncertain\(Parket al\.[2025](https://arxiv.org/html/2608.10004#bib.bib4); Shinet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib8)\)\. The resulting concept states may therefore remain unreliable, compromising both downstream predictions and the explanations presented to users\. Robust CBMs should thus be able to use the reliable evidence that remains to recover concept states that are missing or identified as unreliable\.

Recovering unreliable concept states requires determining which concepts can be trusted and using them to infer the remaining ones\. Concept\-wise uncertainty provides a reliability signal, while dependencies among concepts allow reliable states to support the recovery of missing or unreliable ones\. For this process to remain interpretable, these dependencies should distinguish semantic relations such as co\-occurrence, implication, and exclusion, which provide different forms of evidence for concept refinement\(Denget al\.[2014](https://arxiv.org/html/2608.10004#bib.bib5); Liet al\.[2015](https://arxiv.org/html/2608.10004#bib.bib6); Xionget al\.[2022](https://arxiv.org/html/2608.10004#bib.bib7)\)\. Crucially, reliability and relational reasoning must operate jointly: uncertainty should regulate which concepts propagate evidence and which concepts receive correction\. In interactive settings, the same mechanism should also accommodate reliability feedback when true concept values are unavailable\. As summarized in Table[1](https://arxiv.org/html/2608.10004#S1.T1), existing CBM extensions provide uncertainty estimation or dependency modeling\(Kimet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib14); Vandenhirtzet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib15); Xuet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib16),[2026](https://arxiv.org/html/2608.10004#bib.bib17)\), but lack a unified mechanism for uncertainty\-guided recovery through semantically defined relations\.

To address this gap, we propose ReCBM, an uncertainty\-gated relational reasoning framework that improves the robustness of CBMs\. ReCBM combines concept relations with concept\-wise uncertainty: reliable concepts provide relational evidence, while unreliable concepts are prevented from propagating errors and can be corrected using evidence from related concepts\. An evidence\-aware anchor preserves confident predictions when relational support is insufficient\. This design enables interpretable concept recovery, while supporting uncertainty\-aware intervention and compact concept subset extraction\.

MethodUn\.De\.Se\.U\-ref\.U\-int\.CBM\(Kohet al\.[2020](https://arxiv.org/html/2608.10004#bib.bib9)\)×\\times×\\times×\\times×\\times×\\timesCEM\(Zarlengaet al\.[2022](https://arxiv.org/html/2608.10004#bib.bib10)\)×\\times×\\times×\\times×\\times×\\timesProbCBM\(Kimet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib14)\)✓\\checkmark×\\times×\\times×\\times×\\timesSCBM\(Vandenhirtzet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib15)\)✓\\checkmark✓\\checkmark×\\times×\\times△\\triangleECBM\(Xuet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib16)\)△\\triangle✓\\checkmark×\\times×\\times×\\timesGraphCBM\(Xuet al\.[2026](https://arxiv.org/html/2608.10004#bib.bib17)\)×\\times✓\\checkmark×\\times×\\times×\\timesReCBM✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmark✓\\checkmarkTable 1:Capabilities for recovering unreliable concepts, including concept uncertainty \(Un\.\), concept dependencies \(De\.\), semantic relations \(Se\.\), uncertainty\-gated refinement \(U\-ref\.\), and uncertainty\-only intervention \(U\-int\.\)\. The symbols✓\\checkmark,△\\triangle, and×\\timesindicate full, partial, and no explicit support, respectively\.We summarize our contributions as follows:

- •We propose ReCBM, an uncertainty\-gated relational reasoning framework that refines concept states using semantically defined co\-occurrence, implication, and exclusion relations\.
- •We introduce an uncertainty intervention mechanism that reduces the influence of uncertain concepts and corrects them using relational evidence, without direct concept intervention\.
- •We develop a sufficient concept set extraction strategy that identifies compact global and class\-specific subsets while preserving the predictive behavior of the full model through relational reconstruction\.
- •Extensive experiments on diverse concept\-based benchmarks demonstrated that ReCBM improved concept and task recovery under unreliable concept states, supported uncertainty\-only intervention, and identified compact concept subsets while preserving task performance\.

## 2Related Work

Robust concept refinement requires interpretable concept states, reliability estimates, and mechanisms for exploiting concept dependencies\.

##### Concept Bottleneck Models\.

Concept Bottleneck Models \(CBMs\) improve interpretability by decomposing prediction into concept prediction and label prediction, so that decisions are mediated by human\-interpretable concepts\(Kohet al\.[2020](https://arxiv.org/html/2608.10004#bib.bib9)\)\. This structure enables concept\-level explanations and test\-time interventions, where users can edit predicted concepts to influence the final output\. Subsequent work has extended CBMs along several directions\. Concept Embedding Models \(CEMs\) improve expressiveness by representing each concept with a learnable embedding rather than a scalar variable\(Zarlengaet al\.[2022](https://arxiv.org/html/2608.10004#bib.bib10)\)\. Post\-hoc and label\-free CBMs reduce the need for fully supervised concept annotations by constructing concept bottlenecks from pretrained models or vision\-language representations\(Yuksekgonulet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib13); Oikarinenet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib12)\)\. Other studies analyze limitations of CBMs, including concept leakage, weak concept alignment, and failures of local semantic grounding\(Mahinpeiet al\.[2021](https://arxiv.org/html/2608.10004#bib.bib19); Margeloiuet al\.[2021](https://arxiv.org/html/2608.10004#bib.bib20); Havasiet al\.[2022](https://arxiv.org/html/2608.10004#bib.bib21); Ramanet al\.[2025](https://arxiv.org/html/2608.10004#bib.bib22)\)\. Our work instead addresses recovery when the exposed concept state is corrupted or unreliable\.

##### Reliability and Uncertainty in Concept Prediction\.

Reliable concept estimates are essential for faithful CBM explanations and effective interventions\. Standard CBMs usually produce deterministic concept predictions, which can be inadequate when concepts are ambiguous, noisy, or difficult to annotate\. Probabilistic CBMs model concept uncertainty through probabilistic concept embeddings, allowing explanations to include both predictions and their uncertainty\(Kimet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib14)\)\. Other uncertainty\-aware variants consider concept ambiguity, annotation uncertainty, or confidence\-based intervention policies\(Chauhanet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib18); Gaoet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib11)\)\. ReCBM is complementary to these approaches: it uses concept\-wise uncertainty as an operational control variable inside the bottleneck, directly regulating which concepts may propagate evidence and which concepts should receive relational correction\.

##### Concept Dependencies and Relational Reasoning\.

Many CBM formulations treat concepts as independent intermediate variables, although semantic concepts often exhibit strong dependencies\. Stochastic CBMs capture correlations through a distributional parameterization and allow interventions on one concept to affect related concepts\(Vandenhirtzet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib15)\)\. Energy\-based CBMs define joint energy functions over inputs, concepts, and labels to model higher\-order interactions\(Xuet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib16)\), while GraphCBM introduces graph\-based concept dependencies\(Xuet al\.[2026](https://arxiv.org/html/2608.10004#bib.bib17)\)\. ReCBM differs by organizing dependency propagation into co\-occurrence, implication, and exclusion channels and coupling these channels with uncertainty\-dependent source, receiver, and anchor gates\. This design clarifies how concepts influence one another and how uncertainty regulates this process\.

## 3Method

### 3\.1Framework Overview

ReCBM refines unreliable concept states through uncertainty\-weighted semantic relations before downstream prediction\. Let𝒟=\{\(xn,yn,𝐜n\)\}n=1N\\mathcal\{D\}=\\\{\(x\_\{n\},y\_\{n\},\\mathbf\{c\}\_\{n\}\)\\\}\_\{n=1\}^\{N\}denote the training set, wherexnx\_\{n\}is the input,yn∈\{1,…,K\}y\_\{n\}\\in\\\{1,\\ldots,K\\\}is the task label, and𝐜n∈\{0,1\}C\\mathbf\{c\}\_\{n\}\\in\\\{0,1\\\}^\{C\}is a vector ofCCpredefined binary concepts\.

![Refer to caption](https://arxiv.org/html/2608.10004v1/x1.png)Figure 1:Overview of ReCBM\. Numbered circles denote concepts, with red and blue representing relatively high and low values\. For conceptjj,Ej\+E\_\{j\}^\{\+\}andEj−E\_\{j\}^\{\-\}denote relational evidence that its probability should increase \(↑\\uparrow\) or decrease \(↓\\downarrow\), respectively\.As illustrated in Fig\.[1](https://arxiv.org/html/2608.10004#S3.F1), the framework consists of four components\. First, an encoderfθf\_\{\\theta\}maps the inputxxto a feature representation𝐳=fθ​\(x\)\\mathbf\{z\}=f\_\{\\theta\}\(x\)\. Second, an evidential concept predictorgψg\_\{\\psi\}produces raw concept probabilities and concept\-wise uncertainties,\(𝐩raw,𝐮\)=gψ​\(𝐳\)\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{u\}\)=g\_\{\\psi\}\(\\mathbf\{z\}\)with a Beta distribution\(Sensoyet al\.[2018](https://arxiv.org/html/2608.10004#bib.bib26)\)\. Third, an uncertainty\-gated relational refinement modulerϕr\_\{\\phi\}propagates evidence through semantic concept relations to obtain𝐩ref=rϕ​\(𝐩raw,𝐮\)\\mathbf\{p\}^\{\\mathrm\{ref\}\}=r\_\{\\phi\}\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{u\}\)\. Finally, the task head produces the predictiony^=argmaxk\[hϑ\(𝐩ref\)\]k\\hat\{y\}=\\arg\\max\_\{k\}\\left\[h\_\{\\vartheta\}\(\\mathbf\{p\}^\{\\mathrm\{ref\}\}\)\\right\]\_\{k\}\.

### 3\.2Learning Concept Relations

Reliable concept refinement requires distinguishing how concepts are related\. ReCBM therefore models co\-occurrence, implication, and exclusion separately using three relation matrices𝐀co\\mathbf\{A\}^\{\\mathrm\{co\}\},𝐀imp\\mathbf\{A\}^\{\\mathrm\{imp\}\}, and𝐀exc∈\[0,1\]C×C\\mathbf\{A\}^\{\\mathrm\{exc\}\}\\in\[0,1\]^\{C\\times C\}, where larger matrix entries indicate stronger relations\. Specifically,Ai​jcoA\_\{ij\}^\{\\mathrm\{co\}\}measures the tendency of conceptsiiandjjto share the same state,Ai​jimpA\_\{ij\}^\{\\mathrm\{imp\}\}measures how strongly the presence of conceptiisupports that of conceptjj, andAi​jexcA\_\{ij\}^\{\\mathrm\{exc\}\}measures the tendency of conceptsiiandjjnot to be active simultaneously\. Co\-occurrence and exclusion are symmetric, whereas implication is directed\. All diagonal entries are set to zero\.

The relation matrices are initialized from pairwise concept statistics in the training set and subsequently optimized with the refinement module\. Symmetry is preserved for co\-occurrence and exclusion during optimization\. Details of the initialization and representative learned relations are provided in the supplementary material\.

### 3\.3Uncertainty\-Gated Relational Refinement

Concept relations allow the state of one concept to provide evidence for another\. However, a relational proposal is useful only when it is supported by sufficient evidence from reliable source concepts, and a target concept should accept this proposal only when correction is warranted\. ReCBM addresses these requirements through an iterative refinement process with receiver, source, and anchor gates\.

##### Relational evidence and iterative refinement\.

For target conceptjj, we aggregate positive and negative relational evidence from concept state𝐩\\mathbf\{p\}using nonnegative source weights𝐰\\mathbf\{w\}:

Ej\+​\(𝐰,𝐩\)=\\displaystyle E\_\{j\}^\{\+\}\(\\mathbf\{w\},\\mathbf\{p\}\)=\{\}∑iwi​pi​\(Ai​jco\+Ai​jimp\),\\displaystyle\\sum\_\{i\}w\_\{i\}p\_\{i\}\\left\(A^\{\\mathrm\{co\}\}\_\{ij\}\+A^\{\\mathrm\{imp\}\}\_\{ij\}\\right\),\(1\)Ej−​\(𝐰,𝐩\)=\\displaystyle E\_\{j\}^\{\-\}\(\\mathbf\{w\},\\mathbf\{p\}\)=\{\}∑iwi\[Ai​jco\(1−pi\)\\displaystyle\\sum\_\{i\}w\_\{i\}\\Bigl\[A^\{\\mathrm\{co\}\}\_\{ij\}\(1\-p\_\{i\}\)\+\\displaystyle\+Aj​iimp\(1−pi\)\+Ai​jexc\[pi\+pj−1\]\+\]\.\\displaystyle A^\{\\mathrm\{imp\}\}\_\{ji\}\(1\-p\_\{i\}\)\+A^\{\\mathrm\{exc\}\}\_\{ij\}\[p\_\{i\}\+p\_\{j\}\-1\]\_\{\+\}\\Bigr\]\.\(2\)Here,Ej\+E\_\{j\}^\{\+\}andEj−E\_\{j\}^\{\-\}support the presence and absence of conceptjj, respectively;wiw\_\{i\}weights the contribution of source conceptii, and\[x\]\+=max⁡\(x,0\)\[x\]\_\{\+\}=\\max\(x,0\)\.

Letmj​\(𝐰,𝐩\)=Ej\+​\(𝐰,𝐩\)\+Ej−​\(𝐰,𝐩\)m\_\{j\}\(\\mathbf\{w\},\\mathbf\{p\}\)=E\_\{j\}^\{\+\}\(\\mathbf\{w\},\\mathbf\{p\}\)\+E\_\{j\}^\{\-\}\(\\mathbf\{w\},\\mathbf\{p\}\)denote the total relational evidence\. We define the relational proposal and its evidence\-mass gate as

qj​\(𝐰,𝐩\)\\displaystyle q\_\{j\}\(\\mathbf\{w\},\\mathbf\{p\}\)=Ej\+​\(𝐰,𝐩\)mj​\(𝐰,𝐩\)\+ϵ,\\displaystyle=\\frac\{E\_\{j\}^\{\+\}\(\\mathbf\{w\},\\mathbf\{p\}\)\}\{m\_\{j\}\(\\mathbf\{w\},\\mathbf\{p\}\)\+\\epsilon\},\(3\)gjmass​\(𝐰,𝐩\)\\displaystyle g\_\{j\}^\{\\mathrm\{mass\}\}\(\\mathbf\{w\},\\mathbf\{p\}\)=mj​\(𝐰,𝐩\)mj​\(𝐰,𝐩\)\+1,\\displaystyle=\\frac\{m\_\{j\}\(\\mathbf\{w\},\\mathbf\{p\}\)\}\{m\_\{j\}\(\\mathbf\{w\},\\mathbf\{p\}\)\+1\},\(4\)whereϵ\>0\\epsilon\>0is a small constant for numerical stability\. The proposalqjq\_\{j\}specifies the state suggested by the related concepts, whereas the total evidencemjm\_\{j\}controls how strongly this proposal influences the refinement through the evidence\-mass gategjmassg\_\{j\}^\{\\mathrm\{mass\}\}\. Specifically,gjmassg\_\{j\}^\{\\mathrm\{mass\}\}remains close to zero when the available relational evidence is weak and increases gradually toward one as the evidence mass grows\.

Starting from𝐩\(0\)=𝐩raw\\mathbf\{p\}^\{\(0\)\}=\\mathbf\{p\}^\{\\mathrm\{raw\}\}, ReCBM performsTTrefinement iterations\. At iterationtt, we use thesource gate𝐬\(t\)\\mathbf\{s\}^\{\(t\)\}defined later in Eq\. \([12](https://arxiv.org/html/2608.10004#S3.E12)\) as the source reliability weights, to compute

qj\(t\)\\displaystyle q\_\{j\}^\{\(t\)\}=qj​\(𝐬\(t\),𝐩\(t\)\),\\displaystyle=q\_\{j\}\(\\mathbf\{s\}^\{\(t\)\},\\mathbf\{p\}^\{\(t\)\}\),\(5\)gjmass,\(t\)\\displaystyle g\_\{j\}^\{\\mathrm\{mass\},\(t\)\}=gjmass​\(𝐬\(t\),𝐩\(t\)\)\.\\displaystyle=g\_\{j\}^\{\\mathrm\{mass\}\}\(\\mathbf\{s\}^\{\(t\)\},\\mathbf\{p\}^\{\(t\)\}\)\.\(6\)The target probability is then updated according to

pj\(t\+1\)=clip​\(pj\(t\)\+δ​rj\(t\)​gjmass,\(t\)​\(qj\(t\)−pj\(t\)\),ϵ,1−ϵ\),p\_\{j\}^\{\(t\+1\)\}=\\mathrm\{clip\}\\\!\\left\(p\_\{j\}^\{\(t\)\}\+\\delta r\_\{j\}^\{\(t\)\}g\_\{j\}^\{\\mathrm\{mass\},\(t\)\}\\bigl\(q\_\{j\}^\{\(t\)\}\-p\_\{j\}^\{\(t\)\}\\bigr\),\\epsilon,1\-\\epsilon\\right\),\(7\)whereδ\>0\\delta\>0is a learned step size andrj\(t\)r\_\{j\}^\{\(t\)\}is thereceiver gatedefined below\.

##### Receiver gate\.

The receiver gate determines how strongly target conceptjjaccepts the relational update:

rj\(t\)=σ​\(γu​uj\+γv​gjmass,\(t\)​\|qj\(t\)−pj\(t\)\|\+br\),r\_\{j\}^\{\(t\)\}=\\sigma\\\!\\left\(\\gamma\_\{u\}u\_\{j\}\+\\gamma\_\{v\}g\_\{j\}^\{\\mathrm\{mass\},\(t\)\}\\left\|q\_\{j\}^\{\(t\)\}\-p\_\{j\}^\{\(t\)\}\\right\|\+b\_\{r\}\\right\),\(8\)whereσ​\(⋅\)\\sigma\(\\cdot\)is the sigmoid function,γu,γv\>0\\gamma\_\{u\},\\gamma\_\{v\}\>0are learned coefficients, andbrb\_\{r\}is a learned bias\. A larger update is allowed when the current prediction is uncertain or when it differs substantially from a proposal supported by strong relational evidence\.

##### Source gate\.

While the receiver gate controls the target’s acceptance of an update, the source gate controls how strongly each source concept contributes to the relational proposal\. A reliable source should have low uncertainty and agree with the estimate supported by its reliable neighbors\.

To evaluate this consistency, we first compute an uncertainty\-weighted preliminary proposal and its evidence mass:

q¯i\(t\)\\displaystyle\\bar\{q\}\_\{i\}^\{\(t\)\}=qi​\(𝟏−𝐮,𝐩\(t\)\),\\displaystyle=q\_\{i\}\(\\mathbf\{1\}\-\\mathbf\{u\},\\mathbf\{p\}^\{\(t\)\}\),\(9\)g¯i\(t\)\\displaystyle\\bar\{g\}\_\{i\}^\{\(t\)\}=gimass​\(𝟏−𝐮,𝐩\(t\)\)\.\\displaystyle=g\_\{i\}^\{\\mathrm\{mass\}\}\(\\mathbf\{1\}\-\\mathbf\{u\},\\mathbf\{p\}^\{\(t\)\}\)\.\(10\)Thus, more certain concepts contribute more strongly to the preliminary proposal\. We then measure the disagreement between conceptiiand this proposal:

vi\(t\)=g¯i\(t\)​\|q¯i\(t\)−pi\(t\)\|\.v\_\{i\}^\{\(t\)\}=\\bar\{g\}\_\{i\}^\{\(t\)\}\\left\|\\bar\{q\}\_\{i\}^\{\(t\)\}\-p\_\{i\}^\{\(t\)\}\\right\|\.\(11\)The disagreementvi\(t\)v\_\{i\}^\{\(t\)\}becomes large only when conceptiidiffers from a well\-supported preliminary proposal\. Based on this disagreement, the source gate is defined as

si\(t\)=\(1−ui\)​σ​\(bs−γs​vi\(t\)\),s\_\{i\}^\{\(t\)\}=\(1\-u\_\{i\}\)\\sigma\\\!\\left\(b\_\{s\}\-\\gamma\_\{s\}v\_\{i\}^\{\(t\)\}\\right\),\(12\)whereγs\>0\\gamma\_\{s\}\>0andbsb\_\{s\}are learned parameters\. The resulting𝐬\(t\)\\mathbf\{s\}^\{\(t\)\}is used as the source weights in the relational proposal\.

##### Anchor gate\.

AfterTTiterations, the anchor gate preserves reliable raw predictions that have weak initial relational support:

aj=\(1−gjmass,\(0\)\)​\(1−uj\),a\_\{j\}=\\left\(1\-g\_\{j\}^\{\\mathrm\{mass\},\(0\)\}\\right\)\(1\-u\_\{j\}\),\(13\)where a largeraja\_\{j\}assigns greater importance to the raw prediction\.

Finally, the refined concept probability is obtained by:

pjref=aj​pjraw\+\(1−aj\)​pj\(T\)\.p\_\{j\}^\{\\mathrm\{ref\}\}=a\_\{j\}p\_\{j\}^\{\\mathrm\{raw\}\}\+\(1\-a\_\{j\}\)p\_\{j\}^\{\(T\)\}\.\(14\)

### 3\.4Training Objective

Relational refinement requires stable concept probabilities and informative uncertainty estimates\. We therefore train ReCBM in two stages\.

##### Stage 1: Evidential concept learning\.

In Stage 1, we jointly train the encoderfθf\_\{\\theta\}, evidential concept predictorgψg\_\{\\psi\}, and task predictorhϑh\_\{\\vartheta\}\. LetΩ=\{\(n,i\):cn,i​is observed\}\\Omega=\\\{\(n,i\):c\_\{n,i\}\\text\{ is observed\}\\\}index the available annotations, wherennandiidenote the sample and concept, respectively\. For concept supervision, we use the evidential learning objective proposed in prior work\(Gaoet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib11); Sensoyet al\.[2018](https://arxiv.org/html/2608.10004#bib.bib26)\), which consists of an evidential negative log\-likelihoodℒconcept\\mathcal\{L\}\_\{\\mathrm\{concept\}\}and a KL regularizerℒKL\\mathcal\{L\}\_\{\\mathrm\{KL\}\}\. We weight the latter by an annealing coefficientωKL∈\[0,1\]\\omega\_\{\\mathrm\{KL\}\}\\in\[0,1\]that increases linearly during the early stage of training\. The task loss is

ℒtask=CE​\(hϑ​\(𝐩raw\),y\)\.\\mathcal\{L\}\_\{\\mathrm\{task\}\}=\\mathrm\{CE\}\\\!\\left\(h\_\{\\vartheta\}\(\\mathbf\{p\}^\{\\mathrm\{raw\}\}\),y\\right\)\.\(15\)Here,CE\\mathrm\{CE\}denotes cross\-entropy\. Following the calibrated uncertainty principle used in prior evidential models\(Zouet al\.[2025](https://arxiv.org/html/2608.10004#bib.bib27)\), we encourage incorrect concept predictions to have high uncertainty and correct predictions to have low uncertainty\. For each observed concept, we define the error indicatoren,i=𝕀​\[𝕀​\(pn,iraw≥0\.5\)≠cn,i\]e\_\{n,i\}=\\mathbb\{I\}\\\!\\left\[\\mathbb\{I\}\(p^\{\\mathrm\{raw\}\}\_\{n,i\}\\geq 0\.5\)\\neq c\_\{n,i\}\\right\]\.

LetΩℬ\\Omega\_\{\\mathcal\{B\}\}denote the set of observed sample–concept pairs in the current minibatch, and define the number of pairs with error labelρ∈\{0,1\}\\rho\\in\\\{0,1\\\}as𝒩​\(ρ\)=∑\(n,i\)∈Ωℬ𝕀​\[en,i=ρ\]\.\\mathcal\{N\}\(\\rho\)=\\sum\_\{\(n,i\)\\in\\Omega\_\{\\mathcal\{B\}\}\}\\mathbb\{I\}\[e\_\{n,i\}=\\rho\]\.To prevent the more frequent error group from dominating training, we use the group\-balanced uncertainty loss

ℒu=∑\(n,i\)∈Ωℬ𝒩​\(en,i\)−1​ℓBCE​\(un,i,en,i\)∑\(n,i\)∈Ωℬ𝒩​\(en,i\)−1,\\mathcal\{L\}\_\{u\}=\\frac\{\\sum\_\{\(n,i\)\\in\\Omega\_\{\\mathcal\{B\}\}\}\\mathcal\{N\}\(e\_\{n,i\}\)^\{\-1\}\\ell\_\{\\mathrm\{BCE\}\}\(u\_\{n,i\},e\_\{n,i\}\)\}\{\\sum\_\{\(n,i\)\\in\\Omega\_\{\\mathcal\{B\}\}\}\\mathcal\{N\}\(e\_\{n,i\}\)^\{\-1\}\},\(16\)whereℓBCE​\(u,e\)=−e​log⁡u−\(1−e\)​log⁡\(1−u\)\.\\ell\_\{\\mathrm\{BCE\}\}\(u,e\)=\-e\\log u\-\(1\-e\)\\log\(1\-u\)\.

Finally, the Stage 1 objective is

ℒstage1=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{stage1\}\}=\{\}λc​\(ℒconcept\+ωKL​ℒKL\)\\displaystyle\\lambda\_\{c\}\\left\(\\mathcal\{L\}\_\{\\mathrm\{concept\}\}\+\\omega\_\{\\mathrm\{KL\}\}\\mathcal\{L\}\_\{\\mathrm\{KL\}\}\\right\)\+λy​ℒtask\+λu​ℒu,\\displaystyle\+\\lambda\_\{y\}\\mathcal\{L\}\_\{\\mathrm\{task\}\}\+\\lambda\_\{u\}\\mathcal\{L\}\_\{u\},\(17\)whereλc,λy,λu≥0\\lambda\_\{c\},\\lambda\_\{y\},\\lambda\_\{u\}\\geq 0control the contributions of concept supervision, task prediction, and uncertainty calibration, respectively\.

##### Stage 2: Relational refinement\.

During Stage 2, we freezefθf\_\{\\theta\}andgψg\_\{\\psi\}, and optimize the refinement modulerϕr\_\{\\phi\}together with the task predictorhϑh\_\{\\vartheta\}, initialized from Stage 1\.

To exposerϕr\_\{\\phi\}to unreliable concept states, we sampleΩaug⊆Ω\\Omega\_\{\\mathrm\{aug\}\}\\subseteq\\Omegaand replace each selected probability with the value opposite to its label:

pn,iraw,aug=\{0,cn,i=1,1,cn,i=0,​un,iaug=1,\(n,i\)∈Ωaug\.p\_\{n,i\}^\{\\mathrm\{raw,aug\}\}=\\begin\{cases\}0,&c\_\{n,i\}=1,\\\\ 1,&c\_\{n,i\}=0,\\end\{cases\}u\_\{n,i\}^\{\\mathrm\{aug\}\}=1,\(n,i\)\\in\\Omega\_\{\\mathrm\{aug\}\}\.\(18\)All unselected entries retain their original probabilities and uncertainties\. Let𝐩ref,aug\\mathbf\{p\}^\{\\mathrm\{ref,aug\}\}denote the output obtained by refining the augmented state\. We define the refinement loss as

ℒref=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{ref\}\}=\{\}ℬ​\(𝐩ref,𝐜;Ω\)\+κaug​ℬ​\(𝐩ref,aug,𝐜;Ωaug\),\\displaystyle\\mathcal\{B\}\(\\mathbf\{p\}^\{\\mathrm\{ref\}\},\\mathbf\{c\};\\Omega\)\+\\kappa\_\{\\mathrm\{aug\}\}\\mathcal\{B\}\(\\mathbf\{p\}^\{\\mathrm\{ref,aug\}\},\\mathbf\{c\};\\Omega\_\{\\mathrm\{aug\}\}\),\(19\)where𝐜\\mathbf\{c\}collects the ground\-truth concepts,ℬ​\(𝐩,𝐜;S\)\\mathcal\{B\}\(\\mathbf\{p\},\\mathbf\{c\};S\)denotes binary cross\-entropy over an index setSS, averaged separately over its positive and negative labels, andκaug≥0\\kappa\_\{\\mathrm\{aug\}\}\\geq 0controls the strength of augmentation supervision\.

We further discourage refinement from moving an observed concept away from its label\. Let𝒟𝐠​\(𝝃,𝜼\)\\mathcal\{D\}\_\{\\mathbf\{g\}\}\(\\boldsymbol\{\\xi\},\\boldsymbol\{\\eta\}\)denote a gated directional penalty that measures whether an update from𝝃\\boldsymbol\{\\xi\}to𝜼\\boldsymbol\{\\eta\}moves observed concept probabilities away from their labels\. We apply directional supervision at three points: to the final refined probabilities𝐩ref\\mathbf\{p\}^\{\\mathrm\{ref\}\}, to the state𝐩\(T\)\\mathbf\{p\}^\{\(T\)\}produced by theTTiterative relational updates, and to the relation\-induced proposal𝐪\(t\)\\mathbf\{q\}^\{\(t\)\}at each iteration\. The resulting loss is

ℒdir=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{dir\}\}=\{\}12​𝒟𝐠0​\(𝐩raw,𝐩ref\)\+12​𝒟𝐠0​\(𝐩raw,𝐩\(T\)\)\\displaystyle\\tfrac\{1\}\{2\}\\mathcal\{D\}\_\{\\mathbf\{g\}^\{0\}\}\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{p\}^\{\\mathrm\{ref\}\}\)\+\\tfrac\{1\}\{2\}\\mathcal\{D\}\_\{\\mathbf\{g\}^\{0\}\}\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{p\}^\{\(T\)\}\)\+κmsgT​∑t=0T−1𝒟𝐠mass,\(t\)​\(𝐩\(t\),𝐪\(t\)\),\\displaystyle\+\\frac\{\\kappa\_\{\\mathrm\{msg\}\}\}\{T\}\\sum\_\{t=0\}^\{T\-1\}\\mathcal\{D\}\_\{\\mathbf\{g\}^\{\\mathrm\{mass\},\(t\)\}\}\(\\mathbf\{p\}^\{\(t\)\},\\mathbf\{q\}^\{\(t\)\}\),\(20\)where the first two terms supervise all observed entries and use the uniform gate𝐠0\\mathbf\{g\}^\{0\}, defined bygn,i0=1g^\{0\}\_\{n,i\}=1for every\(n,i\)∈Ω\(n,i\)\\in\\Omega, andκmsg≥0\\kappa\_\{\\mathrm\{msg\}\}\\geq 0controls intermediate\-message supervision\.

To define the gated directional penalty, letΩz=\{\(n,i\)∈Ω:cn,i=z\}\\Omega\_\{z\}=\\\{\(n,i\)\\in\\Omega:c\_\{n,i\}=z\\\}forz∈\{0,1\}z\\in\\\{0,1\\\}, and let𝒵⊆\{0,1\}\\mathcal\{Z\}\\subseteq\\\{0,1\\\}denote the label groups present in the current minibatch\. We define

𝒟𝐠​\(𝝃,𝜼\)\\displaystyle\\mathcal\{D\}\_\{\\mathbf\{g\}\}\(\\boldsymbol\{\\xi\},\\boldsymbol\{\\eta\}\)=1\|𝒵\|​∑z∈𝒵∑\(n,i\)∈Ωzgn,i​\[\(2​z−1\)​\(ξn,i−ηn,i\)\]\+∑\(n,i\)∈Ωzgn,i\+ϵ\.\\displaystyle=\\frac\{1\}\{\|\\mathcal\{Z\}\|\}\\sum\_\{z\\in\\mathcal\{Z\}\}\\frac\{\\sum\_\{\(n,i\)\\in\\Omega\_\{z\}\}g\_\{n,i\}\\big\[\(2z\-1\)\(\\xi\_\{n,i\}\-\\eta\_\{n,i\}\)\\big\]\_\{\+\}\}\{\\sum\_\{\(n,i\)\\in\\Omega\_\{z\}\}g\_\{n,i\}\+\\epsilon\}\.\(21\)Here,𝝃\\boldsymbol\{\\xi\}and𝜼\\boldsymbol\{\\eta\}denote the concept states before and after an update, respectively\. The penalty is nonzero when an update decreases a positive\-concept probability or increases a negative\-concept probability\. The gategn,ig\_\{n,i\}weights the directional penalty for each entry, while averaging over the nonempty label groups prevents either group from dominating the loss\.

The primary Stage 2 objective is

ℒstage2=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{stage2\}\}=\{\}λy​CE​\(hϑ​\(𝐩ref\),y\)\+λref​ℒref\+λdir​ℒdir,\\displaystyle\\lambda\_\{y\}\\mathrm\{CE\}\\\!\\left\(h\_\{\\vartheta\}\(\\mathbf\{p\}^\{\\mathrm\{ref\}\}\),y\\right\)\+\\lambda\_\{\\mathrm\{ref\}\}\\mathcal\{L\}\_\{\\mathrm\{ref\}\}\+\\lambda\_\{\\mathrm\{dir\}\}\\mathcal\{L\}\_\{\\mathrm\{dir\}\},\(22\)whereλref\\lambda\_\{\\mathrm\{ref\}\}andλdir\\lambda\_\{\\mathrm\{dir\}\}are nonnegative loss weights\. Auxiliary optimization terms are detailed in the supplementary material\.

### 3\.5Sufficient Concept Set Extraction

Concept relations may allow part of the bottleneck to be reconstructed from the remaining concepts\. We therefore seek compact concept subsets that preserve both concept reconstruction and downstream prediction\. To identify such subsets, we define an anchor\-weighted task\-gradient importance score for each concept and use this score to guide a greedy search\.

Specifically, the importance of conceptiiis defined as

scorei=\(1Nval​∑n=1Nval\|∂ℓn,yn∂pn,iref\|\)​\(1Nval​∑n=1Nvalan,i\),\\mathrm\{score\}\_\{i\}=\\left\(\\frac\{1\}\{N\_\{\\mathrm\{val\}\}\}\\sum\_\{n=1\}^\{N\_\{\\mathrm\{val\}\}\}\\left\|\\frac\{\\partial\\ell\_\{n,y\_\{n\}\}\}\{\\partial p^\{\\mathrm\{ref\}\}\_\{n,i\}\}\\right\|\\right\)\\left\(\\frac\{1\}\{N\_\{\\mathrm\{val\}\}\}\\sum\_\{n=1\}^\{N\_\{\\mathrm\{val\}\}\}a\_\{n,i\}\\right\),\(23\)whereNvalN\_\{\\mathrm\{val\}\}is the number of validation samples,ℓn,yn\\ell\_\{n,y\_\{n\}\}is the logit of the ground\-truth class, andan,ia\_\{n,i\}is the preservation anchor defined in Eq\. \([13](https://arxiv.org/html/2608.10004#S3.E13)\)\. The first factor measures the sensitivity of the task prediction to conceptii\. The second factor measures how strongly the refined prediction retains the corresponding raw concept value\. Their product therefore prioritizes concepts that are important to the task and cannot be readily reconstructed from the remaining concepts\.

Given a candidate setSS, we construct the refinement input as

\(p~i,u~i\)=\{\(piraw,0\),i∈S,\(0\.5,1\),i∉S\.\(\\tilde\{p\}\_\{i\},\\tilde\{u\}\_\{i\}\)=\\begin\{cases\}\(p\_\{i\}^\{\\mathrm\{raw\}\},0\),&i\\in S,\\\\ \(0\.5,1\),&i\\notin S\.\\end\{cases\}\(24\)Thus, selected concepts retain their raw predictions, while unselected concepts are treated as unknown\. The relation module reconstructs𝐩~ref\\tilde\{\\mathbf\{p\}\}^\{\\mathrm\{ref\}\}from this partial concept state\. We define concept and task retention as the validation concept F1 and task accuracy obtained from𝐩~ref\\tilde\{\\mathbf\{p\}\}^\{\\mathrm\{ref\}\}, respectively, divided by their full\-model counterparts obtained from𝐩ref\\mathbf\{p\}^\{\\mathrm\{ref\}\}\. A setSSis considered sufficient if both retention values are at least100%100\\%\. We rank concepts byscorei\\mathrm\{score\}\_\{i\}and greedily search the resulting order for a compact set satisfying both conditions\. Full search and pruning details are provided in the supplementary material\.

Repeating the search for each classyyyields a class\-specific setSyS\_\{y\}\. At test time, the concepts inSyS\_\{y\}are supplied with their ground\-truth values, while the remaining concepts are treated as unknown\. We evaluate eachSyS\_\{y\}as a one\-vs\-rest classifier, treatingyyas positive and all other classes as negative, and report balanced accuracy averaged over the classes for which a sufficient set is found\.

## 4Experiments

Our experiments assessed whether relational refinement recovered unreliable concepts, which components drove this recovery, and whether uncertainty supported intervention and compact concept\-set extraction\.

### 4\.1Experimental Setup

#### Datasets

We evaluated ReCBM on two image datasets with human\-annotated concepts and a controlled synthetic dataset with predefined concept relations\. Table[2](https://arxiv.org/html/2608.10004#S4.T2)summarizes their statistics\.

DatasetTrainValidationTestConcepts / ClassesWBC6,1691,0303,09924 / 5CUB4,7961,1985,794112 / 200Synthetic12,0003,0003,00012 / 4Table 2:Dataset statistics\.WBC\.The White Blood Cell Attribute dataset \(WBC\) contains peripheral blood cell images from five leukocyte classes\(Tsutsuiet al\.[2026](https://arxiv.org/html/2608.10004#bib.bib23)\)\. We converted its 11 categorical morphological attributes into 24 binary concepts\.

CUB\.The Caltech\-UCSD Birds\-200\-2011 dataset contains 11,788 images from 200 bird species\(Wahet al\.[2011](https://arxiv.org/html/2608.10004#bib.bib24)\)\. Following standard CBM preprocessing\(Kohet al\.[2020](https://arxiv.org/html/2608.10004#bib.bib9)\), we retained the 112 binary attributes as concepts\.

Synthetic\.We constructed a controlled dataset with 12 binary concepts, 4 classes, and predefined co\-occurrence, implication, and exclusion relations\. Full generation details are provided in the supplementary material\.

#### Baselines

We compared ReCBM with five baseline configurations: independently and jointly trained CBMs\(Kohet al\.[2020](https://arxiv.org/html/2608.10004#bib.bib9)\), ProbCBM\(Kimet al\.[2023](https://arxiv.org/html/2608.10004#bib.bib14)\), SCBM\(Vandenhirtzet al\.[2024](https://arxiv.org/html/2608.10004#bib.bib15)\), and GraphCBM\(Xuet al\.[2026](https://arxiv.org/html/2608.10004#bib.bib17)\)\.

#### Implementation Details

ForWBCandCUB, we used a ResNet\-34 concept encoder with images resized to224×224224\\times 224\. ForSynthetic, we used an MLP with hidden dimension 128\. ReCBM was trained for up to 150 epochs, with 70 epochs for Stage 1 and 80 epochs for Stage 2\. The initial learning rate was5×10−45\\times 10^\{\-4\}forWBCandCUBand10−310^\{\-3\}forSynthetic\.

The relation matrices were initialized from training\-set concept statistics and optimized during Stage 2\. We usedT=5T=5refinement iterations forWBCandSyntheticandT=10T=10forCUB\. For Stage 2 augmentation, each observed concept entry was included inΩaug\\Omega\_\{\\mathrm\{aug\}\}with probability0\.10\.1\. In Stage 1, we setλc=λy=1\\lambda\_\{c\}=\\lambda\_\{y\}=1andλu=10\\lambda\_\{u\}=10, and linearly annealedωKL\\omega\_\{\\mathrm\{KL\}\}from 0 to 1 over the first 10 epochs\. In Stage 2, we setλy=1\\lambda\_\{y\}=1,λdir=10\\lambda\_\{\\mathrm\{dir\}\}=10, andκmsg=1\\kappa\_\{\\mathrm\{msg\}\}=1\. We used\(λref,κaug\)=\(20,0\.25\)\(\\lambda\_\{\\mathrm\{ref\}\},\\kappa\_\{\\mathrm\{aug\}\}\)=\(20,0\.25\)forWBCand\(10,1\)\(10,1\)forCUBandSynthetic\. All experiments were conducted on a single NVIDIA GeForce RTX 3090 GPU\. Additional training hyperparameters are provided in the supplementary material\.

![Refer to caption](https://arxiv.org/html/2608.10004v1/x2.png)Figure 2:Recovery from missing concepts\. Concept and task accuracy are reported as the proportion of entries assigned the neutral state increases\.![Refer to caption](https://arxiv.org/html/2608.10004v1/x3.png)Figure 3:Recovery from concept flips\. Concept and task accuracy are reported as the proportion of flipped entries assigned maximal uncertainty increases\.

### 4\.2Robustness to Concept Corruption

We evaluated the robustness of ReCBM at evenly spaced corruption ratios from 0 to 1, defined asrk=k/10r\_\{k\}=k/10for indiceskkranging from 0 to 10\. At test time, selected concept entries were replaced by\(pn,iraw,un,i\)=\(0\.5,1\)\(p^\{\\mathrm\{raw\}\}\_\{n,i\},u\_\{n,i\}\)=\(0\.5,1\)formissingcorruption, or by\(pn,iraw,un,i\)=\(1−c^n,i,1\)\(p^\{\\mathrm\{raw\}\}\_\{n,i\},u\_\{n,i\}\)=\(1\-\\hat\{c\}\_\{n,i\},1\)forflipcorruption, wherec^n,i=𝕀​\[pn,iraw≥0\.5\]\\hat\{c\}\_\{n,i\}=\\mathbb\{I\}\[p^\{\\mathrm\{raw\}\}\_\{n,i\}\\geq 0\.5\]\. Robustness across all ratios was summarized by the trapezoidal Corruption Robustness AUC:

CR​\-​AUC=∑k=0K−1Acc​\(rk\)\+Acc​\(rk\+1\)2​\(rk\+1−rk\)\.\\mathrm\{CR\\text\{\-\}AUC\}=\\sum\_\{k=0\}^\{K\-1\}\\frac\{\\mathrm\{Acc\}\(r\_\{k\}\)\+\\mathrm\{Acc\}\(r\_\{k\+1\}\)\}\{2\}\(r\_\{k\+1\}\-r\_\{k\}\)\.\(25\)Higher CR\-AUC indicates better robustness across corruption severities\.

MethodMissingFlipMacroConceptTaskConceptTaskCR\-AUCGraphCBM63\.2060\.6250\.1318\.6148\.14SCBM61\.0838\.4956\.3836\.0648\.00ProbCBM64\.7068\.8050\.1130\.1053\.43CBM\-Joint58\.6270\.5150\.1328\.5951\.96CBM\-Independent64\.4966\.5350\.1128\.6652\.45ReCBM86\.5973\.7276\.0157\.6473\.49Table 3:CR\-AUC \(%\) averaged across datasets\. Macro averages missing/flip and concept/task CR\-AUC\.Table[3](https://arxiv.org/html/2608.10004#S4.T3)shows that ReCBM achieved the highest concept and task CR\-AUC under both corruption types, demonstrating recovery from both missing and misleading concept evidence\.

Figure[2](https://arxiv.org/html/2608.10004#S4.F2)shows that ReCBM degraded more slowly as concepts were replaced by neutral states\. Relational refinement reconstructed missing states from the remaining reliable concepts, delaying the loss of both concept and task information\. This advantage diminished near full corruption, where little reliable evidence remained for reconstruction\.

Figure[3](https://arxiv.org/html/2608.10004#S4.F3)presents a harder setting in which corrupted concepts provided active but incorrect evidence\. ReCBM limited their influence through the source gate and admitted corrections supported by reliable neighbors, yielding substantially slower degradation across most corruption ratios\. SCBM’s rebound at full corruption reflected its fallback to the training prior when no reliable concepts remained, rather than concept recovery\.

The starting points of Figures[2](https://arxiv.org/html/2608.10004#S4.F2)and[3](https://arxiv.org/html/2608.10004#S4.F3)correspond to performance without corruption \(r=0r=0\)\. ReCBM remained competitive with existing methods in this setting, indicating that its robustness gains did not come at the cost of substantially degraded performance without corruption\.

### 4\.3Ablation Study

We ablated each relation type individually and jointly, as well as adaptive gating\. For the latter, the source and receiver gates were fixed to one and the preservation anchor to zero, yielding fully ungated relational updates\. The ablation was conducted onWBCunder50%50\\%missing and flip corruption\. Results on the remaining datasets are provided in the supplementary material\.

CIEGFlipMissingConceptTaskConceptTask✓✓✓✓80\.10±\\pm3\.1476\.00±\\pm2\.6290\.89±\\pm0\.2195\.31±\\pm0\.49×\\times✓✓✓74\.53±\\pm1\.7059\.75±\\pm1\.4290\.61±\\pm0\.0794\.87±\\pm0\.26✓×\\times✓✓72\.89±\\pm0\.2171\.81±\\pm0\.5088\.70±\\pm0\.1594\.84±\\pm0\.69✓✓×\\times✓78\.09±\\pm5\.0365\.87±\\pm7\.3489\.23±\\pm0\.3694\.85±\\pm0\.53✓×\\times×\\times✓64\.53±\\pm1\.4655\.02±\\pm1\.1286\.42±\\pm0\.1493\.97±\\pm0\.40×\\times✓×\\times✓59\.24±\\pm0\.9730\.54±\\pm1\.8185\.69±\\pm0\.0593\.99±\\pm0\.47×\\times×\\times✓✓66\.02±\\pm0\.2550\.31±\\pm1\.0581\.20±\\pm0\.0794\.30±\\pm0\.41×\\times×\\times×\\times✓49\.94±\\pm0\.1113\.92±\\pm0\.6364\.44±\\pm0\.1491\.84±\\pm0\.86✓✓✓×\\times74\.83±\\pm0\.0614\.71±\\pm0\.9985\.77±\\pm0\.1290\.59±\\pm1\.85Table 4:Component ablation onWBCunder 50% concept corruption\. C, I, E, and G denote co\-occurrence, implication, exclusion, and adaptive gating\. Concept and task accuracy \(%\) are reported as mean±\\pmstandard deviation over three paired seeds\. Best means are in bold\.As presented in Table[4](https://arxiv.org/html/2608.10004#S4.T4), the complete model obtained the highest mean concept and task accuracy under both corruption settings\. The sharp drop in task accuracy without adaptive gating, particularly under flips, showed that controlling unreliable propagation was critical when corrupted concepts provided active evidence\.

The relation ablations further showed that co\-occurrence, implication, and exclusion provided complementary recovery signals\. No individual or pairwise configuration matched the complete model across both corruption types, while removing all relation channels caused the largest loss in concept recovery\. Compared tomissingsetting, the substantially larger task degradation underflipsfurther indicated that relational evidence was most important when the bottleneck contained actively misleading states\.

### 4\.4Uncertainty Intervention

We evaluated whether identifying unreliable concepts enabled relational recovery without providing their true values\. We first corrupted concept predictions using a sampling ratio of0\.50\.5by replacing each sampled binary prediction with its opposite value\. We then marked an increasing fraction of the corrupted entries as unreliable\. For each marked entry, its uncertainty was set tou=1u=1, while its corrupted probability remained unchanged\. Recovery therefore relied on the remaining reliable concepts and their relations\.

![Refer to caption](https://arxiv.org/html/2608.10004v1/x4.png)Figure 4:Uncertainty\-guided recovery onWBC,CUB, andSynthetic\. Dashed lines denote performance before recovery, and solid lines denote performance as increasing fractions of corrupted entries were identified as unreliable\.Figure[4](https://arxiv.org/html/2608.10004#S4.F4)shows that both concept and task accuracy improved as more corrupted entries were identified as unreliable, demonstrating that reliability annotations alone could support relational recovery\. This interface is useful when an annotator or monitoring system can detect an unreliable concept more readily than determine its correct value\. Further experiments in the supplementary material examine how the assigned uncertainty controls the influence of provided concept edits\.

### 4\.5Sufficient Concept Set Extraction

We evaluated whether compact concept subsets could preserve the concept and task performance of ReCBM\. Following the extraction procedure in Section[3\.5](https://arxiv.org/html/2608.10004#S3.SS5), selection was performed exclusively on the validation set, yielding one global set shared by all classes and a separate set for each class\. Each selected set was subsequently evaluated once on the test set\. For class\-specific evaluation, only the selected concepts were supplied with their ground\-truth values\.

Table[5](https://arxiv.org/html/2608.10004#S4.T5)summarizes the global and class\-specific extraction results\. For the global set, Size gives the number of selected concepts relative to the full concept vocabulary, and Task\-Ret gives its test task accuracy as a percentage of that achieved by the full ReCBM\. For the class\-specific sets, Valid reports the number of classes for which the selected subset retained at least100%100\\%of both the full\-model concept F1 and task accuracy on the validation set\. Mean\-Size and BAcc report, respectively, the average subset size and macro one\-vs\-rest balanced accuracy over these valid classes\.

DatasetGlobal SetClass\-Specific SetsSizeTask\-RetValidMean\-SizeBAccWBC18/2418/2499\.90%99\.90\\%5/55/56\.60/246\.60/2498\.67%98\.67\\%CUB99/11299/112101\.03%101\.03\\%166/200166/2008\.42/1128\.42/11297\.38%97\.38\\%Synthetic12/1212/12100\.03%100\.03\\%4/44/410\.50/1210\.50/1298\.56%98\.56\\%Table 5:Global and class\-specific concept sets selected on the validation set and evaluated on the test set\.Table[5](https://arxiv.org/html/2608.10004#S4.T5)reveals global concept redundancy onWBCandCUB, where relational refinement reconstructs unselected concepts without reducing task performance\. In contrast,Syntheticrequires the complete concept set, indicating less global redundancy\.

Class\-specific extraction produces substantially smaller sets while achieving high one\-vs\-rest balanced accuracy, indicating that individual classes can be supported by compact, class\-dependent concept subsets\.

## 5Conclusion

In this work, we introduced ReCBM, an uncertainty\-gated relational reasoning framework that refines concept predictions through semantically defined relations\. By controlling evidence propagation according to concept reliability, ReCBM improved concept and task robustness under missing and flipped concepts\. It also supported uncertainty\-only recovery without ground\-truth concept values and identified compact class\-specific concept subsets\. Future work will explore iterative human\-in\-the\-loop refinement and relations derived from external knowledge\.

## References

- K\. Chauhan, R\. Tiwari, J\. Freyberg, P\. Shenoy, and K\. Dvijotham \(2023\)Interactive concept bottleneck models\.InProceedings of the Thirty\-Seventh AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 5948–5955\.External Links:[Link](https://doi.org/10.1609/aaai.v37i5.25736),[Document](https://dx.doi.org/10.1609/aaai.v37i5.25736)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Deng, N\. Ding, Y\. Jia, A\. Frome, K\. Murphy, S\. Bengio, Y\. Li, H\. Neven, and H\. Adam \(2014\)Large\-scale object classification using label relation graphs\.InComputer Vision – ECCV 2014,Cham,pp\. 48–64\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p2.1)\.
- Y\. Gao, Z\. Gao, X\. Gao, Y\. Liu, B\. Wang, and X\. Zhuang \(2024\)Evidential concept embedding models: towards reliable concept explanations for skin disease diagnosis\.InMedical Image Computing and Computer Assisted Intervention – MICCAI 2024,Cham,pp\. 308–317\.Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px2.p1.1),[§3\.4](https://arxiv.org/html/2608.10004#S3.SS4.SSS0.Px1.p1.9)\.
- M\. Havasi, S\. Parbhoo, and F\. Doshi\-Velez \(2022\)Addressing leakage in concept bottleneck models\.InAdvances in Neural Information Processing Systems,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),Vol\.35,pp\. 23386–23397\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/file/944ecf65a46feb578a43abfd5cddd960-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- E\. Kim, D\. Jung, S\. Park, S\. Kim, and S\. Yoon \(2023\)Probabilistic concept bottleneck models\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 16521–16540\.External Links:[Link](https://proceedings.mlr.press/v202/kim23g.html)Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.15.15.6),[§1](https://arxiv.org/html/2608.10004#S1.p2.1),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx2.p1.1)\.
- P\. W\. Koh, T\. Nguyen, Y\. S\. Tang, S\. Mussmann, E\. Pierson, B\. Kim, and P\. Liang \(2020\)Concept bottleneck models\.InProceedings of the 37th International Conference on Machine Learning,H\. D\. III and A\. Singh \(Eds\.\),Vol\.119,pp\. 5338–5348\.Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.5.5.6),[§1](https://arxiv.org/html/2608.10004#S1.p1.1),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx1.p3.1),[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx2.p1.1)\.
- Y\. LeCun, Y\. Bengio, and G\. Hinton \(2015\)Deep learning\.Nature521\(7553\),pp\. 436–444\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p1.1)\.
- X\. Li, F\. Zhao, and Y\. Guo \(2015\)Conditional Restricted Boltzmann Machines for Multi\-label Learning with Incomplete Labels\.InProceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics,Proceedings of Machine Learning Research, Vol\.38,San Diego, California, USA,pp\. 635–643\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p2.1)\.
- Z\. C\. Lipton \(2018\)The mythos of model interpretability\.Communications of the ACM61\(10\),pp\. 36–43\.External Links:ISSN 0001\-0782,[Link](https://doi.org/10.1145/3233231),[Document](https://dx.doi.org/10.1145/3233231)Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p1.1)\.
- A\. Mahinpei, J\. Clark, I\. Lage, F\. Doshi\-Velez, and W\. Pan \(2021\)Promises and pitfalls of black\-box concept learning models\.External Links:2106\.13314,[Link](https://arxiv.org/abs/2106.13314)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Margeloiu, M\. Ashman, U\. Bhatt, Y\. Chen, M\. Jamnik, and A\. Weller \(2021\)Do concept bottleneck models learn as intended?\.External Links:2105\.04289,[Link](https://arxiv.org/abs/2105.04289)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- T\. Oikarinen, S\. Das, L\. M\. Nguyen, and T\. Weng \(2023\)Label\-free concept bottleneck models\.InThe Eleventh International Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Park, J\. Mun, D\. Oh, and N\. Lee \(2025\)An analysis of concept bottleneck models: measuring, understanding, and mitigating the impact of noisy annotations\.InAdvances in Neural Information Processing Systems,D\. Belgrave, C\. Zhang, H\. Lin, R\. Pascanu, P\. Koniusz, M\. Ghassemi, and N\. Chen \(Eds\.\),Vol\.38,pp\. 5297–5331\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p1.1)\.
- N\. J\. Raman, M\. E\. Zarlenga, J\. Heo, and M\. Jamnik \(2025\)Do concept bottleneck models respect localities?\.Transactions on Machine Learning Research\.Note:External Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=4mCkRbUXOf)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- C\. Rudin \(2019\)Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead\.Nature Machine Intelligence1\(5\),pp\. 206–215\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p1.1)\.
- M\. Sensoy, L\. Kaplan, and M\. Kandemir \(2018\)Evidential deep learning to quantify classification uncertainty\.InAdvances in Neural Information Processing Systems,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\. 3183–3193\.Cited by:[§3\.1](https://arxiv.org/html/2608.10004#S3.SS1.p2.8),[§3\.4](https://arxiv.org/html/2608.10004#S3.SS4.SSS0.Px1.p1.9)\.
- S\. Shin, Y\. Jo, S\. Ahn, and N\. Lee \(2023\)A closer look at the intervention procedure of concept bottleneck models\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 31504–31520\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p1.1)\.
- S\. Tsutsui, W\. Pang, S\. He, and B\. Wen \(2026\)WBCAtt\+: fine\-grained pixel\-level morphological annotations for white blood cell images\.Medical Image Analysis112,pp\. 104137\.External Links:ISSN 1361\-8415Cited by:[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx1.p2.1)\.
- M\. Vandenhirtz, S\. Laguna, R\. Marcinkevičs, and J\. E\. Vogt \(2024\)Stochastic concept bottleneck models\.InAdvances in Neural Information Processing Systems,Vol\.37,pp\. 51787–51810\.Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.20.20.6),[§1](https://arxiv.org/html/2608.10004#S1.p2.1),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx2.p1.1)\.
- C\. Wah, S\. Branson, P\. Welinder, P\. Perona, and S\. Belongie \(2011\)The Caltech\-UCSD Birds\-200\-2011 dataset\.Technical reportTechnical ReportCNS\-TR\-2011\-001,California Institute of Technology\.Cited by:[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx1.p3.1)\.
- E\. B\. Wilson \(1927\)Probable inference, the law of succession, and statistical inference\.Journal of the American Statistical Association22\(158\),pp\. 209–212\.Cited by:[Appendix B](https://arxiv.org/html/2608.10004#A2.p1.10)\.
- B\. Xiong, M\. Cochez, M\. Nayyeri, and S\. Staab \(2022\)Hyperbolic embedding inference for structured multi\-label prediction\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 33016–33028\.Cited by:[§1](https://arxiv.org/html/2608.10004#S1.p2.1)\.
- H\. Xu, T\. Weng, L\. M\. Nguyen, and T\. Ma \(2026\)Graph concept bottleneck models\.External Links:2508\.14255,[Link](https://arxiv.org/abs/2508.14255)Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.30.30.6),[§1](https://arxiv.org/html/2608.10004#S1.p2.1),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.10004#S4.SS1.SSSx2.p1.1)\.
- X\. Xu, Y\. Qin, L\. Mi, H\. Wang, and X\. Li \(2024\)Energy\-based concept bottleneck models: unifying prediction, concept intervention, and probabilistic interpretations\.InThe Twelfth International Conference on Learning Representations,Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.25.25.6),[§1](https://arxiv.org/html/2608.10004#S1.p2.1),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Yuksekgonul, M\. Wang, and J\. Zou \(2023\)Post\-hoc concept bottleneck models\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nA5AZ8CEyow)Cited by:[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- M\. E\. Zarlenga, P\. Barbiero, G\. Ciravegna, G\. Marra, F\. Giannini, M\. Diligenti, Z\. Shams, F\. Precioso, S\. Melacci, A\. Weller, P\. Lio, and M\. Jamnik \(2022\)Concept embedding models\.InAdvances in Neural Information Processing Systems,A\. H\. Oh, A\. Agarwal, D\. Belgrave, and K\. Cho \(Eds\.\),Vol\.35,pp\. 21400–21413\.Cited by:[Table 1](https://arxiv.org/html/2608.10004#S1.T1.10.10.6),[§2](https://arxiv.org/html/2608.10004#S2.SS0.SSS0.Px1.p1.1)\.
- K\. Zou, Y\. Chen, L\. Huang, N\. Zhou, X\. Yuan, X\. Shen, M\. Wang, R\. S\. M\. Goh, Y\. Liu, Y\. C\. Tham, and H\. Fu \(2025\)Toward reliable medical image segmentation by modeling evidential calibrated uncertainty\.IEEE Transactions on Cybernetics55\(12\),pp\. 5975–5988\.External Links:[Document](https://dx.doi.org/10.1109/TCYB.2025.3604432)Cited by:[§3\.4](https://arxiv.org/html/2608.10004#S3.SS4.SSS0.Px1.p1.11)\.

Supplementary Material

## Contents

Guide to the Supplementary Material\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[Guide to the Supplementary Material](https://arxiv.org/html/2608.10004#Ax1)

Appendix A: Synthetic Dataset Details\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A](https://arxiv.org/html/2608.10004#A1)

Appendix B: Relation\-Matrix Initialization\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B](https://arxiv.org/html/2608.10004#A2)

Appendix C: Preserving Relation Semantics during Optimization\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C](https://arxiv.org/html/2608.10004#A3)

Appendix D: Implementation Details\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D](https://arxiv.org/html/2608.10004#A4)

Appendix E: ReCBM Training Pseudocode\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[E](https://arxiv.org/html/2608.10004#A5)

Appendix F: Complete Ablation Results\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[F](https://arxiv.org/html/2608.10004#A6)

Appendix G: Qualitative Refinement Analysis\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[G](https://arxiv.org/html/2608.10004#A7)

Appendix H: Uncertainty\-Only Intervention\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[H](https://arxiv.org/html/2608.10004#A8)

Appendix I: Counterfactual Edits under Different Uncertainty Levels\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[I](https://arxiv.org/html/2608.10004#A9)

Appendix J: Sufficient Concept\-Set Extraction\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[J](https://arxiv.org/html/2608.10004#A10)

## Guide to the Supplementary Material

The supplementary material is organized as follows\. Appendix A specifies the complete Synthetic dataset generator\. Appendix B describes relation\-matrix initialization and presents representative learned relations\. Appendix C describes how the semantic consistency of the relation matrices is preserved during optimization\. Appendix D reports implementation details, including training, corruption, hyperparameters for ReCBM and the comparison methods, random seeds, and the computing environment\. Appendix E gives the complete ReCBM training pseudocode\. Appendix F reports ablation results on all datasets\. Appendix G provides qualitative examples of concept refinement and uncertainty\-only intervention\. Appendix H describes the uncertainty\-only intervention protocol\. Appendix I studies how assigned uncertainty controls counterfactual concept edits\. Finally, Appendix J describes the complete procedure for extracting and pruning sufficient concept sets, together with the resulting class\-specific sets\.

## Appendix ASynthetic Dataset Details

This section documents the Synthetic dataset construction to clarify how its concepts, task labels, and relational structure were generated\. The Synthetic dataset represents access\-control events with 12 binary concepts and 4 task labels\. Table[6](https://arxiv.org/html/2608.10004#A1.T6)defines the complete concept and task\-label vocabularies\. The concepts describe authentication outcome, device and location status, failed\-login behavior, account privilege, and policy enforcement\. The task labels summarize four event types: normal access, credential anomaly, device or location anomaly, and policy block or privilege risk\.

IndexConcept0Password verified1Multi\-Factor Authentication \(MFA\) verified2Registered device3New device4Usual location5Anomalous location6Failed attempts reach threshold7Account locked8Login succeeds9Administrator account10Administrator operation11Login blocked by policyIndexTask label0Normal access1Credential anomaly2Device or location anomaly3Policy block or privilege riskTable 6:Concept and task\-label vocabularies of the Synthetic dataset\.We sampled task labels uniformly and drew concepts independently from label\-dependent Bernoulli distributions\. Table[7](https://arxiv.org/html/2608.10004#A1.T7)reports the complete probability vector in the order of the concept indices\. We then imposed the relations in Table[8](https://arxiv.org/html/2608.10004#A1.T8): for each co\-occurring pair, activation of one concept activated the other with probability0\.850\.85; implications were enforced by activating the consequent whenever the antecedent was active; and exclusion conflicts were resolved using the class\-dependent rules described below\. This procedure produced a concept vector consistent with the specified relational structure\.

Labelc0c\_\{0\}c1c\_\{1\}c2c\_\{2\}c3c\_\{3\}c4c\_\{4\}c5c\_\{5\}c6c\_\{6\}c7c\_\{7\}c8c\_\{8\}c9c\_\{9\}c10c\_\{10\}c11c\_\{11\}Normal access\.98\.96\.92\.02\.90\.02\.01\.01\.94\.06\.02\.01Credential anomaly\.35\.08\.48\.08\.44\.08\.96\.94\.01\.03\.02\.92Device/location anomaly\.92\.88\.02\.94\.02\.92\.04\.02\.78\.03\.02\.03Policy block/privilege risk\.62\.56\.16\.18\.14\.16\.04\.03\.01\.96\.90\.82Table 7:Label\-conditional Bernoulli probabilities used before applying the relations\. Concept indices follow Table[6](https://arxiv.org/html/2608.10004#A1.T6)\.After independently sampling the 12 initial concept states from Table[7](https://arxiv.org/html/2608.10004#A1.T7), the generator applied the relations in the following order\.

1. 1\.For each co\-occurrence pair, if exactly one concept was active, the other was activated with probability0\.850\.85\.
2. 2\.All implication rules were applied repeatedly until no implication could activate an additional concept\. For example, an active “Failed attempts reach threshold” concept activated “Account locked,” which subsequently activated “Login blocked by policy\.”
3. 3\.For each exclusion pair, mutual exclusivity was enforced by retaining one active concept according to a class\-conditional selection rule\. If both “Registered device” and “New device” were active, the latter was retained for the device/location\-anomaly class\. For either of the other non\-normal classes, it was retained with probability0\.350\.35; otherwise, “Registered device” was retained\. The same rule was used for “Usual location” and “Anomalous location\.” Independently of the class, an active “Account locked” or “Login blocked by policy” concept deactivated “Login succeeds\.”
4. 4\.Implication closure and exclusion resolution were applied once more because resolving one relation could change whether another relation was satisfied\. The resulting vector was used as the final concept annotation\.

The training, validation, and test splits contained 12,000, 3,000, and 3,000 examples, respectively\. The generator mapped each final concept vector to a 32\-dimensional input in two sampling stages\. First, it sampled one projection matrix and four class\-specific offset vectors:

𝐖d​k\\displaystyle\\mathbf\{W\}\_\{dk\}∼i\.i\.d\.​𝒩​\(0,1\),\\displaystyle\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\mathcal\{N\}\(0,1\),d∈\{1,…,32\},k∈\{1,…,12\},\\displaystyle d\\in\\\{1,\\ldots,32\\\},\\quad k\\in\\\{1,\\ldots,12\\\},\(26\)𝐞y\\displaystyle\\mathbf\{e\}\_\{y\}∼i\.i\.d\.​𝒩​\(𝟎,σe2​𝐈\),\\displaystyle\\overset\{\\mathrm\{i\.i\.d\.\}\}\{\\sim\}\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{e\}^\{2\}\\mathbf\{I\}\),y∈\{0,1,2,3\},\\displaystyle y\\in\\\{0,1,2,3\\\},\(27\)where𝐖∈ℝ32×12\\mathbf\{W\}\\in\\mathbb\{R\}^\{32\\times 12\},𝐞y∈ℝ32\\mathbf\{e\}\_\{y\}\\in\\mathbb\{R\}^\{32\}, and𝐈∈ℝ32×32\\mathbf\{I\}\\in\\mathbb\{R\}^\{32\\times 32\}is the identity matrix\. We setσe=0\.35\\sigma\_\{e\}=0\.35\. The sampled𝐖\\mathbf\{W\}and the four vectors\{𝐞y\}y=03\\\{\\mathbf\{e\}\_\{y\}\\\}\_\{y=0\}^\{3\}were fixed and shared by the training, validation, and test splits\.

Second, for every examplenn, the generator independently sampled

ϵn∼𝒩​\(𝟎,σϵ2​𝐈\),σϵ=0\.35,\\boldsymbol\{\\epsilon\}\_\{n\}\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\sigma\_\{\\epsilon\}^\{2\}\\mathbf\{I\}\),\\qquad\\sigma\_\{\\epsilon\}=0\.35,and constructed the unstandardized input

𝐱~n=𝐖𝐜n\+𝐞yn\+ϵn\.\\widetilde\{\\mathbf\{x\}\}\_\{n\}=\\mathbf\{W\}\\mathbf\{c\}\_\{n\}\+\\mathbf\{e\}\_\{y\_\{n\}\}\+\\boldsymbol\{\\epsilon\}\_\{n\}\.\(28\)Here,𝐜n∈\{0,1\}12\\mathbf\{c\}\_\{n\}\\in\\\{0,1\\\}^\{12\}andyn∈\{0,1,2,3\}y\_\{n\}\\in\\\{0,1,2,3\\\}are the concept vector and task label of examplenn, respectively\. The term𝐖𝐜n\\mathbf\{W\}\\mathbf\{c\}\_\{n\}encodes the concept state,𝐞yn\\mathbf\{e\}\_\{y\_\{n\}\}adds the offset shared by all examples of classyny\_\{n\}, andϵn\\boldsymbol\{\\epsilon\}\_\{n\}introduces example\-specific variation\.

Finally, each splits∈\{train,validation,test\}s\\in\\\{\\mathrm\{train\},\\mathrm\{validation\},\\mathrm\{test\}\\\}was standardized independently in every feature dimension:

xn,d=x~n,d−μs,dσs,d\+10−6,x\_\{n,d\}=\\frac\{\\widetilde\{x\}\_\{n,d\}\-\\mu\_\{s,d\}\}\{\\sigma\_\{s,d\}\+10^\{\-6\}\},\(29\)whereμs,d\\mu\_\{s,d\}andσs,d\\sigma\_\{s,d\}are the empirical mean and standard deviation of dimensionddin splitss\. All random quantities were generated using a NumPy random generator initialized with seed 42\.

RelationConcept pairsCo\-occurrencePassword verified↔\\leftrightarrowMFA verifiedRegistered device↔\\leftrightarrowUsual locationNew device↔\\leftrightarrowAnomalous locationFailed attempts reach threshold↔\\leftrightarrowAccount lockedAdministrator account↔\\leftrightarrowAdministrator operationImplicationLogin succeeds→\\rightarrowPassword verifiedLogin succeeds→\\rightarrowMFA verifiedFailed attempts reach threshold→\\rightarrowAccount lockedAccount locked→\\rightarrowLogin blocked by policyAdministrator operation→\\rightarrowAdministrator accountExclusionRegistered device⟂\\perpNew deviceUsual location⟂\\perpAnomalous locationAccount locked⟂\\perpLogin succeedsLogin blocked by policy⟂\\perpLogin succeedsTable 8:Ground\-truth relations in the Synthetic dataset\.
## Appendix BRelation\-Matrix Initialization

This section describes the complete initialization procedure\. For a pair of binary concepts\(ci,cj\)\(c\_\{i\},c\_\{j\}\), letν^a​bi​j=na​bi​j/Ni​j\\hat\{\\nu\}\_\{ab\}^\{ij\}=n\_\{ab\}^\{ij\}/N\_\{ij\}denote the empirical frequency of\(ci,cj\)=\(a,b\)\(c\_\{i\},c\_\{j\}\)=\(a,b\)among theNi​jN\_\{ij\}training examples in which both concepts are observed, wherea,b∈\{0,1\}a,b\\in\\\{0,1\\\},na​bi​jn\_\{ab\}^\{ij\}is the corresponding joint count, andi,j∈\{1,…,C\}i,j\\in\\\{1,\\ldots,C\\\}index concepts\. Under independence, the expected frequencies are

ν11i​j=pi​pj,ν10i​j=pi​\(1−pj\),ν01i​j=\(1−pi\)​pj,\\nu\_\{11\}^\{ij\}=p\_\{i\}p\_\{j\},\\qquad\\nu\_\{10\}^\{ij\}=p\_\{i\}\(1\-p\_\{j\}\),\\qquad\\nu\_\{01\}^\{ij\}=\(1\-p\_\{i\}\)p\_\{j\},\(30\)wherepi=P​\(ci=1\)p\_\{i\}=P\(c\_\{i\}=1\)andpj=P​\(cj=1\)p\_\{j\}=P\(c\_\{j\}=1\)\. The normalized reduction in violation frequency and its expected\-support factor are

ρa​bi​j=\[νa​bi​j−UCB​\(ν^a​bi​j\)max⁡\(νa​bi​j,ϵ\)\]\+,Ra​bi​j=Ni​j​νa​bi​jNi​j​νa​bi​j\+τ\+ϵ\.\\rho\_\{ab\}^\{ij\}=\\left\[\\frac\{\\nu\_\{ab\}^\{ij\}\-\\mathrm\{UCB\}\(\\hat\{\\nu\}\_\{ab\}^\{ij\}\)\}\{\\max\(\\nu\_\{ab\}^\{ij\},\\epsilon\)\}\\right\]\_\{\+\},\\qquad R\_\{ab\}^\{ij\}=\\frac\{N\_\{ij\}\\nu\_\{ab\}^\{ij\}\}\{N\_\{ij\}\\nu\_\{ab\}^\{ij\}\+\\tau\+\\epsilon\}\.\(31\)For an empirical frequencyν^=n/N\\hat\{\\nu\}=n/N, the Wilson upper confidence bound\(Wilson[1927](https://arxiv.org/html/2608.10004#bib.bib25)\)used in our implementation is

UCB​\(ν^\)=ν^\+z22​N\+z​ν^​\(1−ν^\)N\+z24​N21\+z2N\.\\mathrm\{UCB\}\(\\hat\{\\nu\}\)=\\frac\{\\hat\{\\nu\}\+\\frac\{z^\{2\}\}\{2N\}\+z\\sqrt\{\\frac\{\\hat\{\\nu\}\(1\-\\hat\{\\nu\}\)\}\{N\}\+\\frac\{z^\{2\}\}\{4N^\{2\}\}\}\}\{1\+\\frac\{z^\{2\}\}\{N\}\}\.\(32\)We usedz=1\.96z=1\.96, corresponding to a two\-sided 95% interval, and set the support constantτ=5\\tau=5and numerical constantϵ=10−6\\epsilon=10^\{\-6\}\. The operator\[x\]\+=max⁡\(x,0\)\[x\]\_\{\+\}=\\max\(x,0\)\. The initial relation strengths are

Ai​jco\\displaystyle A\_\{ij\}^\{\\mathrm\{co\}\}=ρ10i​j​ρ01i​j​R10i​j​R01i​j,\\displaystyle=\{\}\\sqrt\{\\rho\_\{10\}^\{ij\}\\rho\_\{01\}^\{ij\}R\_\{10\}^\{ij\}R\_\{01\}^\{ij\}\},Ai​jimp\\displaystyle A\_\{ij\}^\{\\mathrm\{imp\}\}=\[ρ10i​j−ρ01i​j\]\+​R10i​j,\\displaystyle=\{\}\[\\rho\_\{10\}^\{ij\}\-\\rho\_\{01\}^\{ij\}\]\_\{\+\}R\_\{10\}^\{ij\},Ai​jexc\\displaystyle A\_\{ij\}^\{\\mathrm\{exc\}\}=ρ11i​j​R11i​j\.\\displaystyle=\{\}\\rho\_\{11\}^\{ij\}R\_\{11\}^\{ij\}\.\(33\)Co\-occurrence and exclusion were symmetrized by averaging each matrix with its transpose\. All diagonal entries were set to zero, and all strengths were clipped to\[0,1\]\[0,1\]\. Co\-occurrence required both one\-sided mismatch cells to be rarer than independence predicted, implication required the forward violation cell to be selectively rare, and exclusion required joint activation to be rare\. These values initialized trainable relation parameters\. The matrices were subsequently optimized with the rest of the refinement module\.

##### Representative Learned Relations

Table[9](https://arxiv.org/html/2608.10004#A2.T9)lists the two learned relations with the highest weights for each relation type and dataset\.

DatasetTypeStrongest learned relationsWBCCo\-occurrenceround granule type↔\\leftrightarrowred granule colour \(0\.993\)nil granule type↔\\leftrightarrownil granule colour \(0\.991\)Implicationcoarse granule type→\\rightarrowgranularity \(0\.928\)segmented multilobed nucleus shape→\\rightarrowgranularity \(0\.928\)Exclusionnil granule type⟂\\perpgranularity \(0\.991\)nil granule colour⟂\\perpgranularity \(0\.991\)CUBCo\-occurrencebelly color: yellow↔\\leftrightarrowprimary color: yellow \(0\.971\)underparts color: yellow↔\\leftrightarrowprimary color: yellow \(0\.971\)Implicationforehead color: yellow→\\rightarrowbill length: shorter than head \(0\.919\)forehead color: yellow→\\rightarrowshape: perching\-like \(0\.913\)Exclusionbill length: about the same as head⟂\\perpbill length: shorter than head \(0\.994\)size: small⟂\\perpsize: medium \(0\.987\)SyntheticCo\-occurrenceFailed attempts reach threshold↔\\leftrightarrowAccount locked \(0\.993\)Administrator account↔\\leftrightarrowAdministrator operation \(0\.987\)ImplicationLogin succeeds→\\rightarrowPassword verified \(0\.915\)Login succeeds→\\rightarrowMFA verified \(0\.909\)ExclusionLogin succeeds⟂\\perpLogin blocked by policy \(0\.998\)Registered device⟂\\perpNew device \(0\.997\)Table 9:Highest\-weight learned concept relations\. Values in parentheses are the corresponding learned relation weights\.

## Appendix CPreserving Relation Semantics during Optimization

To preserve the intended distinction among co\-occurrence, implication, and exclusion during optimization, we used a logic\-conflict regularizer\. LetCCbe the number of concepts and let𝐀co,𝐀imp,𝐀exc∈\[0,1\]C×C\\mathbf\{A\}^\{\\mathrm\{co\}\},\\mathbf\{A\}^\{\\mathrm\{imp\}\},\\mathbf\{A\}^\{\\mathrm\{exc\}\}\\in\[0,1\]^\{C\\times C\}denote the three relation matrices defined in the main paper\. For matrices𝐗,𝐘\\mathbf\{X\},\\mathbf\{Y\},𝐗⊙𝐘\\mathbf\{X\}\\odot\\mathbf\{Y\}denotes element\-wise multiplication,𝐗⊤\\mathbf\{X\}^\{\\top\}denotes transpose, and⟨𝐗⟩=C−2​∑i=1C∑j=1CXi​j\\langle\\mathbf\{X\}\\rangle=C^\{\-2\}\\sum\_\{i=1\}^\{C\}\\sum\_\{j=1\}^\{C\}X\_\{ij\}denotes the mean of all entries\. We write𝟏∈ℝC×C\\mathbf\{1\}\\in\\mathbb\{R\}^\{C\\times C\}for the all\-ones matrix\. The regularizer is

ℒlogic=\\displaystyle\\mathcal\{L\}\_\{\\mathrm\{logic\}\}=\{\}⟨𝐀co⊙𝐀exc⟩\+⟨\(𝐀imp\+𝐀imp⊤\)⊙𝐀exc⟩\\displaystyle\\left\\langle\\mathbf\{A\}^\{\\mathrm\{co\}\}\\odot\\mathbf\{A\}^\{\\mathrm\{exc\}\}\\right\\rangle\+\\left\\langle\(\\mathbf\{A\}^\{\\mathrm\{imp\}\}\+\\mathbf\{A\}^\{\\mathrm\{imp\}\\top\}\)\\odot\\mathbf\{A\}^\{\\mathrm\{exc\}\}\\right\\rangle\+⟨\(𝐀imp⊙𝐀imp⊤\)⊙\(𝟏−𝐀co\)⟩\.\\displaystyle\+\\left\\langle\(\\mathbf\{A\}^\{\\mathrm\{imp\}\}\\odot\\mathbf\{A\}^\{\\mathrm\{imp\}\\top\}\)\\odot\(\\mathbf\{1\}\-\\mathbf\{A\}^\{\\mathrm\{co\}\}\)\\right\\rangle\.\(34\)The first term penalizes overlap between co\-occurrence and exclusion for the same concept pair\. The second penalizes overlap between either direction of an implication and exclusion\. The third maps strong bidirectional implication to co\-occurrence by penalizing disagreement between the two relation types\.

Stage 1 optimizedhϑh\_\{\\vartheta\}using raw concept probabilities\. At the beginning of Stage 2, the learned parametersϑ\\varthetainitialized the task predictor applied to refined concept probabilities\. The task predictor and relational refinement module were subsequently optimized using the complete Stage 2 objective:

ℒstage2complete=ℒstage2\+λlogic​ℒlogic,\\mathcal\{L\}\_\{\\mathrm\{stage2\}\}^\{\\mathrm\{complete\}\}=\\mathcal\{L\}\_\{\\mathrm\{stage2\}\}\+\\lambda\_\{\\mathrm\{logic\}\}\\mathcal\{L\}\_\{\\mathrm\{logic\}\},\(35\)where we setλlogic=0\.1\\lambda\_\{\\mathrm\{logic\}\}=0\.1\.

## Appendix DImplementation Details

This section provides the implementation and evaluation settings needed to reproduce the experiments\. The three datasets were selected to cover complementary experimental settings\. WBC provided a compact concept vocabulary for evaluating relational refinement on medical images with interpretable morphological attributes\. CUB provided a larger setting with 112 concepts and 200 fine\-grained classes, allowing us to evaluate the method with a substantially larger concept and task space\. Synthetic provided known concept relations and a controlled generation process, allowing relational recovery to be examined when the underlying co\-occurrence, implication, and exclusion structures were explicitly defined\.

For WBC and CUB, the encoderfθf\_\{\\theta\}was a ResNet\-34, and images were resized to224×224224\\times 224\. For Synthetic,fθf\_\{\\theta\}was an MLP with hidden dimension 128\.

ReCBM was trained for 150 epochs\. Stage 1 lasted 70 epochs and jointly trained the encoderfθf\_\{\\theta\}, evidential concept predictorgψg\_\{\\psi\}, and task predictorhϑh\_\{\\vartheta\}\. At the beginning of Stage 2, the task predictor was initialized from Stage 1\. The encoder and evidential concept predictor were then frozen\. During the remaining 80 epochs, the relation module and task predictor were optimized\.

### D\.1Training and Hyperparameter Configuration

This subsection specifies the optimization schedule and the final settings used for each dataset\. Hyperparameters were selected on the validation split by task accuracy, with concept accuracy used to break ties\. Refinement configurations were evaluated under the corrupted validation protocol described in Section[D\.3](https://arxiv.org/html/2608.10004#A4.SS3)\. Table[10](https://arxiv.org/html/2608.10004#A4.T10)reports the candidate values and final configuration for each dataset\. In Stage 1, the encoder and concept predictor used the dataset\-specific initial learning rate in Table[10](https://arxiv.org/html/2608.10004#A4.T10), and the task predictor used the same rate\. In Stage 2, the relation module and task predictor started at half this rate\. Stage 1 and Stage 2 used independent cosine annealing schedules with horizons of 70 and 80 epochs, respectively, and a minimum learning rate of zero\.

HyperparameterCandidate valuesWBCCUBSyntheticInitial learning rate\{5×10−4,10−3\}\\\{5\{\\times\}10^\{\-4\},10^\{\-3\}\\\}5×10−45\{\\times\}10^\{\-4\}5×10−45\{\\times\}10^\{\-4\}10−310^\{\-3\}Batch size\{64,192,256\}\\\{64,192,256\\\}6464256Refinement iterations\{5,10,15\}\\\{5,10,15\\\}5105Stage 1 epochs\{50,70,100\}\\\{50,70,100\\\}707070Total epochs\{100,150\}\\\{100,150\\\}150150150Stage 2 concept\-state augmentation ratio\{0,0\.1\}\\\{0,0\.1\\\}0\.10\.10\.1Natural refined\-concept BCE weight\{5,10,20\}\\\{5,10,20\\\}201010Augmented refined\-concept BCE weight\{1,5,10\}\\\{1,5,10\\\}51010Endpoint direction\-loss weight\{1,10\}\\\{1,10\\\}101010Message direction\-loss weight\{1,10\}\\\{1,10\\\}101010Uncertainty\-error loss weight\{1,10\}\\\{1,10\\\}101010Logic\-conflict weightλlogic\\lambda\_\{\\mathrm\{logic\}\}\{0\.1,1\}\\\{0\.1,1\\\}0\.10\.10\.1OptimizerAdamWAdamWAdamWAdamWWeight decay5×10−65\{\\times\}10^\{\-6\}5×10−65\{\\times\}10^\{\-6\}5×10−65\{\\times\}10^\{\-6\}5×10−65\{\\times\}10^\{\-6\}Learning\-rate schedulercosine annealingcosine annealingcosine annealingcosine annealingKL coefficient / annealing horizon1/101/10epochs1/101/10epochs1/101/10epochs1/101/10epochsStage 1 / Stage 2 task\-loss weights1/11/11/11/11/11/11/11/1Concept\-loss weight1111Minimum support countτ\\tau5555Wilson interval normal quantile1\.961\.961\.961\.96Table 10:Candidate hyperparameter values and final configurations selected on the validation split for each dataset\.
### D\.2Comparison Method Configuration

The comparison methods used the same dataset splits, input preprocessing, backbones, and batch sizes as ReCBM\. Image models used ResNet\-34, and models for Synthetic used the MLP encoder with hidden dimension 128\. Table[11](https://arxiv.org/html/2608.10004#A4.T11)reports the final training configuration of each comparison method\.

MethodOptimizerLearning rateTraining epochsAdditional configurationIndependent CBMSGD0\.01/0\.010\.01/0\.01100/100100/100Concept and task stages used weight decay4×10−54\{\\times\}10^\{\-5\}\. The StepLR step size was 1000 with a decay factor of 0\.1\.Joint CBMSGD0\.010\.01100Weight decay was4×10−54\{\\times\}10^\{\-5\}\. The concept loss weight was 0\.01, and the normalized joint loss was used\. The StepLR step size was 1000 with a decay factor of 0\.1\.ProbCBMAdamP10−310^\{\-3\}50/2050/20The concept and task stages lasted 50 and 20 epochs\. The first five epochs used encoder warmup\. Newly introduced parameter groups used a learning rate multiplier of 10\. The model used 50 Monte Carlo samples, concept hidden dimension 16, task hidden dimension 128, intervention probability 0\.5, and variational weight5×10−55\{\\times\}10^\{\-5\}\. Weight decay was zero, and each stage used cosine annealing\.SCBMAdam10−410^\{\-4\}100/300/300100/300/300The epoch counts correspond to WBC, CUB, and Synthetic, respectively\. SCBM used amortized covariance, 100 Monte Carlo samples, hard concept learning with the straight through estimator, and anℓ1\\ell\_\{1\}precision regularizer with weight 1\. Weight decay was zero\. StepLR used a step size of 150 and a decay factor of 0\.5\.GraphCBMAdam10−310^\{\-3\}150Weight decay was4×10−54\{\\times\}10^\{\-5\}\. The model used three graph layers, concept loss weight 1, graph regularization weight 0\.1, and gradient clipping at 0\.5\. Validation was performed every five epochs with patience 20\. The StepLR step size was 1000 with a decay factor of 0\.1\.Table 11:Final training configurations of the comparison methods\. For Independent CBM and ProbCBM, values separated by a slash correspond to the concept and task training stages\. For SCBM, the three epoch counts correspond to WBC, CUB, and Synthetic\.
### D\.3Corruption and Evaluation Protocol

For the corruption robustness experiments, the same entry\-level corruption mask was used for all methods or ablation variants compared in a run\. A missing entry was replaced by\(p,u\)=\(0\.5,1\)\(p,u\)=\(0\.5,1\)\. To flip an entry, we first converted its raw probability into a binary prediction using a threshold of 0\.5 and then reversed the prediction, changing 0 to 1 and 1 to 0\. Its uncertainty was subsequently set tou=1u=1\. Entries outside the corruption mask retained their original probabilities and uncertainties\. Thus, these experiments evaluated recovery when the locations of the corrupted concepts were explicitly identified as unreliable\.

Concept accuracy was computed by thresholding concept probabilities at 0\.5, whereas task accuracy was computed from the final class prediction\. Standard performance results used five seeds \(42\-46\)\. Ablations used three paired seeds \(0\-2\), shared corruption masks, and a 50% corruption ratio\.

### D\.4Random Seeds and Deterministic Execution

ReCBM training used the seed specified in each configuration through the PyTorch Lightning seeding utility\. The comparison method and evaluation utilities explicitly seeded Python, NumPy, PyTorch, and all available CUDA devices\. For evaluation, cuDNN benchmarking was disabled, deterministic cuDNN execution was enabled, and deterministic PyTorch operations were requested\.

### D\.5Computing Environment

Experiments were conducted on Ubuntu Linux using an AMD EPYC 7742 CPU, 256 GiB of system memory, and one NVIDIA GeForce RTX 3090 GPU with 24 GiB of memory\. The software environment consisted of Python 3\.9\.20, PyTorch 2\.5\.0, torchvision 0\.20\.0, PyTorch Lightning 2\.6\.0, CUDA 12\.4, cuDNN 9\.1, NumPy 2\.0\.1, and scikit\-learn 1\.5\.2\. The code archive contains a version\-pinned requirements file\.

## Appendix EReCBM Training Pseudocode

This section presents the ReCBM training procedure with two stages\. In Algorithm[1](https://arxiv.org/html/2608.10004#alg1),𝒟=\{\(xn,yn,𝐜n\)\}n=1N\\mathcal\{D\}=\\\{\(x\_\{n\},y\_\{n\},\\mathbf\{c\}\_\{n\}\)\\\}\_\{n=1\}^\{N\}is the training set,𝒜=\{𝐀co,𝐀imp,𝐀exc\}\\mathcal\{A\}=\\\{\\mathbf\{A\}^\{\\mathrm\{co\}\},\\mathbf\{A\}^\{\\mathrm\{imp\}\},\\mathbf\{A\}^\{\\mathrm\{exc\}\}\\\},E1E\_\{1\}andEEare the numbers of Stage 1 and total epochs, andTTis the number of refinement iterations\. The encoder, concept predictor, relation module, and task predictor are denoted byfθf\_\{\\theta\},gψg\_\{\\psi\},rϕr\_\{\\phi\}, andhϑh\_\{\\vartheta\}, respectively\. The concept predictor outputs raw probabilities𝐩raw\\mathbf\{p\}^\{\\mathrm\{raw\}\}and uncertainties𝐮\\mathbf\{u\}, while𝐩\(t\)\\mathbf\{p\}^\{\(t\)\}denotes the concept state at refinement iterationtt\.

Algorithm 1Training of ReCBM in two stages0:Data

𝒟\\mathcal\{D\}; relation matrices

𝒜\\mathcal\{A\}; Stage\-1 epochs

E1E\_\{1\}; total epochs

EE; refinement steps

TT
1:Initialize encoder

fθf\_\{\\theta\}, concept predictor

gψg\_\{\\psi\}, task predictor

hϑh\_\{\\vartheta\}, and relation module

rϕr\_\{\\phi\}
2:for

e=1,…,E1e=1,\\ldots,E\_\{1\}do

3:foreach minibatch

\(x,y,𝐜\)\(x,y,\\mathbf\{c\}\)do

4:

\(𝐩raw,𝐮\)←gψ​\(fθ​\(x\)\)\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{u\}\)\\leftarrow g\_\{\\psi\}\(f\_\{\\theta\}\(x\)\)
5:Update

\(θ,ψ,ϑ\)\(\\theta,\\psi,\\vartheta\)with the Stage\-1 loss

6:endfor

7:endfor

8:Retain the learned

ϑ\\varthetaas the Stage 2 initialization and freeze

fθf\_\{\\theta\}and

gψg\_\{\\psi\}
9:for

e=E1\+1,…,Ee=E\_\{1\}\+1,\\ldots,Edo

10:foreach minibatch

\(x,y,𝐜\)\(x,y,\\mathbf\{c\}\)do

11:

\(𝐩raw,𝐮\)←gψ​\(fθ​\(x\)\)\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{u\}\)\\leftarrow g\_\{\\psi\}\(f\_\{\\theta\}\(x\)\);

\(𝐩\(0\),𝐮\(0\)\)←Augment⁡\(𝐩raw,𝐮\)\(\\mathbf\{p\}^\{\(0\)\},\\mathbf\{u\}^\{\(0\)\}\)\\leftarrow\\operatorname\{Augment\}\(\\mathbf\{p\}^\{\\mathrm\{raw\}\},\\mathbf\{u\}\)
12:for

t=0,…,T−1t=0,\\ldots,T\-1do

13:

𝐩\(t\+1\)←rϕ​\(𝐩\(t\),𝐮\(0\),𝒜\)\\mathbf\{p\}^\{\(t\+1\)\}\\leftarrow r\_\{\\phi\}\(\\mathbf\{p\}^\{\(t\)\},\\mathbf\{u\}^\{\(0\)\},\\mathcal\{A\}\)
14:endfor

15:

𝐩ref←Anchor⁡\(𝐩\(0\),𝐩\(T\),𝐮\(0\),𝒜\)\\mathbf\{p\}^\{\\mathrm\{ref\}\}\\leftarrow\\operatorname\{Anchor\}\(\\mathbf\{p\}^\{\(0\)\},\\mathbf\{p\}^\{\(T\)\},\\mathbf\{u\}^\{\(0\)\},\\mathcal\{A\}\)
16:Update

\(ϕ,ϑ\)\(\\phi,\\vartheta\)with Eq\. \([35](https://arxiv.org/html/2608.10004#A3.E35)\)

17:endfor

18:endfor

19:returnTrained parameters

θ,ψ,ϕ,ϑ\\theta,\\psi,\\phi,\\vartheta

## Appendix FComplete Ablation Results

This section evaluates the contribution of each relation type and the uncertainty gate across all three datasets\. Table[12](https://arxiv.org/html/2608.10004#A6.T12)extends the WBC table in the main paper to all three datasets while retaining both corruption types\. It reports results using the same three paired seeds and a corruption ratio of 50%\.

DatasetVariantCo\.Im\.Ex\.GateFlipMissingWBCFull ReCBM✓✓✓✓80\.10±\\pm3\.14 / 76\.00±\\pm2\.6290\.89±\\pm0\.21 / 95\.31±\\pm0\.49w/o Co\.×\\times✓✓✓74\.53±\\pm1\.70 / 59\.75±\\pm1\.4290\.61±\\pm0\.07 / 94\.87±\\pm0\.26w/o Im\.✓×\\times✓✓72\.89±\\pm0\.21 / 71\.81±\\pm0\.5088\.70±\\pm0\.15 / 94\.84±\\pm0\.69w/o Ex\.✓✓×\\times✓78\.09±\\pm5\.03 / 65\.87±\\pm7\.3489\.23±\\pm0\.36 / 94\.85±\\pm0\.53Co\. only✓×\\times×\\times✓64\.53±\\pm1\.46 / 55\.02±\\pm1\.1286\.42±\\pm0\.14 / 93\.97±\\pm0\.40Im\. only×\\times✓×\\times✓59\.24±\\pm0\.97 / 30\.54±\\pm1\.8185\.69±\\pm0\.05 / 93\.99±\\pm0\.47Ex\. only×\\times×\\times✓✓66\.02±\\pm0\.25 / 50\.31±\\pm1\.0581\.20±\\pm0\.07 / 94\.30±\\pm0\.41w/o Relations×\\times×\\times×\\times✓49\.94±\\pm0\.11 / 13\.92±\\pm0\.6364\.44±\\pm0\.14 / 91\.84±\\pm0\.86w/o Gating✓✓✓×\\times74\.83±\\pm0\.06 / 14\.71±\\pm0\.9985\.77±\\pm0\.12 / 90\.59±\\pm1\.85CUBFull ReCBM✓✓✓✓87\.65±\\pm0\.07 / 40\.09±\\pm0\.2791\.51±\\pm0\.14 /67\.31±\\pm0\.28w/o Co\.×\\times✓✓✓84\.45±\\pm0\.49 / 28\.74±\\pm2\.2890\.24±\\pm0\.07 / 67\.24±\\pm0\.53w/o Im\.✓×\\times✓✓85\.81±\\pm0\.25 / 38\.58±\\pm2\.4690\.00±\\pm0\.48 / 65\.69±\\pm0\.84w/o Ex\.✓✓×\\times✓86\.25±\\pm0\.24 / 37\.13±\\pm3\.2091\.54±\\pm0\.18/ 67\.13±\\pm1\.13Co\. only✓×\\times×\\times✓83\.00±\\pm0\.23 / 33\.70±\\pm0\.6489\.92±\\pm0\.06 / 66\.22±\\pm0\.68Im\. only×\\times✓×\\times✓78\.72±\\pm1\.39 / 17\.32±\\pm4\.4387\.84±\\pm0\.38 / 66\.35±\\pm0\.54Ex\. only×\\times×\\times✓✓79\.05±\\pm0\.96 / 10\.73±\\pm1\.1687\.78±\\pm0\.06 / 66\.80±\\pm1\.40w/o Relations×\\times×\\times×\\times✓50\.03±\\pm0\.02 / 0\.12±\\pm0\.0658\.02±\\pm0\.29 / 66\.24±\\pm1\.47w/o Gating✓✓✓×\\times74\.52±\\pm0\.07 / 0\.32±\\pm0\.0586\.18±\\pm0\.06 / 16\.57±\\pm0\.99SyntheticFull ReCBM✓✓✓✓84\.23±\\pm1\.06 / 77\.34±\\pm2\.0791\.27±\\pm0\.28 / 89\.97±\\pm0\.15w/o Co\.×\\times✓✓✓68\.85±\\pm8\.97 / 58\.00±\\pm9\.7485\.82±\\pm0\.13 / 89\.19±\\pm0\.52w/o Im\.✓×\\times✓✓82\.55±\\pm1\.35 / 76\.49±\\pm2\.4490\.69±\\pm0\.23 / 89\.76±\\pm0\.16w/o Ex\.✓✓×\\times✓79\.23±\\pm0\.56 / 65\.78±\\pm2\.4587\.75±\\pm0\.86 / 89\.53±\\pm0\.28Co\. only✓×\\times×\\times✓77\.21±\\pm0\.50 / 64\.80±\\pm2\.9487\.39±\\pm0\.82 / 89\.07±\\pm0\.73Im\. only×\\times✓×\\times✓55\.67±\\pm4\.63 / 35\.13±\\pm4\.3281\.29±\\pm0\.52 / 89\.09±\\pm0\.41Ex\. only×\\times×\\times✓✓61\.16±\\pm8\.61 / 49\.53±\\pm9\.3178\.42±\\pm0\.13 / 88\.93±\\pm0\.37w/o Relations×\\times×\\times×\\times✓49\.99±\\pm0\.35 / 24\.74±\\pm2\.3570\.72±\\pm1\.55 / 88\.97±\\pm0\.15w/o Gating✓✓✓×\\times57\.83±\\pm1\.37 / 24\.64±\\pm2\.9089\.28±\\pm0\.10 / 85\.16±\\pm0\.69Table 12:Complete component ablation under 50% concept corruption\. Entries are concept/task accuracy \(%\), reported as mean±\\pmstandard deviation over three paired seeds\. Bold denotes the best mean for each dataset and corruption type\.
## Appendix GQualitative Refinement Analysis

This section illustrates how relational refinement changed concept and task predictions in representative examples\. Figures[5\(a\)](https://arxiv.org/html/2608.10004#A7.F5.sf1)and[5\(b\)](https://arxiv.org/html/2608.10004#A7.F5.sf2)present representative refinement cases\. Each dataset is shown in a separate row to preserve readability\. Without concept corruption, refinement increased concept F1 from 0\.667 to 0\.833 on WBC, from 0\.787 to 0\.915 on CUB, and from 0\.889 to 1\.000 on Synthetic\. It corrected the WBC and CUB task predictions while leaving the already correct Synthetic prediction unchanged\.

For the uncertainty\-only cases, only the uncertainty assigned to unreliable concepts was changed, and no correct concept values were supplied\. Relational refinement increased concept F1 from 0 to 0\.941 on WBC, from 0\.038 to 0\.857 on CUB, and from 0 to 0\.500 on Synthetic, while correcting the task prediction in all three cases\.

![Refer to caption](https://arxiv.org/html/2608.10004v1/x5.png)

![Refer to caption](https://arxiv.org/html/2608.10004v1/x6.png)

![Refer to caption](https://arxiv.org/html/2608.10004v1/x7.png)

\(a\)Refinement of uncorrupted concept predictions\.
![Refer to caption](https://arxiv.org/html/2608.10004v1/x8.png)

![Refer to caption](https://arxiv.org/html/2608.10004v1/x9.png)

![Refer to caption](https://arxiv.org/html/2608.10004v1/x10.png)

\(b\)Refinement of corrupted concept predictions after marking corrupted entries as uncertain\.

Figure 5:Representative qualitative refinement examples on WBC, CUB, and Synthetic\. Left: concept predictions were refined without applying corruption\. Refinement improved concept F1 from 0\.667 to 0\.833 on WBC, from 0\.787 to 0\.915 on CUB, and from 0\.889 to 1\.000 on Synthetic\. It corrected the WBC and CUB task predictions while leaving the already correct Synthetic prediction unchanged\. Right: concept predictions were first corrupted, after which the corrupted entries were marked as unreliable by setting only their uncertainty to one, and their correct concept values were not provided\. Refinement improved concept F1 from 0 to 0\.941 on WBC, from 0\.038 to 0\.857 on CUB, and from 0 to 0\.500 on Synthetic, and corrected the task predictions in all three cases\.
## Appendix HUncertainty\-Only Intervention

This section examines whether changes in uncertainty alone could improve concept recovery while leaving the concept probabilities unchanged\. We first divided raw concept predictions into two candidate groups\. The first group contained correct predictions with uncertainty no greater than 0\.2, and the second contained incorrect predictions with uncertainty no less than 0\.8\. Within each group, 30% of the eligible predictions were sampled and flipped by reversing their thresholded binary values\. This produced incorrect predictions with low uncertainty and correct predictions with high uncertainty\.

For each intervention ratio, we sampled the corresponding proportion of the first group and set their uncertainty tou=1u=1\. We sampled the same proportion of the second group and set their uncertainty tou=0u=0\. The corrupted concept probabilities remained fixed throughout the intervention\. Any improvement therefore resulted from changing the reliability information provided to the refinement module rather than supplying corrected concept values\.

The uncertainty intervention figure in the main paper reports the results for WBC, CUB, and Synthetic\. It shows that concept and task performance improved as the intervention adjusted the uncertainty of a larger proportion of the sampled concepts, assigning high uncertainty to incorrect predictions\. Figure[5\(b\)](https://arxiv.org/html/2608.10004#A7.F5.sf2)provides individual examples of how these changes affected the refined concept predictions\.

## Appendix ICounterfactual Edits under Different Uncertainty Levels

This section tests whether uncertainty regulated the effect of identical counterfactual concept edits on task predictions\.

We considered test examples that were correctly classified by the complete model\. For each example, the alternative class was the class other than the original prediction that had the highest original probability\. For binary classification, this was the class opposite to the original prediction\.

Candidate concepts were scored using both their relational strength and their effect on the task logits\. For each concept, we first reversed its thresholded raw prediction and set its uncertainty to zero\. We then measured the sum of the decrease in the original class logit and the increase in the alternative class logit, with negative values clipped to zero\. The relational score was the strongest learned relation involving that concept, considering co\-occurrence, exclusion, and both directions of implication\. After separately normalizing the logit shift and relational scores to\[0,1\]\[0,1\], we multiplied them and selected theK∈\{1,3,5\}K\\in\\\{1,3,5\\\}concepts with the highest resulting scores\.

For each example and edit budget, the selected concept indices and their edited values were identical in the two uncertainty settings\. Each selected raw prediction was converted to a binary value using a threshold of 0\.5 and then reversed\. The uncertainty of every edited concept was set either tou=0u=0, indicating a reliable edit, or tou=1u=1, indicating an unreliable edit\. Target success was the fraction of examples whose final prediction changed to the selected alternative class\. Table[13](https://arxiv.org/html/2608.10004#A9.T13)reports the results\.

DatasetKKLowuuHighuuDifferenceWBC17\.021\.825\.20397\.0120\.9276\.09599\.9784\.7015\.27CUB118\.445\.2713\.17348\.5716\.0932\.48580\.1229\.7250\.40Synthetic119\.493\.6515\.84399\.0130\.8568\.16598\.8191\.067\.75Table 13:Target class success \(%\) of counterfactual concept edits under low and high uncertainty\.KKdenotes the number of edited concepts, and Difference denotes the low uncertainty result minus the high uncertainty result\.The desired behavior differed between the two settings\. Under low uncertainty, a high target success rate indicated that reliable concept edits effectively controlled the task prediction\. Under high uncertainty, a low target success rate indicated that the model resisted the same edits after they were marked as unreliable\. Target success generally increased withKK, while remaining lower under high uncertainty\. The gap between the two settings therefore measured how strongly uncertainty regulated the influence of identical concept edits\. This behavior also provided a degree of robustness to potentially incorrect human interventions: when an edit was marked as uncertain, the model reduced its influence and relied more on the remaining reliable concepts and learned relations\.

## Appendix JSufficient Concept\-Set Extraction

This section identifies compact concept sets that retained the predictive performance obtained from the complete concept representation\. A global set provides a compact concept interface for the overall task, whereas a set selected for an individual class identifies the concepts sufficient for predictions associated with that class\. If a small subset achieves performance comparable to that obtained using all concepts, users need to inspect or provide feedback for fewer concepts, reducing the burden of human interaction\. All set selection was performed on the validation split\.

During set selection, the selected dimensions retained their predicted raw probabilities and were assignedu=0u=0, while all unselected dimensions were replaced by the unknown state\(p,u\)=\(0\.5,1\)\(p,u\)=\(0\.5,1\)\. ReCBM then reconstructed the complete refined concept vector before applying the task predictor\. The same input protocol was used to evaluate the global set on the test split\. For the final evaluation of each class\-specific set, however, the selected dimensions were supplied with their ground truth binary values and assignedu=0u=0\. This setting more closely represented practical human intervention, in which users provide definite concept values rather than continuous prediction probabilities\. The unselected dimensions remained in the unknown state\.

To obtain the global set, we ranked concept dimensions by the product of their mean absolute task logit gradient and mean anchor gate\. The gradient was computed for the logit corresponding to the ground truth task label on the validation split\. We searched the ranked prefixes for the smallest prefix satisfying both retention criteria\. Starting from this prefix, we considered concepts in increasing order of their ranking scores and removed a concept whenever both criteria remained satisfied\.

For a candidate setSS, concept retention and task retention were defined as

Rconcept​\(S\)=F1concept⁡\(S\)F1concept⁡\(full\),Rtask​\(S\)=Acctask⁡\(S\)Acctask⁡\(full\)\.R\_\{\\mathrm\{concept\}\}\(S\)=\\frac\{\\operatorname\{F1\}\_\{\\mathrm\{concept\}\}\(S\)\}\{\\operatorname\{F1\}\_\{\\mathrm\{concept\}\}\(\\mathrm\{full\}\)\},\\qquad R\_\{\\mathrm\{task\}\}\(S\)=\\frac\{\\operatorname\{Acc\}\_\{\\mathrm\{task\}\}\(S\)\}\{\\operatorname\{Acc\}\_\{\\mathrm\{task\}\}\(\\mathrm\{full\}\)\}\.\(36\)A set satisfied a retention targetγ\\gammaonly whenRconcept​\(S\)≥γR\_\{\\mathrm\{concept\}\}\(S\)\\geq\\gammaandRtask​\(S\)≥γR\_\{\\mathrm\{task\}\}\(S\)\\geq\\gamma\.

For each individual class, selection was performed using only validation examples from that class\. Starting from an empty set, we evaluated every unselected concept as the next candidate and added the concept that produced the greatest minimum of concept F1 retention and task accuracy retention\. This process continued until both retention criteria reachedγ=1\\gamma=1\. We then removed selected concepts one at a time whenever both criteria remained satisfied\. The resulting concept sets were fixed before their final evaluation on the test split\.

Dataset90%95%98%100%WBC16/2416/2417/2418/24CUB64/11288/11294/11299/112Synthetic8/1211/1211/1212/12Table 14:Sizes of the global concept sets selected on the validation split at additional retention targets\. Selection required both concept F1 and task accuracy retention to meet the stated target\.Table[14](https://arxiv.org/html/2608.10004#A10.T14)shows how the number of globally selected concepts changed with the required performance retention level across the three datasets\. Tables[15](https://arxiv.org/html/2608.10004#A10.T15),[16](https://arxiv.org/html/2608.10004#A10.T16), andLABEL:tab:supp\_class\_sets\_cub112report the concept sets selected separately for each class\. Selection was performed on the validation split, and only sets that satisfied the retention criteria were included\. Valid sets were obtained for all five WBC classes, all four Synthetic classes, and 166 of the 200 CUB classes\. These sets were subsequently evaluated on the test split\. The results show that ReCBM identified compact concept sets that satisfied both retention criteria on the validation split and largely preserved performance on the test split\. These sets required users to inspect or intervene on only a subset of the available concepts\. The global sets provided a compact concept interface for the overall task, whereas the sets selected for individual classes provided more focused interfaces for inspection and intervention within each class\.

ClassSizeSelected conceptsBasophil6purple granule colour; big cell size; irregular nucleus shape; segmented multilobed nucleus shape; round cell shape; densely chromatin densityEosinophil5red granule colour; purple granule colour; segmented multilobed nucleus shape; round cell shape; small granule typeLymphocyte11high nuclear cytoplasmic ratio; purple blue cytoplasm colour; big cell size; round cell shape; irregular nucleus shape; clear cytoplasm texture; round granule type; red granule colour; light blue cytoplasm colour; granularity; densely chromatin densityMonocyte8unsegmented indented nucleus shape; high nuclear cytoplasmic ratio; cytoplasm vacuole; nil granule type; round cell shape; purple blue cytoplasm colour; unsegmented round nucleus shape; big cell sizeNeutrophil3unsegmented band nucleus shape; small granule type; clear cytoplasm textureTable 15:Class\-specific sufficient concept sets for WBC\.ClassSizeSelected conceptsNormal access11Login blocked by policy; Usual location; Password verified; Administrator account; New device; Account locked; MFA verified; Administrator operation; Login succeeds; Registered device; Anomalous locationCredential anomaly8Failed attempts reach threshold; Anomalous location; Administrator operation; Registered device; Login blocked by policy; Password verified; New device; Account lockedDevice or location anomaly12Anomalous location; Login blocked by policy; Usual location; Administrator account; Account locked; Password verified; Login succeeds; MFA verified; Administrator operation; Registered device; Failed attempts reach threshold; New devicePolicy block or privilege risk11Administrator operation; Password verified; Registered device; Anomalous location; Account locked; Login blocked by policy; Login succeeds; Administrator account; Usual location; New device; MFA verifiedTable 16:Class\-specific sufficient concept sets for Synthetic\.Table 17:Class\-specific sufficient concept sets for CUB\.ClassSizeSelected conceptsBlack footed Albatross10has\_bill\_shape: all\-purpose; has\_wing\_color: buff; has\_wing\_color: white; has\_upperparts\_color: black; has\_underparts\_color: white; has\_back\_color: white; has\_breast\_pattern: solid; has\_upper\_tail\_color: grey; has\_bill\_shape: hooked\_seabird; has\_head\_pattern: plainLaysan Albatross11has\_bill\_shape: hooked\_seabird; has\_nape\_color: grey; has\_wing\_color: grey; has\_wing\_color: black; has\_wing\_color: brown; has\_wing\_color: yellow; has\_upperparts\_color: brown; has\_upperparts\_color: yellow; has\_underparts\_color: brown; has\_underparts\_color: grey; has\_size: small\_\(5\_\-\_9\_in\)Sooty Albatross8has\_belly\_pattern: solid; has\_throat\_color: grey; has\_bill\_length: shorter\_than\_head; has\_upperparts\_color: buff; has\_upper\_tail\_color: brown; has\_under\_tail\_color: brown; has\_upper\_tail\_color: black; has\_bill\_shape: daggerCrested Auklet3has\_throat\_color: black; has\_bill\_length: about\_the\_same\_as\_head; has\_underparts\_color: whiteLeast Auklet4has\_wing\_color: black; has\_eye\_color: black; has\_breast\_pattern: striped; has\_wing\_color: whiteParakeet Auklet10has\_forehead\_color: black; has\_leg\_color: black; has\_eye\_color: black; has\_forehead\_color: blue; has\_underparts\_color: buff; has\_back\_color: grey; has\_primary\_color: grey; has\_underparts\_color: brown; has\_throat\_color: white; has\_wing\_pattern: solidRhinoceros Auklet7has\_underparts\_color: white; has\_throat\_color: white; has\_belly\_color: white; has\_size: medium\_\(9\_\-\_16\_in\); has\_wing\_pattern: spotted; has\_wing\_color: grey; has\_bill\_length: shorter\_than\_headRed winged Blackbird9has\_wing\_color: black; has\_bill\_color: buff; has\_belly\_color: white; has\_breast\_color: yellow; has\_wing\_pattern: multi\-colored; has\_leg\_color: black; has\_eye\_color: black; has\_breast\_pattern: solid; has\_tail\_shape: notched\_tailRusty Blackbird14has\_eye\_color: black; has\_upperparts\_color: white; has\_breast\_pattern: multi\-colored; has\_wing\_pattern: spotted; has\_back\_pattern: multi\-colored; has\_belly\_pattern: solid; has\_bill\_shape: hooked\_seabird; has\_wing\_color: buff; has\_upperparts\_color: grey; has\_breast\_color: grey; has\_underparts\_color: white; has\_wing\_shape: rounded\-wings; has\_size: small\_\(5\_\-\_9\_in\); has\_upperparts\_color: buffBobolink10has\_leg\_color: grey; has\_crown\_color: white; has\_throat\_color: black; has\_size: small\_\(5\_\-\_9\_in\); has\_back\_pattern: solid; has\_belly\_color: yellow; has\_primary\_color: black; has\_bill\_shape: dagger; has\_back\_pattern: multi\-colored; has\_wing\_color: whiteIndigo Bunting5has\_forehead\_color: blue; has\_head\_pattern: plain; has\_bill\_color: black; has\_wing\_pattern: multi\-colored; has\_crown\_color: blueLazuli Bunting12has\_forehead\_color: blue; has\_bill\_shape: all\-purpose; has\_crown\_color: blue; has\_breast\_pattern: striped; has\_breast\_color: grey; has\_throat\_color: grey; has\_underparts\_color: grey; has\_belly\_color: grey; has\_primary\_color: brown; has\_wing\_shape: rounded\-wings; has\_wing\_pattern: multi\-colored; has\_breast\_color: buffPainted Bunting7has\_forehead\_color: blue; has\_nape\_color: grey; has\_breast\_color: white; has\_forehead\_color: brown; has\_forehead\_color: black; has\_bill\_shape: all\-purpose; has\_shape: perching\-likeCardinal10has\_bill\_color: black; has\_upper\_tail\_color: black; has\_wing\_pattern: multi\-colored; has\_forehead\_color: black; has\_breast\_color: black; has\_throat\_color: white; has\_forehead\_color: blue; has\_back\_pattern: multi\-colored; has\_underparts\_color: yellow; has\_wing\_color: greySpotted Catbird10has\_throat\_color: buff; has\_bill\_color: black; has\_bill\_shape: dagger; has\_wing\_color: black; has\_leg\_color: grey; has\_breast\_pattern: solid; has\_size: small\_\(5\_\-\_9\_in\); has\_belly\_color: buff; has\_under\_tail\_color: grey; has\_wing\_color: yellowGray Catbird5has\_nape\_color: grey; has\_underparts\_color: yellow; has\_bill\_shape: hooked\_seabird; has\_upperparts\_color: brown; has\_underparts\_color: whiteYellow breasted Chat8has\_wing\_color: black; has\_leg\_color: grey; has\_forehead\_color: blue; has\_upper\_tail\_color: grey; has\_tail\_pattern: solid; has\_primary\_color: grey; has\_throat\_color: yellow; has\_wing\_color: yellowEastern Towhee7has\_breast\_pattern: multi\-colored; has\_bill\_shape: all\-purpose; has\_underparts\_color: grey; has\_primary\_color: grey; has\_back\_color: grey; has\_throat\_color: black; has\_size: small\_\(5\_\-\_9\_in\)Chuck will Widow6has\_underparts\_color: brown; has\_primary\_color: buff; has\_upperparts\_color: grey; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_wing\_shape: rounded\-wings; has\_wing\_shape: pointed\-wingsBrandt Cormorant2has\_throat\_color: black; has\_bill\_shape: all\-purposeRed faced Cormorant15has\_belly\_color: black; has\_primary\_color: brown; has\_upper\_tail\_color: white; has\_bill\_shape: all\-purpose; has\_wing\_shape: rounded\-wings; has\_bill\_shape: cone; has\_wing\_color: brown; has\_breast\_pattern: solid; has\_wing\_color: grey; has\_head\_pattern: plain; has\_underparts\_color: black; has\_wing\_color: buff; has\_upperparts\_color: brown; has\_upperparts\_color: grey; has\_throat\_color: whiteFish Crow14has\_tail\_pattern: solid; has\_wing\_color: grey; has\_leg\_color: black; has\_bill\_shape: cone; has\_underparts\_color: black; has\_wing\_shape: rounded\-wings; has\_upperparts\_color: grey; has\_breast\_pattern: solid; has\_breast\_pattern: multi\-colored; has\_back\_color: brown; has\_back\_color: grey; has\_back\_color: black; has\_wing\_color: white; has\_throat\_color: blackBlack billed Cuckoo8has\_eye\_color: black; has\_belly\_color: brown; has\_primary\_color: buff; has\_upperparts\_color: black; has\_head\_pattern: eyebrow; has\_breast\_color: grey; has\_primary\_color: brown; has\_back\_pattern: solidMangrove Cuckoo6has\_throat\_color: yellow; has\_leg\_color: buff; has\_breast\_color: grey; has\_forehead\_color: grey; has\_underparts\_color: buff; has\_bill\_shape: coneYellow billed Cuckoo5has\_eye\_color: black; has\_back\_color: black; has\_leg\_color: grey; has\_belly\_color: white; has\_crown\_color: whitePurple Finch7has\_bill\_color: buff; has\_upperparts\_color: grey; has\_upperparts\_color: black; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_tail\_shape: notched\_tail; has\_underparts\_color: white; has\_nape\_color: brownNorthern Flicker11has\_throat\_color: black; has\_wing\_color: black; has\_wing\_pattern: multi\-colored; has\_underparts\_color: white; has\_forehead\_color: yellow; has\_underparts\_color: brown; has\_wing\_pattern: spotted; has\_leg\_color: black; has\_leg\_color: grey; has\_wing\_color: buff; has\_wing\_shape: pointed\-wingsGreat Crested Flycatcher8has\_forehead\_color: blue; has\_underparts\_color: white; has\_wing\_pattern: spotted; has\_belly\_color: white; has\_wing\_color: grey; has\_breast\_color: grey; has\_back\_color: black; has\_throat\_color: greyLeast Flycatcher15has\_eye\_color: black; has\_belly\_color: white; has\_primary\_color: black; has\_leg\_color: grey; has\_under\_tail\_color: grey; has\_underparts\_color: grey; has\_bill\_color: black; has\_nape\_color: buff; has\_wing\_pattern: striped; has\_belly\_pattern: solid; has\_bill\_shape: hooked\_seabird; has\_upper\_tail\_color: white; has\_bill\_length: about\_the\_same\_as\_head; has\_bill\_length: shorter\_than\_head; has\_breast\_color: yellowOlive sided Flycatcher5has\_eye\_color: black; has\_wing\_color: white; has\_under\_tail\_color: buff; has\_breast\_pattern: multi\-colored; has\_crown\_color: yellowScissor tailed Flycatcher28has\_bill\_shape: hooked\_seabird; has\_bill\_color: buff; has\_nape\_color: buff; has\_breast\_color: brown; has\_throat\_color: buff; has\_under\_tail\_color: buff; has\_belly\_color: brown; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_tail\_pattern: striped; has\_underparts\_color: grey; has\_belly\_color: grey; has\_forehead\_color: blue; has\_crown\_color: blue; has\_back\_pattern: multi\-colored; has\_bill\_color: grey; has\_back\_pattern: striped; has\_bill\_shape: cone; has\_wing\_color: brown; has\_upperparts\_color: brown; has\_underparts\_color: buff; has\_back\_color: buff; has\_breast\_color: buff; has\_throat\_color: yellow; has\_belly\_color: buff; has\_tail\_shape: notched\_tail; has\_size: small\_\(5\_\-\_9\_in\); has\_under\_tail\_color: grey; has\_back\_color: blackVermilion Flycatcher12has\_breast\_pattern: solid; has\_size: small\_\(5\_\-\_9\_in\); has\_bill\_shape: hooked\_seabird; has\_primary\_color: white; has\_wing\_color: yellow; has\_wing\_color: grey; has\_forehead\_color: black; has\_wing\_pattern: solid; has\_tail\_pattern: solid; has\_tail\_shape: notched\_tail; has\_leg\_color: black; has\_wing\_color: whiteYellow bellied Flycatcher14has\_forehead\_color: black; has\_belly\_pattern: solid; has\_upper\_tail\_color: white; has\_bill\_length: about\_the\_same\_as\_head; has\_forehead\_color: brown; has\_throat\_color: buff; has\_nape\_color: brown; has\_nape\_color: white; has\_primary\_color: white; has\_bill\_color: buff; has\_primary\_color: grey; has\_breast\_color: black; has\_back\_color: yellow; has\_leg\_color: greyFrigatebird7has\_leg\_color: black; has\_under\_tail\_color: grey; has\_breast\_color: white; has\_belly\_pattern: solid; has\_wing\_pattern: multi\-colored; has\_bill\_color: black; has\_wing\_color: brownNorthern Fulmar1has\_bill\_shape: hooked\_seabirdGadwall2has\_shape: duck\-like; has\_wing\_color: blackAmerican Goldfinch14has\_back\_color: yellow; has\_upperparts\_color: black; has\_bill\_shape: cone; has\_breast\_color: black; has\_wing\_color: yellow; has\_bill\_color: grey; has\_belly\_pattern: solid; has\_leg\_color: buff; has\_underparts\_color: white; has\_back\_color: brown; has\_back\_color: white; has\_tail\_shape: notched\_tail; has\_under\_tail\_color: white; has\_throat\_color: yellowEuropean Goldfinch11has\_forehead\_color: grey; has\_tail\_pattern: multi\-colored; has\_breast\_pattern: solid; has\_bill\_shape: all\-purpose; has\_bill\_shape: cone; has\_wing\_color: yellow; has\_belly\_color: white; has\_upperparts\_color: yellow; has\_underparts\_color: buff; has\_crown\_color: black; has\_belly\_pattern: solidBoat tailed Grackle14has\_tail\_pattern: solid; has\_upperparts\_color: white; has\_wing\_pattern: spotted; has\_forehead\_color: yellow; has\_wing\_color: buff; has\_upperparts\_color: grey; has\_back\_color: grey; has\_breast\_color: grey; has\_belly\_color: brown; has\_belly\_color: white; has\_back\_pattern: multi\-colored; has\_bill\_color: buff; has\_breast\_pattern: solid; has\_leg\_color: blackEared Grebe2has\_size: small\_\(5\_\-\_9\_in\); has\_belly\_color: greyHorned Grebe4has\_wing\_color: white; has\_eye\_color: black; has\_nape\_color: grey; has\_shape: duck\-likePied billed Grebe4has\_underparts\_color: grey; has\_nape\_color: grey; has\_forehead\_color: grey; has\_belly\_color: yellowWestern Grebe2has\_eye\_color: black; has\_under\_tail\_color: whiteBlue Grosbeak7has\_back\_pattern: multi\-colored; has\_crown\_color: blue; has\_upperparts\_color: black; has\_nape\_color: yellow; has\_wing\_pattern: solid; has\_underparts\_color: white; has\_nape\_color: blackEvening Grosbeak6has\_upperparts\_color: yellow; has\_nape\_color: brown; has\_belly\_color: brown; has\_wing\_pattern: multi\-colored; has\_throat\_color: white; has\_wing\_color: greyPine Grosbeak14has\_breast\_color: grey; has\_belly\_color: brown; has\_size: medium\_\(9\_\-\_16\_in\); has\_throat\_color: grey; has\_underparts\_color: white; has\_underparts\_color: brown; has\_bill\_color: buff; has\_back\_color: grey; has\_belly\_pattern: solid; has\_primary\_color: grey; has\_upperparts\_color: white; has\_bill\_length: about\_the\_same\_as\_head; has\_belly\_color: grey; has\_wing\_color: whiteRose breasted Grosbeak14has\_breast\_pattern: multi\-colored; has\_belly\_color: yellow; has\_primary\_color: grey; has\_back\_color: grey; has\_forehead\_color: blue; has\_nape\_color: brown; has\_underparts\_color: brown; has\_shape: duck\-like; has\_leg\_color: grey; has\_head\_pattern: eyebrow; has\_wing\_shape: pointed\-wings; has\_throat\_color: buff; has\_wing\_pattern: multi\-colored; has\_wing\_pattern: spottedPigeon Guillemot8has\_under\_tail\_color: black; has\_bill\_shape: all\-purpose; has\_back\_pattern: solid; has\_wing\_pattern: multi\-colored; has\_bill\_shape: hooked\_seabird; has\_tail\_pattern: multi\-colored; has\_primary\_color: grey; has\_crown\_color: greyCalifornia Gull1has\_forehead\_color: whiteGlaucous winged Gull9has\_crown\_color: white; has\_throat\_color: grey; has\_wing\_color: black; has\_bill\_shape: dagger; has\_crown\_color: black; has\_upperparts\_color: black; has\_back\_color: brown; has\_back\_color: yellow; has\_upperparts\_color: greyHeermann Gull3has\_primary\_color: grey; has\_wing\_color: brown; has\_bill\_shape: daggerIvory Gull2has\_forehead\_color: white; has\_back\_color: greySlaty backed Gull3has\_crown\_color: white; has\_breast\_pattern: solid; has\_wing\_pattern: solidWestern Gull1has\_crown\_color: whiteAnna Hummingbird9has\_upperparts\_color: black; has\_breast\_pattern: solid; has\_wing\_shape: pointed\-wings; has\_leg\_color: black; has\_bill\_shape: dagger; has\_upper\_tail\_color: white; has\_wing\_color: grey; has\_tail\_pattern: multi\-colored; has\_size: small\_\(5\_\-\_9\_in\)Ruby throated Hummingbird19has\_underparts\_color: yellow; has\_wing\_color: brown; has\_wing\_color: yellow; has\_bill\_shape: cone; has\_upperparts\_color: brown; has\_upperparts\_color: white; has\_upperparts\_color: buff; has\_underparts\_color: brown; has\_breast\_pattern: striped; has\_back\_color: brown; has\_back\_color: buff; has\_upper\_tail\_color: brown; has\_breast\_color: yellow; has\_breast\_color: white; has\_throat\_color: grey; has\_throat\_color: yellow; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_wing\_color: black; has\_leg\_color: blackRufous Hummingbird4has\_bill\_shape: hooked\_seabird; has\_belly\_pattern: solid; has\_upper\_tail\_color: black; has\_back\_pattern: solidGreen Violetear8has\_wing\_color: white; has\_wing\_color: grey; has\_crown\_color: yellow; has\_wing\_shape: rounded\-wings; has\_bill\_length: shorter\_than\_head; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_upperparts\_color: brown; has\_nape\_color: yellowBlue Jay14has\_forehead\_color: blue; has\_upper\_tail\_color: buff; has\_breast\_color: black; has\_crown\_color: blue; has\_back\_pattern: striped; has\_bill\_shape: cone; has\_nape\_color: buff; has\_primary\_color: black; has\_leg\_color: buff; has\_wing\_pattern: solid; has\_tail\_pattern: striped; has\_back\_color: grey; has\_eye\_color: black; has\_bill\_color: greyFlorida Jay7has\_tail\_shape: notched\_tail; has\_crown\_color: grey; has\_crown\_color: blue; has\_wing\_color: white; has\_leg\_color: black; has\_leg\_color: grey; has\_size: small\_\(5\_\-\_9\_in\)Green Jay10has\_wing\_color: white; has\_head\_pattern: eyebrow; has\_crown\_color: blue; has\_breast\_pattern: multi\-colored; has\_under\_tail\_color: buff; has\_primary\_color: grey; has\_wing\_shape: rounded\-wings; has\_belly\_pattern: solid; has\_wing\_pattern: multi\-colored; has\_primary\_color: yellowDark eyed Junco12has\_throat\_color: grey; has\_breast\_color: white; has\_belly\_color: grey; has\_primary\_color: grey; has\_belly\_pattern: solid; has\_bill\_length: about\_the\_same\_as\_head; has\_belly\_color: white; has\_upper\_tail\_color: grey; has\_wing\_color: black; has\_wing\_shape: rounded\-wings; has\_size: small\_\(5\_\-\_9\_in\); has\_primary\_color: yellowTropical Kingbird12has\_crown\_color: grey; has\_underparts\_color: grey; has\_throat\_color: white; has\_breast\_pattern: striped; has\_under\_tail\_color: buff; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_primary\_color: buff; has\_crown\_color: brown; has\_upperparts\_color: black; has\_wing\_pattern: multi\-colored; has\_bill\_shape: all\-purpose; has\_nape\_color: greyGray Kingbird12has\_nape\_color: grey; has\_wing\_color: yellow; has\_throat\_color: buff; has\_forehead\_color: blue; has\_wing\_color: buff; has\_nape\_color: brown; has\_nape\_color: buff; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_forehead\_color: grey; has\_bill\_shape: hooked\_seabird; has\_back\_color: black; has\_breast\_color: whiteGreen Kingfisher9has\_upperparts\_color: white; has\_bill\_shape: dagger; has\_wing\_shape: rounded\-wings; has\_breast\_color: white; has\_belly\_color: white; has\_nape\_color: grey; has\_belly\_pattern: solid; has\_wing\_shape: pointed\-wings; has\_forehead\_color: blackPied Kingfisher3has\_bill\_shape: dagger; has\_breast\_pattern: multi\-colored; has\_wing\_pattern: spottedWhite breasted Kingfisher11has\_bill\_shape: hooked\_seabird; has\_under\_tail\_color: buff; has\_upperparts\_color: black; has\_wing\_pattern: multi\-colored; has\_wing\_color: yellow; has\_wing\_color: black; has\_head\_pattern: plain; has\_wing\_color: brown; has\_underparts\_color: brown; has\_forehead\_color: brown; has\_bill\_shape: daggerRed legged Kittiwake6has\_forehead\_color: white; has\_forehead\_color: grey; has\_bill\_shape: hooked\_seabird; has\_bill\_shape: all\-purpose; has\_wing\_color: yellow; has\_wing\_pattern: stripedHorned Lark15has\_upperparts\_color: buff; has\_back\_pattern: striped; has\_wing\_shape: rounded\-wings; has\_throat\_color: yellow; has\_nape\_color: buff; has\_underparts\_color: brown; has\_bill\_shape: all\-purpose; has\_underparts\_color: white; has\_primary\_color: white; has\_back\_color: yellow; has\_primary\_color: yellow; has\_wing\_color: grey; has\_underparts\_color: black; has\_bill\_color: black; has\_breast\_pattern: multi\-coloredPacific Loon10has\_wing\_pattern: spotted; has\_belly\_color: grey; has\_throat\_color: grey; has\_underparts\_color: buff; has\_belly\_color: brown; has\_bill\_length: shorter\_than\_head; has\_back\_pattern: multi\-colored; has\_bill\_color: grey; has\_shape: perching\-like; has\_nape\_color: greyMallard6has\_back\_pattern: multi\-colored; has\_wing\_color: grey; has\_shape: duck\-like; has\_wing\_color: black; has\_head\_pattern: plain; has\_wing\_color: brownHooded Merganser12has\_primary\_color: black; has\_eye\_color: black; has\_leg\_color: black; has\_belly\_color: black; has\_wing\_color: brown; has\_underparts\_color: black; has\_back\_color: grey; has\_wing\_shape: pointed\-wings; has\_upperparts\_color: black; has\_breast\_color: black; has\_bill\_length: about\_the\_same\_as\_head; has\_wing\_pattern: solidRed breasted Merganser4has\_size: medium\_\(9\_\-\_16\_in\); has\_size: small\_\(5\_\-\_9\_in\); has\_back\_pattern: solid; has\_crown\_color: whiteMockingbird15has\_eye\_color: black; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_underparts\_color: yellow; has\_bill\_shape: all\-purpose; has\_breast\_color: yellow; has\_underparts\_color: brown; has\_forehead\_color: blue; has\_belly\_color: brown; has\_shape: duck\-like; has\_breast\_pattern: multi\-colored; has\_wing\_color: white; has\_wing\_shape: pointed\-wings; has\_upperparts\_color: grey; has\_forehead\_color: grey; has\_size: medium\_\(9\_\-\_16\_in\)Nighthawk12has\_wing\_color: grey; has\_upperparts\_color: black; has\_upperparts\_color: white; has\_underparts\_color: white; has\_bill\_shape: cone; has\_tail\_shape: notched\_tail; has\_wing\_pattern: spotted; has\_wing\_pattern: multi\-colored; has\_back\_pattern: solid; has\_under\_tail\_color: brown; has\_wing\_color: buff; has\_throat\_color: whiteClark Nutcracker20has\_nape\_color: grey; has\_under\_tail\_color: black; has\_breast\_color: black; has\_underparts\_color: black; has\_bill\_color: grey; has\_underparts\_color: grey; has\_belly\_color: yellow; has\_tail\_shape: notched\_tail; has\_underparts\_color: white; has\_under\_tail\_color: white; has\_wing\_shape: pointed\-wings; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_under\_tail\_color: buff; has\_leg\_color: buff; has\_bill\_color: buff; has\_nape\_color: buff; has\_leg\_color: grey; has\_under\_tail\_color: brown; has\_primary\_color: buff; has\_wing\_shape: rounded\-wingsWhite breasted Nuthatch11has\_throat\_color: grey; has\_wing\_pattern: spotted; has\_crown\_color: yellow; has\_bill\_shape: cone; has\_upperparts\_color: black; has\_shape: duck\-like; has\_head\_pattern: plain; has\_wing\_shape: rounded\-wings; has\_bill\_color: grey; has\_nape\_color: white; has\_wing\_pattern: solidBaltimore Oriole13has\_leg\_color: grey; has\_crown\_color: grey; has\_breast\_color: white; has\_leg\_color: black; has\_wing\_color: brown; has\_underparts\_color: white; has\_back\_color: white; has\_wing\_pattern: striped; has\_upperparts\_color: buff; has\_wing\_color: black; has\_upperparts\_color: white; has\_crown\_color: yellow; has\_primary\_color: yellowOrchard Oriole19has\_wing\_color: black; has\_wing\_shape: pointed\-wings; has\_eye\_color: black; has\_upperparts\_color: brown; has\_underparts\_color: buff; has\_back\_color: brown; has\_back\_color: white; has\_throat\_color: grey; has\_forehead\_color: brown; has\_nape\_color: buff; has\_belly\_color: grey; has\_belly\_color: white; has\_throat\_color: black; has\_breast\_color: black; has\_shape: perching\-like; has\_leg\_color: black; has\_back\_color: black; has\_wing\_pattern: multi\-colored; has\_wing\_color: greyScott Oriole3has\_breast\_pattern: multi\-colored; has\_throat\_color: black; has\_primary\_color: brownOvenbird10has\_bill\_shape: hooked\_seabird; has\_eye\_color: black; has\_leg\_color: grey; has\_back\_pattern: multi\-colored; has\_tail\_pattern: striped; has\_upperparts\_color: grey; has\_upperparts\_color: black; has\_breast\_color: buff; has\_head\_pattern: eyebrow; has\_size: small\_\(5\_\-\_9\_in\)White Pelican5has\_forehead\_color: white; has\_back\_color: grey; has\_bill\_length: about\_the\_same\_as\_head; has\_back\_color: white; has\_wing\_color: blackSayornis16has\_breast\_pattern: multi\-colored; has\_crown\_color: yellow; has\_back\_pattern: multi\-colored; has\_bill\_color: buff; has\_belly\_color: white; has\_crown\_color: blue; has\_underparts\_color: brown; has\_under\_tail\_color: white; has\_back\_color: white; has\_forehead\_color: blue; has\_nape\_color: yellow; has\_crown\_color: grey; has\_nape\_color: white; has\_size: small\_\(5\_\-\_9\_in\); has\_leg\_color: black; has\_forehead\_color: blackWhip poor Will6has\_belly\_color: brown; has\_throat\_color: buff; has\_bill\_color: black; has\_shape: duck\-like; has\_underparts\_color: brown; has\_leg\_color: buffHorned Puffin16has\_forehead\_color: black; has\_leg\_color: black; has\_breast\_pattern: multi\-colored; has\_bill\_shape: cone; has\_wing\_pattern: spotted; has\_nape\_color: brown; has\_upperparts\_color: brown; has\_wing\_color: buff; has\_underparts\_color: brown; has\_breast\_color: grey; has\_back\_pattern: solid; has\_wing\_shape: rounded\-wings; has\_bill\_length: about\_the\_same\_as\_head; has\_bill\_shape: hooked\_seabird; has\_wing\_pattern: solid; has\_underparts\_color: whiteWhite necked Raven5has\_underparts\_color: black; has\_upperparts\_color: white; has\_wing\_color: brown; has\_belly\_color: black; has\_nape\_color: whiteAmerican Redstart10has\_breast\_pattern: multi\-colored; has\_back\_color: white; has\_wing\_color: black; has\_belly\_pattern: solid; has\_throat\_color: grey; has\_underparts\_color: brown; has\_belly\_color: brown; has\_bill\_shape: hooked\_seabird; has\_upperparts\_color: white; has\_size: small\_\(5\_\-\_9\_in\)Geococcyx10has\_wing\_pattern: spotted; has\_underparts\_color: white; has\_breast\_pattern: solid; has\_shape: perching\-like; has\_size: small\_\(5\_\-\_9\_in\); has\_underparts\_color: buff; has\_bill\_color: grey; has\_forehead\_color: black; has\_belly\_pattern: solid; has\_bill\_length: about\_the\_same\_as\_headLoggerhead Shrike7has\_breast\_color: grey; has\_forehead\_color: grey; has\_bill\_shape: all\-purpose; has\_back\_color: grey; has\_underparts\_color: grey; has\_primary\_color: grey; has\_bill\_color: blackGreat Grey Shrike24has\_crown\_color: yellow; has\_bill\_shape: hooked\_seabird; has\_back\_color: grey; has\_belly\_color: grey; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_forehead\_color: blue; has\_belly\_color: buff; has\_eye\_color: black; has\_nape\_color: yellow; has\_wing\_shape: rounded\-wings; has\_wing\_color: yellow; has\_belly\_pattern: solid; has\_head\_pattern: plain; has\_wing\_pattern: multi\-colored; has\_leg\_color: black; has\_wing\_color: brown; has\_wing\_color: buff; has\_primary\_color: black; has\_breast\_pattern: striped; has\_upperparts\_color: brown; has\_head\_pattern: eyebrow; has\_breast\_color: buff; has\_upperparts\_color: yellow; has\_breast\_color: whiteBaird Sparrow6has\_bill\_color: buff; has\_size: small\_\(5\_\-\_9\_in\); has\_bill\_shape: all\-purpose; has\_bill\_length: shorter\_than\_head; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_nape\_color: yellowBlack throated Sparrow3has\_underparts\_color: grey; has\_forehead\_color: grey; has\_upperparts\_color: yellowBrewer Sparrow3has\_belly\_color: buff; has\_breast\_pattern: solid; has\_back\_color: buffChipping Sparrow9has\_upper\_tail\_color: brown; has\_back\_pattern: striped; has\_forehead\_color: brown; has\_breast\_pattern: solid; has\_bill\_color: buff; has\_underparts\_color: brown; has\_leg\_color: buff; has\_tail\_pattern: solid; has\_primary\_color: blackClay colored Sparrow6has\_bill\_color: buff; has\_upper\_tail\_color: white; has\_bill\_length: about\_the\_same\_as\_head; has\_wing\_color: grey; has\_belly\_color: brown; has\_back\_pattern: stripedHouse Sparrow10has\_leg\_color: buff; has\_belly\_color: buff; has\_crown\_color: grey; has\_breast\_color: grey; has\_tail\_shape: notched\_tail; has\_back\_color: grey; has\_under\_tail\_color: grey; has\_primary\_color: grey; has\_underparts\_color: buff; has\_wing\_color: brownField Sparrow3has\_under\_tail\_color: buff; has\_breast\_pattern: solid; has\_underparts\_color: yellowFox Sparrow4has\_upper\_tail\_color: brown; has\_underparts\_color: buff; has\_underparts\_color: brown; has\_back\_color: buffGrasshopper Sparrow2has\_bill\_color: buff; has\_upper\_tail\_color: buffHarris Sparrow13has\_breast\_pattern: striped; has\_back\_pattern: striped; has\_throat\_color: black; has\_underparts\_color: white; has\_upper\_tail\_color: black; has\_wing\_shape: rounded\-wings; has\_leg\_color: black; has\_bill\_shape: hooked\_seabird; has\_wing\_color: yellow; has\_upperparts\_color: yellow; has\_bill\_length: shorter\_than\_head; has\_back\_color: grey; has\_bill\_shape: all\-purposeHenslow Sparrow6has\_leg\_color: grey; has\_under\_tail\_color: grey; has\_shape: duck\-like; has\_breast\_pattern: solid; has\_nape\_color: black; has\_upper\_tail\_color: buffLe Conte Sparrow4has\_back\_color: black; has\_nape\_color: brown; has\_bill\_color: black; has\_upperparts\_color: buffLincoln Sparrow3has\_upper\_tail\_color: buff; has\_bill\_color: buff; has\_wing\_color: whiteNelson Sharp tailed Sparrow13has\_breast\_color: buff; has\_belly\_color: buff; has\_shape: duck\-like; has\_wing\_shape: pointed\-wings; has\_tail\_pattern: solid; has\_primary\_color: grey; has\_wing\_pattern: solid; has\_belly\_color: brown; has\_leg\_color: buff; has\_back\_color: grey; has\_breast\_pattern: striped; has\_belly\_color: white; has\_under\_tail\_color: brownSavannah Sparrow4has\_back\_pattern: striped; has\_upper\_tail\_color: buff; has\_breast\_pattern: solid; has\_bill\_shape: coneSeaside Sparrow11has\_eye\_color: black; has\_wing\_color: grey; has\_upperparts\_color: grey; has\_underparts\_color: white; has\_breast\_pattern: multi\-colored; has\_upper\_tail\_color: grey; has\_breast\_color: black; has\_forehead\_color: white; has\_under\_tail\_color: white; has\_bill\_color: grey; has\_nape\_color: greySong Sparrow4has\_primary\_color: brown; has\_wing\_shape: pointed\-wings; has\_primary\_color: buff; has\_breast\_color: buffTree Sparrow12has\_back\_pattern: striped; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_breast\_color: black; has\_forehead\_color: blue; has\_belly\_color: grey; has\_wing\_color: black; has\_throat\_color: yellow; has\_underparts\_color: grey; has\_forehead\_color: grey; has\_forehead\_color: black; has\_breast\_pattern: solid; has\_leg\_color: buffVesper Sparrow2has\_bill\_color: buff; has\_underparts\_color: buffWhite crowned Sparrow3has\_throat\_color: grey; has\_under\_tail\_color: grey; has\_breast\_color: buffWhite throated Sparrow3has\_throat\_color: grey; has\_under\_tail\_color: brown; has\_forehead\_color: blackCape Glossy Starling6has\_tail\_pattern: solid; has\_underparts\_color: black; has\_forehead\_color: blue; has\_head\_pattern: plain; has\_bill\_shape: cone; has\_back\_pattern: multi\-coloredBank Swallow14has\_size: very\_small\_\(3\_\-\_5\_in\); has\_throat\_color: grey; has\_belly\_color: brown; has\_wing\_pattern: striped; has\_wing\_shape: rounded\-wings; has\_tail\_shape: notched\_tail; has\_primary\_color: yellow; has\_shape: duck\-like; has\_nape\_color: black; has\_breast\_color: black; has\_wing\_pattern: spotted; has\_primary\_color: black; has\_under\_tail\_color: brown; has\_back\_color: brownBarn Swallow17has\_bill\_color: black; has\_wing\_pattern: striped; has\_forehead\_color: white; has\_crown\_color: white; has\_breast\_color: white; has\_back\_pattern: striped; has\_wing\_color: buff; has\_crown\_color: blue; has\_under\_tail\_color: brown; has\_under\_tail\_color: buff; has\_wing\_color: brown; has\_tail\_pattern: multi\-colored; has\_primary\_color: buff; has\_crown\_color: yellow; has\_wing\_color: grey; has\_belly\_color: buff; has\_belly\_pattern: solidCliff Swallow9has\_eye\_color: black; has\_belly\_color: yellow; has\_crown\_color: blue; has\_throat\_color: yellow; has\_primary\_color: yellow; has\_nape\_color: grey; has\_breast\_color: brown; has\_bill\_length: shorter\_than\_head; has\_underparts\_color: whiteScarlet Tanager18has\_bill\_color: buff; has\_wing\_pattern: spotted; has\_underparts\_color: brown; has\_belly\_color: white; has\_shape: duck\-like; has\_nape\_color: white; has\_primary\_color: grey; has\_belly\_pattern: solid; has\_breast\_color: grey; has\_bill\_length: about\_the\_same\_as\_head; has\_belly\_color: brown; has\_upperparts\_color: grey; has\_leg\_color: grey; has\_bill\_color: grey; has\_eye\_color: black; has\_breast\_pattern: solid; has\_tail\_shape: notched\_tail; has\_bill\_shape: all\-purposeSummer Tanager18has\_crown\_color: blue; has\_breast\_color: white; has\_belly\_color: brown; has\_belly\_color: white; has\_underparts\_color: brown; has\_nape\_color: brown; has\_breast\_color: brown; has\_belly\_pattern: solid; has\_leg\_color: black; has\_wing\_color: grey; has\_upper\_tail\_color: grey; has\_back\_color: grey; has\_belly\_color: grey; has\_bill\_length: shorter\_than\_head; has\_breast\_pattern: solid; has\_bill\_shape: all\-purpose; has\_back\_pattern: solid; has\_wing\_color: blackBlack Tern6has\_leg\_color: black; has\_wing\_color: black; has\_upper\_tail\_color: black; has\_belly\_pattern: solid; has\_underparts\_color: black; has\_back\_color: brownElegant Tern3has\_under\_tail\_color: white; has\_wing\_color: black; has\_primary\_color: greyLeast Tern4has\_nape\_color: white; has\_wing\_color: black; has\_bill\_shape: hooked\_seabird; has\_wing\_color: whiteGreen tailed Towhee2has\_throat\_color: grey; has\_wing\_color: blackBrown Thrasher4has\_back\_pattern: striped; has\_belly\_color: white; has\_wing\_shape: rounded\-wings; has\_upper\_tail\_color: brownSage Thrasher8has\_shape: duck\-like; has\_back\_pattern: multi\-colored; has\_breast\_pattern: solid; has\_nape\_color: black; has\_forehead\_color: grey; has\_size: small\_\(5\_\-\_9\_in\); has\_wing\_color: buff; has\_upper\_tail\_color: brownBlack capped Vireo14has\_breast\_pattern: multi\-colored; has\_wing\_pattern: solid; has\_upper\_tail\_color: grey; has\_back\_pattern: multi\-colored; has\_bill\_shape: hooked\_seabird; has\_belly\_pattern: solid; has\_belly\_color: grey; has\_wing\_pattern: spotted; has\_wing\_color: brown; has\_upperparts\_color: brown; has\_forehead\_color: blue; has\_forehead\_color: yellow; has\_under\_tail\_color: grey; has\_belly\_color: whiteBlue headed Vireo9has\_bill\_shape: hooked\_seabird; has\_forehead\_color: grey; has\_back\_color: grey; has\_wing\_color: grey; has\_upperparts\_color: grey; has\_throat\_color: yellow; has\_wing\_color: brown; has\_underparts\_color: grey; has\_underparts\_color: buffPhiladelphia Vireo4has\_crown\_color: grey; has\_back\_pattern: solid; has\_under\_tail\_color: black; has\_tail\_pattern: solidRed eyed Vireo14has\_upper\_tail\_color: black; has\_breast\_color: black; has\_throat\_color: buff; has\_underparts\_color: grey; has\_breast\_color: grey; has\_forehead\_color: black; has\_under\_tail\_color: black; has\_nape\_color: brown; has\_nape\_color: black; has\_belly\_color: brown; has\_eye\_color: black; has\_leg\_color: grey; has\_nape\_color: buff; has\_belly\_color: whiteWarbling Vireo4has\_nape\_color: grey; has\_head\_pattern: eyebrow; has\_underparts\_color: white; has\_back\_color: blackWhite eyed Vireo11has\_nape\_color: grey; has\_throat\_color: grey; has\_wing\_pattern: striped; has\_primary\_color: grey; has\_bill\_shape: cone; has\_bill\_length: about\_the\_same\_as\_head; has\_tail\_pattern: multi\-colored; has\_wing\_color: grey; has\_head\_pattern: plain; has\_leg\_color: grey; has\_leg\_color: blackBlack and white Warbler4has\_upperparts\_color: white; has\_back\_pattern: striped; has\_upperparts\_color: black; has\_upper\_tail\_color: buffBlack throated Blue Warbler16has\_breast\_pattern: multi\-colored; has\_bill\_shape: hooked\_seabird; has\_belly\_color: buff; has\_back\_pattern: multi\-colored; has\_bill\_color: buff; has\_upperparts\_color: grey; has\_wing\_shape: pointed\-wings; has\_nape\_color: black; has\_throat\_color: black; has\_breast\_color: white; has\_belly\_color: black; has\_breast\_pattern: solid; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_wing\_color: brown; has\_nape\_color: white; has\_wing\_pattern: multi\-coloredBlue winged Warbler6has\_breast\_color: brown; has\_under\_tail\_color: buff; has\_throat\_color: buff; has\_wing\_color: yellow; has\_head\_pattern: plain; has\_crown\_color: yellowCanada Warbler7has\_back\_pattern: multi\-colored; has\_bill\_shape: all\-purpose; has\_wing\_shape: pointed\-wings; has\_throat\_color: grey; has\_upperparts\_color: black; has\_breast\_pattern: multi\-colored; has\_under\_tail\_color: greyCerulean Warbler4has\_forehead\_color: blue; has\_underparts\_color: white; has\_size: very\_small\_\(3\_\-\_5\_in\); has\_bill\_shape: coneChestnut sided Warbler13has\_breast\_pattern: multi\-colored; has\_breast\_color: grey; has\_throat\_color: buff; has\_wing\_pattern: spotted; has\_shape: duck\-like; has\_tail\_pattern: striped; has\_belly\_color: buff; has\_underparts\_color: brown; has\_breast\_color: brown; has\_crown\_color: white; has\_crown\_color: grey; has\_bill\_color: buff; has\_eye\_color: blackGolden winged Warbler9has\_bill\_color: black; has\_leg\_color: grey; has\_wing\_shape: pointed\-wings; has\_back\_color: grey; has\_belly\_color: white; has\_wing\_color: black; has\_wing\_pattern: multi\-colored; has\_underparts\_color: white; has\_belly\_color: greyHooded Warbler8has\_breast\_color: yellow; has\_leg\_color: buff; has\_wing\_color: buff; has\_nape\_color: grey; has\_back\_pattern: multi\-colored; has\_belly\_color: white; has\_wing\_color: brown; has\_primary\_color: whiteKentucky Warbler8has\_back\_color: yellow; has\_upperparts\_color: grey; has\_underparts\_color: black; has\_wing\_shape: pointed\-wings; has\_upperparts\_color: white; has\_leg\_color: grey; has\_size: small\_\(5\_\-\_9\_in\); has\_forehead\_color: blackMourning Warbler8has\_eye\_color: black; has\_underparts\_color: grey; has\_upperparts\_color: black; has\_under\_tail\_color: white; has\_throat\_color: buff; has\_breast\_pattern: multi\-colored; has\_throat\_color: grey; has\_primary\_color: greyMyrtle Warbler11has\_forehead\_color: brown; has\_breast\_pattern: striped; has\_under\_tail\_color: buff; has\_nape\_color: brown; has\_back\_pattern: multi\-colored; has\_wing\_pattern: solid; has\_bill\_shape: dagger; has\_underparts\_color: yellow; has\_bill\_length: shorter\_than\_head; has\_wing\_shape: pointed\-wings; has\_leg\_color: blackNashville Warbler3has\_bill\_color: grey; has\_underparts\_color: buff; has\_forehead\_color: blueOrange crowned Warbler11has\_throat\_color: buff; has\_bill\_color: black; has\_bill\_shape: dagger; has\_underparts\_color: grey; has\_upperparts\_color: yellow; has\_back\_color: yellow; has\_wing\_pattern: multi\-colored; has\_back\_color: black; has\_back\_color: grey; has\_bill\_shape: all\-purpose; has\_primary\_color: blackPalm Warbler26has\_bill\_shape: dagger; has\_upperparts\_color: white; has\_wing\_color: white; has\_back\_color: black; has\_back\_color: white; has\_upper\_tail\_color: black; has\_wing\_color: black; has\_head\_pattern: eyebrow; has\_eye\_color: black; has\_breast\_pattern: multi\-colored; has\_underparts\_color: buff; has\_breast\_color: black; has\_throat\_color: white; has\_under\_tail\_color: black; has\_nape\_color: black; has\_belly\_color: brown; has\_belly\_color: white; has\_tail\_pattern: striped; has\_primary\_color: black; has\_wing\_pattern: spotted; has\_under\_tail\_color: white; has\_size: small\_\(5\_\-\_9\_in\); has\_forehead\_color: blue; has\_wing\_pattern: striped; has\_wing\_pattern: multi\-colored; has\_throat\_color: greyPine Warbler1has\_crown\_color: yellowPrairie Warbler4has\_crown\_color: yellow; has\_size: medium\_\(9\_\-\_16\_in\); has\_nape\_color: yellow; has\_primary\_color: whiteSwainson Warbler6has\_bill\_shape: hooked\_seabird; has\_tail\_pattern: striped; has\_upperparts\_color: black; has\_wing\_pattern: solid; has\_forehead\_color: brown; has\_breast\_pattern: solidTennessee Warbler5has\_crown\_color: grey; has\_wing\_shape: pointed\-wings; has\_bill\_color: black; has\_back\_pattern: multi\-colored; has\_crown\_color: brownWilson Warbler2has\_back\_color: yellow; has\_leg\_color: blackWorm eating Warbler6has\_eye\_color: black; has\_forehead\_color: white; has\_wing\_pattern: solid; has\_upperparts\_color: black; has\_forehead\_color: black; has\_forehead\_color: yellowYellow Warbler12has\_back\_color: yellow; has\_breast\_pattern: solid; has\_bill\_color: grey; has\_primary\_color: yellow; has\_bill\_shape: cone; has\_forehead\_color: grey; has\_wing\_pattern: striped; has\_breast\_color: grey; has\_breast\_color: buff; has\_throat\_color: grey; has\_eye\_color: black; has\_back\_pattern: multi\-coloredNorthern Waterthrush5has\_head\_pattern: eyebrow; has\_under\_tail\_color: black; has\_underparts\_color: buff; has\_wing\_color: grey; has\_wing\_shape: pointed\-wingsLouisiana Waterthrush7has\_head\_pattern: eyebrow; has\_breast\_color: black; has\_underparts\_color: buff; has\_tail\_shape: notched\_tail; has\_underparts\_color: brown; has\_wing\_color: grey; has\_tail\_pattern: stripedBohemian Waxwing6has\_upperparts\_color: grey; has\_wing\_shape: pointed\-wings; has\_belly\_color: grey; has\_wing\_shape: rounded\-wings; has\_tail\_pattern: multi\-colored; has\_forehead\_color: greyCedar Waxwing3has\_tail\_pattern: multi\-colored; has\_nape\_color: buff; has\_upper\_tail\_color: brownAmerican Three toed Woodpecker10has\_bill\_shape: dagger; has\_breast\_pattern: solid; has\_leg\_color: grey; has\_head\_pattern: eyebrow; has\_wing\_color: black; has\_back\_pattern: multi\-colored; has\_wing\_shape: pointed\-wings; has\_tail\_pattern: striped; has\_belly\_pattern: solid; has\_tail\_pattern: solidPileated Woodpecker17has\_under\_tail\_color: black; has\_bill\_color: grey; has\_upper\_tail\_color: black; has\_wing\_shape: pointed\-wings; has\_forehead\_color: white; has\_leg\_color: grey; has\_primary\_color: grey; has\_upper\_tail\_color: buff; has\_breast\_color: yellow; has\_throat\_color: yellow; has\_bill\_length: shorter\_than\_head; has\_eye\_color: black; has\_wing\_color: grey; has\_forehead\_color: brown; has\_size: medium\_\(9\_\-\_16\_in\); has\_underparts\_color: white; has\_nape\_color: whiteRed bellied Woodpecker6has\_bill\_shape: dagger; has\_tail\_pattern: striped; has\_wing\_color: brown; has\_back\_pattern: solid; has\_crown\_color: black; has\_wing\_shape: pointed\-wingsRed cockaded Woodpecker5has\_nape\_color: black; has\_head\_pattern: eyebrow; has\_belly\_pattern: solid; has\_wing\_shape: pointed\-wings; has\_throat\_color: whiteDowny Woodpecker25has\_breast\_pattern: multi\-colored; has\_underparts\_color: yellow; has\_upperparts\_color: brown; has\_upperparts\_color: black; has\_underparts\_color: black; has\_eye\_color: black; has\_back\_pattern: multi\-colored; has\_bill\_length: about\_the\_same\_as\_head; has\_primary\_color: white; has\_wing\_color: black; has\_back\_pattern: solid; has\_back\_color: white; has\_forehead\_color: white; has\_size: small\_\(5\_\-\_9\_in\); has\_wing\_shape: pointed\-wings; has\_shape: perching\-like; has\_bill\_shape: all\-purpose; has\_upperparts\_color: white; has\_wing\_pattern: striped; has\_underparts\_color: grey; has\_underparts\_color: buff; has\_bill\_length: shorter\_than\_head; has\_bill\_shape: cone; has\_wing\_color: brown; has\_breast\_pattern: solidBewick Wren4has\_tail\_pattern: striped; has\_wing\_pattern: solid; has\_tail\_shape: notched\_tail; has\_bill\_shape: all\-purposeCactus Wren3has\_wing\_pattern: spotted; has\_crown\_color: white; has\_head\_pattern: eyebrowCarolina Wren8has\_tail\_pattern: striped; has\_underparts\_color: grey; has\_nape\_color: buff; has\_forehead\_color: brown; has\_bill\_color: grey; has\_primary\_color: white; has\_forehead\_color: blue; has\_crown\_color: blueHouse Wren3has\_tail\_pattern: striped; has\_belly\_color: buff; has\_belly\_color: whiteMarsh Wren10has\_tail\_pattern: striped; has\_wing\_pattern: spotted; has\_back\_color: grey; has\_forehead\_color: yellow; has\_back\_pattern: multi\-colored; has\_eye\_color: black; has\_wing\_shape: rounded\-wings; has\_forehead\_color: brown; has\_breast\_pattern: solid; has\_back\_color: blackRock Wren4has\_back\_pattern: striped; has\_crown\_color: brown; has\_belly\_color: brown; has\_under\_tail\_color: blackWinter Wren3has\_breast\_color: brown; has\_breast\_pattern: striped; has\_size: medium\_\(9\_\-\_16\_in\)

Similar Articles

Towards Fine-Grained and Verifiable Concept Bottleneck Models

arXiv cs.LG

This paper proposes a fine-grained concept bottleneck model framework that grounds each concept in localized visual evidence, enabling direct verification of concept correctness and improving transparency in medical imaging tasks.

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Hugging Face Daily Papers

This paper introduces BDH-CQ, a 150M-parameter reasoning model that combines in-context learning with recurrent latent reasoning, achieving 29.5% pass@2 on ARC-AGI-1 at very low inference cost and establishing a new cost-accuracy frontier.

Hoeffding Concept Bottleneck Models with Applications to Overhead Images

arXiv cs.LG

Introduces Hoeffding Concept Bottleneck Models (HCBM), a nonlinear and sparse aggregation of concept scores using Hoeffding functional decomposition of gradient-boosted trees, for improved explainability and accuracy in classification and object detection tasks, with applications to overhead images.

Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation

arXiv cs.LG

This paper introduces concept modulation models (CMMs), a unified framework for identifiability and extrapolation in conditional generative models. It shows that feature agreement on observed attributes induces constraints through attribute potentials, enabling algebraic extrapolation criteria that recover and generalize existing results.