Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks

arXiv cs.AI Papers

Summary

This paper introduces contrastive explanations for Quantitative Bipolar Argumentation Frameworks, explaining differences between two topic arguments to enhance AI explainability, with applications in healthcare and bias identification.

arXiv:2609.02399v1 Announce Type: new Abstract: Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e.g. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability. In this paper, we introduce contrastive explanations for Quantitative Bipolar Argumentation Frameworks (QBAFs), one such formalism. Unlike most existing explanations for QBAFs, which explain the reasoning outcome of a single argument of interest (i.e. a topic argument), contrastive explanations explain the difference between two topic arguments. We introduce a general form of contrastive attribution functions (CAFs) and establish a set of general properties they should satisfy. We introduce CAFs based on removal, gradients and Shapley-values, and study their properties. Finally, to illustrate contrastive explanations, we demonstrate their usefulness in healthcare and bias identification settings.
Original Article
View Cached Full Text

Cached at: 09/03/26, 06:05 AM

# Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks
Source: [https://arxiv.org/html/2609.02399](https://arxiv.org/html/2609.02399)
###### Abstract

Argumentation frameworks are useful tools for representing and reasoning with information in a variety of settings, e\.g\. in supplementing AI models as they perform classification tasks, with a notable benefit of providing additional explainability\. In this paper, we introduce*contrastive explanations*for Quantitative Bipolar Argumentation Frameworks \(QBAFs\), one such formalism\. Unlike most existing explanations for QBAFs, which explain the reasoning outcome of a single argument of interest \(i\.e\. a*topic argument*\), contrastive explanations explain the difference between two topic arguments\. We introduce a general form of contrastive attribution functions \(CAFs\) and establish a set of general properties they should satisfy\. We introduce CAFs based on removal, gradients and Shapley values, and study their properties\. Finally, to illustrate contrastive explanations, we demonstrate their usefulness in healthcare and bias identification settings\.

1Imperial College London, UK

2Cardiff University, UK

3King’s College London, UK

\{x\.yin20, f\.toni\}@imperial\.ac\.uk, potykan@cardiff\.ac\.uk, antonio\.rago@kcl\.ac\.uk

## 1Introduction

Argumentation frameworks have recently emerged as useful tools for representing and reasoning with information in a range of settings, particularly those where explainability may otherwise be lacking, e\.g\. in image classification\([Ayoobi et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib6);[Kori et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib2)\), recommendation\([Rago et al\. 2021](https://arxiv.org/html/2609.02399#bib.bib5);[Rago et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib7)\)and claim verification\([Freedman et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib3);[Zhu et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib4)\)\. Quantitative Bipolar Argumentation Frameworks \(QBAFs\)\([Baroni et al\. 2015](https://arxiv.org/html/2609.02399#bib.bib26)\)are one such formal model for reasoning with conflicting and supporting information\. A QBAF consists of a set of*arguments*,*attack*and*support*relations among them, and a*base score function*assigning each argument a prior strength\. To evaluate a QBAF, a gradual semantics \(e\.g\.,\([Rago et al\. 2016](https://arxiv.org/html/2609.02399#bib.bib18)\)\) is applied to update the base scores by aggregating the influence of attackers and supporters, resulting in a final*strength*for each argument\. QBAFs are well suited to model decision\-support tasks \(e\.g\.,\([Cocarascu et al\. 2019](https://arxiv.org/html/2609.02399#bib.bib24);[Rago et al\. 2021](https://arxiv.org/html/2609.02399#bib.bib5)\)\): the final strengths can serve as decision scores, while the explicit attack and support structure provides a transparent account of the reasoning process \(e\.g\.\([Cocarascu et al\. 2019](https://arxiv.org/html/2609.02399#bib.bib24)\)\)\. Figure[1](https://arxiv.org/html/2609.02399#S1.F1)shows a toy QBAF for an animal classification problem\. The possible classes \(zebra,horse,tiger\), and the relevant features \(stripes,herbivore\) are represented as arguments\. Support and attack relations indicate whether a feature supports or attacks a class\. All arguments are assigned a neutral base score of0\.50\.5, and the DF\-QuAD semantics\([Rago et al\. 2016](https://arxiv.org/html/2609.02399#bib.bib18)\)is applied to compute the strengths of the class arguments\. Here,zebrais predicted as it has the highest strength of0\.8750\.875, whilehorseandtigerboth have strength0\.50\.5\.

![Refer to caption](https://arxiv.org/html/2609.02399v1/figures/image.png)Figure 1:An example QBAF for an animal classification task\. Squares represent arguments; green and red edges indicate support and attack relations, respectively\. The contrastive explanation for zebra over horse is stripes, whereas that for zebra over tiger is being a herbivore\.Given this prediction, a user may ask: whyzebra? A possible explanation is that the animal hasstripesand isherbivorous\. However, research in philosophy, psychology, and social science shows that explanations are contrastive: when people ask a “WhyPP?” question, they often implicitly ask “whyPPrather thanQQ?”, wherePPis a*fact*whileQQis some*contrast case*\([Lipton 1990](https://arxiv.org/html/2609.02399#bib.bib27);[Hilton 1990](https://arxiv.org/html/2609.02399#bib.bib28);[Van Bouwel and Weber 2002](https://arxiv.org/html/2609.02399#bib.bib29);[Ylikoski 2007](https://arxiv.org/html/2609.02399#bib.bib36);[Chin\-Parker and Cantelon 2017](https://arxiv.org/html/2609.02399#bib.bib30);[Miller 2019](https://arxiv.org/html/2609.02399#bib.bib37);[Miller 2021](https://arxiv.org/html/2609.02399#bib.bib12)\)\. In our example, if the question is why the prediction iszebrarather thanhorse, the*contrastive explanation*should highlightstripes, since bothzebraandhorseareherbivorous\. In contrast, if the question is why the prediction iszebrarather thantiger, the contrastive explanation should highlightherbivore, since this feature distinguisheszebrafromtiger\(in this example\), whereasstripesdoes not\. Thus, contrastive explanations aim to identify the discriminative reasons \(referred to as*difference conditions*in\([Lipton 1990](https://arxiv.org/html/2609.02399#bib.bib27)\)\) that distinguishPPfrom a contrast caseQQ, rather than listing all explanatory reasons\. This makes contrastive explanations simpler, more feasible, and cognitively less demanding\([Miller 2021](https://arxiv.org/html/2609.02399#bib.bib12)\)\.

However, existing explanation methods for QBAFs, such as attribution explanations\([Yin et al\. 2023](https://arxiv.org/html/2609.02399#bib.bib9)\)and counterfactual explanations\([Yin et al\. 2024a](https://arxiv.org/html/2609.02399#bib.bib11)\), mainly focus on explaining the strength of a single argument of interest \(*topic argument*\)\. In other words, they answer “whyPP” but overlook the contrast question “whyPPrather thanQQ”, and thus fail to capture the “why notQQ” aspect\. This motivates the central research question of this paper:Given two topic arguments in a QBAF, how can we explain the difference between their strengths?

To address this question, we introduce attribution\-based*contrastive explanations*for QBAFs\. Our key idea is to assign attribution scores to all non\-topic arguments with respect to the strength gap between two topic arguments\. Attribution scores are well suited to this purpose because they provide an intuitive quantitative measure of each argument’s influence on the contrast\. By comparing these scores, we can identify arguments that discriminate between the two topic arguments without manually inspecting all reasoning paths\. For example, when explaining whyzebrarather thanhorse,stripesshould receive a larger attribution score thanherbivore, as the latter supports both classes\.

The contributions of this paper are as follows:

- •We introduce contrastive explanations and desirable properties that they should satisfy \(Sections[3](https://arxiv.org/html/2609.02399#S3),[4](https://arxiv.org/html/2609.02399#S4)\);
- •We study three concrete contrastive explanation methods and their properties \(Sections[5](https://arxiv.org/html/2609.02399#S5),[6](https://arxiv.org/html/2609.02399#S6),[7](https://arxiv.org/html/2609.02399#S7)\);
- •We illustrate the use of contrastive explanations in healthcare \(Section[8](https://arxiv.org/html/2609.02399#S8)\) and bias detection \(Section[9](https://arxiv.org/html/2609.02399#S9)\)\.

## 2Preliminaries

We recall the definition of QBAFs\([Baroni et al\. 2015](https://arxiv.org/html/2609.02399#bib.bib26)\)\.

###### Definition 1\(QBAF\)\.

A*QBAF*is a quadruple𝒬=⟨𝒜,ℛ−,ℛ\+,τ⟩\\mathcal\{Q\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\\right\\ranglewhere𝒜\\mathcal\{A\}is a finite set of*arguments*;ℛ−,ℛ\+⊆𝒜×𝒜\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\}\\subseteq\\mathcal\{A\}\\times\\mathcal\{A\}are*attack*and*support*relations such thatℛ−∩ℛ\+=∅\\mathcal\{R\}^\{\-\}\\cap\\mathcal\{R\}^\{\+\}=\\emptyset;τ:𝒜→\[0,1\]\\tau:\\mathcal\{A\}\\rightarrow\[0,1\]is a*base score function*\.

In the remainder, unless specified otherwise, we will assume a QBAF𝒬=⟨𝒜,ℛ−,ℛ\+,τ⟩\\mathcal\{Q\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\\right\\rangleas given\. At several places, we will consider QBAF restrictions\([Kampik et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib23)\)induced by subsets of arguments\.

###### Definition 2\(QBAF Restriction\)\.

Given a subset of arguments𝒜′⊆𝒜\\mathcal\{A\}^\{\\prime\}\\subseteq\\mathcal\{A\}, the QBAF restriction of𝒬\\mathcal\{Q\}to𝒜′\\mathcal\{A\}^\{\\prime\}is𝒬↓𝒜′=⟨𝒜′,ℛ−∩\(𝒜′×𝒜′\),ℛ\+∩\(𝒜′×𝒜′\),τ∩\(𝒜′×\[0,1\]\)⟩\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}=\\left\\langle\\mathcal\{A\}^\{\\prime\},\\mathcal\{R\}^\{\-\}\\cap\(\\mathcal\{A\}^\{\\prime\}\\times\\mathcal\{A\}^\{\\prime\}\),\\mathcal\{R\}^\{\+\}\\cap\(\\mathcal\{A\}^\{\\prime\}\\times\\mathcal\{A\}^\{\\prime\}\),\\tau\\cap\(\\mathcal\{A\}^\{\\prime\}\\times\[0,1\]\)\\right\\rangle\.

We use gradual semantics to assign a*dialectical strength*to each argument in a QBAF\. Examples of gradual semantics include DF\-QuAD\([Rago et al\. 2016](https://arxiv.org/html/2609.02399#bib.bib18)\), Restricted Euler\-based semantics \(REB\)\([Amgoud and Ben\-Naim 2018](https://arxiv.org/html/2609.02399#bib.bib17)\), and Quadratic Energy semantics \(QE\)\([Potyka 2018](https://arxiv.org/html/2609.02399#bib.bib16)\)\.

###### Definition 3\(Gradual Semantics\)\.

A*gradual semantics*is a functionσ:𝒜→\[0,1\]∪\{⊥\}\\sigma:\\mathcal\{A\}\\rightarrow\[0,1\]\\cup\\\{\\bot\\\}\. We callσ⁡\(α\)\\sigma\(\\alpha\)the*strength*ofα\\alphaand say that it is*undefined*iffσ\(α\)=⊥\\sigma\(\\alpha\)=\\bot\.

All gradual semantics that we are aware of are instances of the class of*modular semantics*\([Mossakowski and Neuhaus 2018](https://arxiv.org/html/2609.02399#bib.bib38)\)and we will make use of some of their properties later\. Roughly speaking, modular semantics compute strength values by an update process that starts from the base scores, and repeatedly updates the strengths of arguments based on the strengths of their attackers and supporters\. The update function of modular semantics can be decomposed into an aggregation function that aggregates the strengths of attackers and supporters, and an influence function that uses the aggregate to adapt the base score\.

An undefined strength value \(σ\(α\)=⊥\\sigma\(\\alpha\)=\\bot\) may arise in some cyclic QBAFs when the update process fails to converge\([Mossakowski and Neuhaus 2018](https://arxiv.org/html/2609.02399#bib.bib38)\)\. However, in all known cases, the convergence problems can be solved by continuizing the semantics\([Potyka 2019](https://arxiv.org/html/2609.02399#bib.bib39);[Potyka and Booth 2024a](https://arxiv.org/html/2609.02399#bib.bib40)\)\. In the following, we will focus on QBAFs for which all strength values are defined\.

To explain the strength of a*topic argument*α∈𝒜\\alpha\\in\\mathcal\{A\}, attribution functions \(AFs\)ϕα:𝒜∖\{α\}→ℝ\\phi\_\{\\alpha\}:\\mathcal\{A\}\\setminus\\\{\\alpha\\\}\\rightarrow\\mathbb\{R\}quantify the*influence*of arguments onα\\alpha\([Kampik et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib23)\)\. In general, a larger magnitude ofϕα​\(β\)\\phi\_\{\\alpha\}\(\\beta\)indicates a stronger influence, while its sign indicates whether the influence is*positive*or*negative*\.

Removal\-based AFs measure by how much the strength ofα\\alphachanges whenβ\\betais removed from𝒬\\mathcal\{Q\}\.

###### Definition 4\(Removal\-based AF\)\.

For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andα≠β\\alpha\\neq\\beta, let𝒜′=𝒜∖\{β\}\\mathcal\{A\}^\{\\prime\}=\\mathcal\{A\}\\setminus\\\{\\beta\\\}, the*removal\-based attribution*fromβ\\betatoα\\alphaisϕαR\(β\)=σ𝒬\(α\)−σ𝒬↓𝒜′\(α\)\\phi\_\{\\alpha\}^\{R\}\(\\beta\)=\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\.

Gradient\-based AFs measure the sensitivity ofα\\alphawith respect to small changes in the base score ofβ\\beta\.

###### Definition 5\(Gradient\-based AF\)\.

For allα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\},α≠β\\alpha\\neq\\beta, andϵ∈\[−τ\(β\),0\)∪\(0,1−τ\(β\)\]\\epsilon\\in\[\-\\tau\(\\beta\),0\)\\cup\(0,1\-\\tau\(\\beta\)\], let𝒬ϵ=⟨𝒜,ℛ−,ℛ\+,τϵ⟩\\mathcal\{Q\}\_\{\\epsilon\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\_\{\\epsilon\}\\right\\rangle, whereτϵ​\(β\)=τ⁡\(β\)\+ϵ\\tau\_\{\\epsilon\}\(\\beta\)=\\tau\(\\beta\)\+\\epsilonandτϵ​\(γ\)=τ⁡\(γ\)\\tau\_\{\\epsilon\}\(\\gamma\)=\\tau\(\\gamma\)for allγ∈𝒜∖\{β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\beta\\\}\. The*gradient\-based attribution*fromβ\\betatoα\\alphais defined asϕαG​\(β\)=limϵ→0σ𝒬ϵ​\(α\)−σ𝒬​\(α\)ϵ\\phi\_\{\\alpha\}^\{G\}\(\\beta\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\sigma\_\{\\mathcal\{Q\}\_\{\\epsilon\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\}\{\\epsilon\}\.

Shapley\-based AFs measure the average contribution ofβ\\betato the strength ofα\\alphawhen addingβ\\betato a subgraph of𝒬\\mathcal\{Q\}\.

###### Definition 6\(Shapley\-based AF\)\.

Forα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andα≠β\\alpha\\neq\\beta, the*Shapley\-based attribution*fromβ\\betatoα\\alphais defined as

ϕαS​\(β\)=∑B⊆𝒜α¯∖\{β\}w⁡\(B\)⋅cβ​\(B\),\\phi\_\{\\alpha\}^\{S\}\(\\beta\)=\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}w\(B\)\\cdot c\_\{\\beta\}\(B\),where𝒜α¯=𝒜∖\{α\}\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}=\\mathcal\{A\}\\setminus\\\{\\alpha\\\},w⁡\(B\)=\|B\|\!⋅\(\|𝒜α¯\|−\|B\|−1\)\!\|𝒜α¯\|\!w\(B\)=\\frac\{\|B\|\!\\cdot\\left\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\-1\\right\)\!\}\{\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\!\}andcβ\(B\)=σ𝒬↓B∪\{α,β\}\(α\)−σ𝒬↓B∪\{α\}\(α\)\.c\_\{\\beta\}\(B\)=\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{B\\cup\\\{\\alpha,\\beta\\\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{B\\cup\\\{\\alpha\\\}\}\}\}\(\\alpha\)\.

## 3Contrastive Explanations

To explain the strength difference between two topic argumentsα\\alphaandβ\\beta, we introduce*contrastive attribution functions \(CAFs\)*, which quantify the influence of arguments on their relative strength\.

###### Notation 1\.

For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}, we writeα⪰β\\alpha\\succeq\\betato denote the contrastα\\alphais stronger thanβ\\beta\.

###### Definition 7\(Contrastive Attribution Function \(CAF\)\)\.

A CAF forα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}is a functionΦα⪰β:𝒜∖\{α,β\}→ℝ\\Phi\_\{\\alpha\\succeq\\beta\}:\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}\\rightarrow\\mathbb\{R\}\.

Intuitively, a*positive influence*Φα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0means thatγ\\gammacontributes more favourably toα\\alphathan toβ\\beta\. This may occur whenγ\\gammasupportsα\\alphamore strongly thanβ\\beta, attacksα\\alphaless strongly thanβ\\beta, or supportsα\\alphawhile attackingβ\\beta\. Conversely, a*negative influence*means thatγ\\gammacontributes more favourably toβ\\betarelative toα\\alpha\. IfΦα⪰β​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=0,γ\\gammacontributes equally toα\\alphaandβ\\betaand we call the influence*neutral*\.

To narrow down the choice of a CAF, we suggest some desirable properties that a CAF should satisfy\.

## 4Desirable Properties of CAFs

To begin with, if an argument positively affectsα⪰β\\alpha\\succeq\\beta, then it should negatively affectβ⪰α\\beta\\succeq\\alphaby the same amount\.

###### Property 1\(Antisymmetry\)\.

For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\},Φα⪰β​\(γ\)=−Φβ⪰α​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)\.

###### Proposition 1\.

IfΦ\\Phisatisfies antisymmetry, thenΦα⪰α​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\)\\\!\\\!=\\\!\\\!0\.

If an attribution functionϕ\\phiis well suited to measure the influence of arguments in a domain, we may want that a CAF is calibrated w\.r\.t\.ϕ\\phiin the following sense\.

###### Property 2\(ϕ\\phi\-Calibration\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, ifϕα​\(γ\)=a\\phi\_\{\\alpha\}\(\\gamma\)=aandϕβ​\(γ\)=0\\phi\_\{\\beta\}\(\\gamma\)=0, thenΦα⪰β​\(γ\)=a\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=a\.

Let us note that antisymmetry implies that a symmetrical calibration property holds for the second argument\.

###### Proposition 2\(Inverseϕ\\phi\-Calibration\)\.

IfΦ\\Phisatisfies antisymmetry andϕ\\phi\-Calibration, thenϕα​\(γ\)=0\\phi\_\{\\alpha\}\(\\gamma\)=0andϕβ​\(γ\)=b\\phi\_\{\\beta\}\(\\gamma\)=bimpliesΦα⪰β​\(γ\)=−b\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-b\.

If the influence ofγ\\gammaon bothα\\alphaandβ\\betais equal according to a reference attribution functionϕ\\phi, then the influence ofγ\\gammaonα⪰β\\alpha\\succeq\\betashould be00\.

###### Property 3\(ϕ\\phi\-Neutrality\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}ifϕα​\(γ\)=ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)=\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=0\.

Similarly, ifγ\\gammainfluencesα\\alphastronger thanβ\\betaaccording to a reference attribution functionϕ\\phi, then the influence ofγ\\gammaonα⪰β\\alpha\\succeq\\betashould be positive\.

###### Property 4\(ϕ\\phi\-Monotonicity\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, ifϕα​\(γ\)\>ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)\>\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0\.

Antisymmetry guarantees again a symmetric behaviour for the case thatγ\\gammainfluencesα\\alphaless thanβ\\beta\.

###### Proposition 3\.

IfΦ\\Phisatisfies antisymmetry andϕ\\phi\-Monotonicity, thenϕα​\(γ\)<ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)<\\phi\_\{\\beta\}\(\\gamma\)impliesΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0\.

## 5Derived CAFs

Since attribution functions measure the influence of arguments on an individual argument, one natural idea to define CAFs is to combine the individual attribution values by a binary function\.

###### Definition 8\(Derived CAFs\)\.

A CAFΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is called derived from an attribution functionϕ\\phiif there is a functionf:ℝ2→ℝf:\\mathbb\{R\}^\{2\}\\rightarrow\\mathbb\{R\}such that for anyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\},Φα⪰β​\(γ\)=f⁡\(ϕα​\(γ\),ϕβ​\(γ\)\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=f\(\\phi\_\{\\alpha\}\(\\gamma\),\\phi\_\{\\beta\}\(\\gamma\)\)\.

As we show next, the properties from the previous section can be satisfied by inducing a CAF from a reference attribution function using subtraction\.

###### Proposition 4\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived from an attribution functionϕ\\phiusing subtraction, thenΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies Antisymmetry,ϕ\\phi\-Calibration,ϕ\\phi\-Neutrality andϕ\\phi\-Monotonicity\.

While subtraction\-derived CAFs satisfy all previously proposed properties, there could still be other functions that lead to the same properties\. We can characterise subtraction\-derived measures by adding the following additivity property\.

###### Property 5\(Additivity\)\.

For anyα,β,η∈𝒜\\alpha,\\beta,\\eta\\in\\mathcal\{A\}and anyγ∈𝒜∖\{α,β,η\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\eta\\\},Φα⪰β​\(γ\)=Φα⪰η​\(γ\)\+Φη⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\Phi\_\{\\alpha\\succeq\\eta\}\(\\gamma\)\+\\Phi\_\{\\eta\\succeq\\beta\}\(\\gamma\)\.

###### Proposition 5\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived using subtraction, thenΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies additivity\.

Additivity can be used to fully characterise subtraction\-derived CAFs for all*plausible*attribution functions\. By plausible we mean the following: all gradual semantics that we are aware of belong to the class of*modular semantics*and all modular semantics satisfy the*independence*property which states that arguments can only affect each other if there is a directed path between them\([Potyka and Booth 2024b](https://arxiv.org/html/2609.02399#bib.bib41)\)\. Consequently, when adding a new argument to a graph without connecting it to any other arguments, the attribution value of this argument should be00for all other arguments\.

###### Definition 9\(Plausibility\)\.

An attribution functionϕ\\phiis called*plausible*if for each QBAF𝒬\\mathcal\{Q\}and the QBAF𝒬′\\mathcal\{Q\}^\{\\prime\}resulting from𝒬\\mathcal\{Q\}by adding a single argumentη\\eta\(and no edges\), it holds that \(1\) the attribution values of all arguments from𝒬\\mathcal\{Q\}are equal in both𝒬\\mathcal\{Q\}and𝒬′\\mathcal\{Q\}^\{\\prime\}, and \(2\)ϕα​\(η\)=0\\phi\_\{\\alpha\}\(\\eta\)=0for all argumentsα\\alphafrom𝒬\\mathcal\{Q\}\.

All AFs introduced previously are plausible\.

###### Proposition 6\.

ϕαR,ϕαG\\phi\_\{\\alpha\}^\{R\},\\phi\_\{\\alpha\}^\{G\}andϕαS\\phi\_\{\\alpha\}^\{S\}are plausible AFs under all modular semantics\.

We have the following characterisation of subtraction\-derived measures\.

###### Proposition 7\.

IfΦ\\Phiis a CAF derived from a plausible attribution functionϕ\\phi, andΦ\\Phisatisfies antisymmetry,ϕ\\phi\-calibration and additivity, then, under all modular gradual semantics,Φ\\Phiis equal to the CAF derived fromϕ\\phiusing subtraction\.

Note that, based on the choice ofϕ\\phi, the derived CAF will be calibrated differently, and so CAFs derived from different attribution functions will usually lead to different CAFs\. In the following proposition, we present compact formulas for CAFs derived from our previously introduced AFs\.

###### Proposition 8\.

The CAFsΦα⪰βR,Φα⪰βG,Φα⪰βS\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\},\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\},\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}derived from the removal\-based, gradient\-based and Shapley\-based AFs using subtraction are defined as follows:

Φα⪰βR\(γ\)=\(σ𝒬\(α\)−σ𝒬\(β\)\)−\(σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\)\.\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)=\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)\.Φα⪰βG​\(γ\)=limϵ→0\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)ϵ\.\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{\\epsilon\}\.Φα⪰βS​\(γ\)=∑B⊆𝒜∖\{α,β,γ\}\(w⁡\(B\)⋅\(cγ→α​\(B\)−cγ→β​\(B\)\)CLOSE\\displaystyle\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\gamma\\\}\}\\big\(w\(B\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\)\)\+w\(B∪\{β\}\)⋅\(cγ→α\(B∪\{β\}\)−cγ→β\(B∪\{α\}\)\)\),\\displaystyle\\ \\ \+w\(B\\cup\\\{\\beta\\\}\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\\cup\\\{\\beta\\\}\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\\cup\\\{\\alpha\\\}\)\)\\big\),wherecγ→x\(X\)=σ𝒬↓X∪\{x,γ\}\(x\)−σ𝒬↓X∪\{x\}\(x\)c\_\{\\gamma\\rightarrow x\}\(X\)=\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{X\\cup\\\{x,\\gamma\\\}\}\}\}\(x\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{X\\cup\\\{x\\\}\}\}\}\(x\)\.

## 6Computing Subtraction\-Derived CAFs

In applications, our topic arguments often correspond to alternatives that we can choose from\. For example, Figure[2](https://arxiv.org/html/2609.02399#S8.F2)in Section[8](https://arxiv.org/html/2609.02399#S8)shows a medical decision QBAF where we can choose one of three different treatments \(antiviral, antibiotic or bronchodilator\)\. Now consider a problem withTTtopic argumentst1,…,tTt\_\{1\},\\dots,t\_\{T\}\. When computing the impact of an argument on all preferencesti⪰tjt\_\{i\}\\succeq t\_\{j\}naively, we requireT⋅\(T−1\)=O⁡\(T2\)T\\cdot\(T\-1\)=O\(T^\{2\}\)CAF calls\. In applications like healthcare, where we may have dozens of different diagnoses or treatments, this can be too expensive\. We can exploit properties of subtraction\-derived CAFs to reduce the number of CAF calls significantly\.

To begin with, note that anti\-symmetry allows us to computeΦti⪰tj​\(γ\)\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)fromΦtj⪰ti​\(γ\)\\Phi\_\{t\_\{j\}\\succeq t\_\{i\}\}\(\\gamma\)\. Hence, it suffices to consider only preferencesti⪰tjt\_\{i\}\\succeq t\_\{j\}withi<ji<j\. While this reduces the number of computations toT⋅\(T−1\)2=O⁡\(T2\)\\frac\{T\\cdot\(T\-1\)\}\{2\}=O\(T^\{2\}\), it remains quadratic asymptotically\.

Additivity allows us to design a dynamic programming algorithm that requires only a linear number of CAF calls\.

Algorithm 1Dynamic Programming CAF ComputationInput: A QBAF𝒬\\mathcal\{Q\}, CAFΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}, set of topic arguments\{t1,…,tT\}\\\{t\_\{1\},\\dots,t\_\{T\}\\\}, query argumentγ\\gamma\. Output: MapMMsuch thatM⁡\[i,j\]=Φti⪰tj​\(γ\)M\[i,j\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)for all1≤i<j≤T1\\leq i<j\\leq T\.

1:Initialise empty map

MM
2:for

i=1i=1to

T−1T\-1do

3:

M⁡\[i,i\+1\]←Φti⪰ti\+1​\(γ\)M\[i,i\+1\]\\leftarrow\\Phi\_\{t\_\{i\}\\succeq t\_\{i\+1\}\}\(\\gamma\)
4:for

i=1i=1to

T−2T\-2do

5:for

j=i\+2j=i\+2to

TTdo

6:

M⁡\[i,j\]=M⁡\[i,j−1\]\+M⁡\[j−1,j\]M\[i,j\]=M\[i,j\-1\]\+M\[j\-1,j\]\.

7:return

MM

###### Proposition 9\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies additivity, Algorithm[1](https://arxiv.org/html/2609.02399#alg1)computesΦti⪰tj​\(γ\)\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)for all1≤i<j≤T1\\leq i<j\\leq TwithO⁡\(T\)O\(T\)Φα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}\-calls\. IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}can be computed in timeO⁡\(C\)O\(C\), the overall time complexity of the algorithm isO⁡\(T⋅C\+T2\)O\(T\\cdot C\+T^\{2\}\)\.

## 7Method\-Specific Properties

We regard the properties introduced in Section[4](https://arxiv.org/html/2609.02399#S4)as desirable for all CAFs and they are indeed satisfied by all previously introduced CAFs\. To distinguish our CAFs axiomatically, we introduce some additional properties inspired by\([Kampik et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib23)\)to separate them\.

*Counterfactuality*captures the intuition that removing an argument with positive \(negative\) attribution should decrease \(increase\) the strength difference between the topic arguments\.

###### Property 6\(Counterfactuality\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Counterfactuality*iff, for anyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, letting𝒜′=𝒜∖\{γ\}\\mathcal\{A\}^\{\\prime\}=\\mathcal\{A\}\\setminus\\\{\\gamma\\\}, the following statements hold: 1\. IfΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0, thenσ𝒬\(α\)−σ𝒬\(β\)<σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)<\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\); 2\. IfΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0, thenσ𝒬\(α\)−σ𝒬\(β\)\>σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\>\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\.

###### Proposition 10\.

Φα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)satisfies Counterfactuality, whileΦα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)andΦα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)can violate Counterfactuality\.

Local faithfulness constrains the influence of an argument on the strength gap between two topic arguments when perturbing its base score\. An argument with positive \(negative\) influence will locally increase \(decrease\) the strength gap between two topic arguments when its base score is slightly increased\.

###### Property 7\(Local Faithfulness\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Local Faithfulness*wrt\.σ\\sigmaiff, for anyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, there existsδ\>0\\delta\>0such that, for alle∈\[τ⁡\(γ\)−δ,τ⁡\(γ\)\+δ\]∩\[0,1\]e\\in\[\\tau\(\\gamma\)\-\\delta,\\tau\(\\gamma\)\+\\delta\]\\cap\[0,1\], letting𝒬′=⟨𝒜,ℛ−,ℛ\+,τ′⟩\\mathcal\{Q\}^\{\\prime\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau^\{\\prime\}\\right\\ranglebe the QBAF such thatτ′​\(γ\)=e\\tau^\{\\prime\}\(\\gamma\)=eandτ′​\(η\)=τ​\(η\)\\tau^\{\\prime\}\(\\eta\)=\\tau\(\\eta\)for allη∈𝒜∖\{γ\}\\eta\\in\\mathcal\{A\}\\setminus\\\{\\gamma\\\}\. the following statements hold: 1\. IfΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0, thenσ𝒬​\(α\)−σ𝒬​\(β\)≤σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\leq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere<τ⁡\(γ\)e<\\tau\(\\gamma\), andσ𝒬​\(α\)−σ𝒬​\(β\)≥σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\geq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere\>τ⁡\(γ\)e\>\\tau\(\\gamma\); 2\. IfΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0, thenσ𝒬​\(α\)−σ𝒬​\(β\)≥σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\geq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere<τ⁡\(γ\)e<\\tau\(\\gamma\), andσ𝒬​\(α\)−σ𝒬​\(β\)≤σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\leq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere\>τ⁡\(γ\)e\>\\tau\(\\gamma\)\.

###### Proposition 11\.

Φα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)satisfies Local Faithfulness, whileΦα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)andΦα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)can violate Local Faithfulness\.

*Cross\-topic\-adjusted Efficiency*states that the overall contrastive effect should be fully explained by two components: the aggregate attribution of the non\-topic arguments and a correction capturing the imbalance between the attributions of the two topic arguments to each other\. The sum of these components should correspond to the difference between the final\-strength gap and the base\-score gap of the two topic arguments\. When the cross\-topic attributions are equal, the correction vanishes and the property reduces to standard Efficiency\([Shapley 1951](https://arxiv.org/html/2609.02399#bib.bib1)\)\.

###### Property 8\(Cross\-topic\-adjusted Efficiency\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Cross\-topic\-adjusted Efficiency*iff∑γ∈𝒜∖\{α,β\}Φα⪰β​\(γ\)\+\(ϕα​\(β\)−ϕβ​\(α\)\)=\(σ𝒬​\(α\)−σ𝒬​\(β\)\)−\(τ⁡\(α\)−τ⁡\(β\)\)\\sum\_\{\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}\}\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\+\\bigl\(\\phi\_\{\\alpha\}\(\\beta\)\-\\phi\_\{\\beta\}\(\\alpha\)\\bigr\)=\\bigl\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\bigr\)\-\\bigl\(\\tau\(\\alpha\)\-\\tau\(\\beta\)\\bigr\)\.

###### Proposition 12\.

Φα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)satisfies Cross\-topic\-adjusted Efficiency, whileΦα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)andΦα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)can violate it\.

### Discussion

While we do not regard Counterfactuality, Local Faithfulness and Cross\-topic\-adjusted Efficiency as essential, we showed that they allow us to distinguish our CAFs and can therefore guide the choice of the CAF in practice\. Intuitively, the removal\-based CAF measures how the strength difference changes when an argument is removed and is therefore appropriate when a counterfactual interpretation of attribution values is desirable\. The gradient\-based CAF captures the local sensitivity of the strength difference to changes in arguments’ base scores, and can be used when we are interested in the robustness to small changes\. Finally, the Shapley\-based CAF measures an argument’s average marginal contribution and is preferable when the difference should be fully explained by the attribution values and the cross\-topic component\.

## 8CAFs for Healthcare

Figure 2:A QBAF for treatment recommendation in healthcare \(taken from\([Yin et al\. 2026a](https://arxiv.org/html/2609.02399#bib.bib13)\)\)\. Blue nodes denote arguments and green/red edges, resp\., support/attack relations\.![Refer to caption](https://arxiv.org/html/2609.02399v1/figures/QE_SHAP_CaseStudy1.png)Figure 3:Shapley\-based contrastive and individual explanations for treatment selection\. Green bars indicate positive influence, while red bars indicate negative influence\.We now present an illustrative example of CAFs in a healthcare setting\. Figure[2](https://arxiv.org/html/2609.02399#S8.F2)\(taken from\([Yin et al\. 2026a](https://arxiv.org/html/2609.02399#bib.bib13)\)\) shows how a QBAF can support practitioners in selecting suitable treatments for a patient based on observed symptoms111This is a simplified, non\-clinically validated example\.\. This QBAF consists of three layers\. The bottom layer represents observable symptoms \(e\.g\.,*cough*,*wheezing*\)\. The intermediate layer represents possible diagnoses \(e\.g\.,*viral/bacterial pneumonia*\)\. The top layer represents possible treatments \(e\.g\.,*antiviral/antibiotic therapy*\)\. These layers are connected hierarchically via*attack*and*support*relations, which capture relationships between symptoms, diagnoses, and treatments\. For example,*wheezing*is commonly associated with*asthma*\(support\)\.*Asthma*can be treated by*bronchodilator therapy*\(support\) but not by*antiviral therapy*\(attack\)\. Treatment arguments are connected by mutual attacks because they correspond to competing treatment choices\.

Base scores can reflect prior information such as the severity of observed symptoms, the plausibility of diagnoses based on a patient’s medical history, or the suitability of candidate treatments\. For instance, a higher body temperature may induce a higher base score for the*high fever*argument\. In this example, we assign a base score of0\.50\.5to all arguments\.

After specifying the QBAF structure and base scores, we compute the strengths of arguments under QE semantics\([Potyka 2018](https://arxiv.org/html/2609.02399#bib.bib16)\), following the setting of\([Yin et al\. 2026a](https://arxiv.org/html/2609.02399#bib.bib13)\)\. Since the symptom arguments have neither attackers nor supporters, their strengths remain equal to their base scores, namely0\.50\.5\. The strengths of the diagnosis arguments are0\.750\.75for*viral pneumonia*,0\.60\.6for*bacterial pneumonia*, and0\.40\.4for*asthma*\. The strengths of the treatment arguments are0\.35640\.3564for*antiviral therapy*,0\.23680\.2368for*antibiotic therapy*, and0\.14790\.1479for*bronchodilator therapy*\. Thus,*antiviral therapy*is selected as the most suitable treatment, as it has the highest strength among the candidate treatments\.

The strength ranking identifies the recommended treatment, but does not explain which symptoms are responsible for this choice\. For example, a patient may ask why*antiviral therapy*is recommended instead of*antibiotic therapy*\. This is a contrastive question and we can use Shapley\-based CAFs to compute an answer\. We present gradient\-based explanations in Section[B](https://arxiv.org/html/2609.02399#A2)of the Supplementary Material\.

### Contrastive Explanations\.

Figure[3](https://arxiv.org/html/2609.02399#S8.F3)\(a\) shows the Shapley\-based contrastive explanations for our example\. We first analyse the attributions for the five symptom arguments\. Among them,*Positive PCR Test*\(abbreviated as*PCR*\) has the largest positive influence on the contrast of the selected treatment \(*antiviral therapy*\) over the alternative \(*antibiotic therapy*\)\. The transparent QBAF structure allows us to visualise this influence through paths connecting*PCR*to the two treatment arguments\. Through*viral pneumonia*,*PCR*supports a diagnosis that supports*antiviral therapy*and attacks*antibiotic therapy*\. Thus, these paths both strengthen the selected treatment and weaken the alternative\. Furthermore, through*bacterial pneumonia*,*PCR*attacks a diagnosis that supports*antibiotic therapy*and attacks*antiviral therapy*\. The paths through*Asthma*are less discriminative for this contrast, since weakening*Asthma*benefits both treatments\.

We next analyse the influence of arguments in the diagnosis layer\.*Viral*and*bacterial pneumonia*have the strongest positive and negative influence on the contrast, respectively\. This is intuitive because*viral pneumonia*directly supports the selected treatment and attacks the alternative; whereas*bacterial pneumonia*directly attacks the selected treatment and supports the alternative\.

### Individual Explanations\.

To clarify what we gain by using CAFs, we compare them to individual AFs\. This comparison highlights several notable differences\. First, an argument that has equal or similar influences on two individual treatments may become much less influential in the contrastive explanation\. For example,*asthma*has a negative influence on both*antiviral therapy*and*antibiotic therapy*in the individual explanations\. However, in the contrastive explanation, its attribution score is close to00\. This indicates that although*asthma*negatively influences the selected treatment, it is not a distinguishing argument in favour of*antiviral therapy*over*antibiotic therapy*because it also negatively influences the alternative in a symmetric manner\. Second, an argument that is not very influential on the selected treatment in the individual explanation may become influential in the contrastive explanation\. For instance,*elevated WBC count*has little influence on*antiviral therapy*, but it has the largest negative attribution score in the contrastive explanation due to its strong positive influence on*antibiotic therapy*\.

### Property Illustration\.

We illustrate*cross\-topic\-adjusted efficiency*, a key property uniquely guaranteed by Shapley\-based CAFs\. For the contrastive explanation in Figure[3](https://arxiv.org/html/2609.02399#S8.F3)\(a\), the final strength difference between*antiviral*and*antibiotic therapy*arguments is0\.3564−0\.2368=0\.11960\.3564\-0\.2368=0\.1196, while the base score difference is0\.5−0\.5=00\.5\-0\.5=0\. The non\-topic contrastive attributions sum to0\.11230\.1123, while the cross\-topic component is−0\.0705−\(−0\.0779\)=0\.0074\-0\.0705\-\(\-0\.0779\)=0\.0074\. Using unrounded values, their adjusted sum is0\.11960\.1196, as required\. For the individual explanations, the attribution scores sum to−0\.1436\-0\.1436for*antiviral therapy*in Figure[3](https://arxiv.org/html/2609.02399#S8.F3)\(b\) and−0\.2632\-0\.2632for*antibiotic therapy*in Figure[3](https://arxiv.org/html/2609.02399#S8.F3)\(c\)\. These sums match the changes from the empty coalition to the full player set:0\.5−0\.1436=0\.35640\.5\-0\.1436=0\.3564and0\.5−0\.2632=0\.23680\.5\-0\.2632=0\.2368\.

## 9CAFs for Bias Detection

In this section, we demonstrate how contrastive explanations can help uncover biases in classification tasks\([Jacovi et al\. 2021](https://arxiv.org/html/2609.02399#bib.bib32)\)\. We use gradient\-based CAFs to detect bias in multilayer perceptrons \(MLPs\), as analysing local sensitivity through gradient\-based methods is a natural and widely used approach to explaining differentiable neural networks \(e\.g\., Integrated Gradients\([Sundararajan et al\. 2017](https://arxiv.org/html/2609.02399#bib.bib33)\)\)\. In addition, gradient\-based CAFs are computationally efficient, as they avoid the combinatorial calculations required by Shapley\-based CAFs\.

Dataset Selection and Preprocessing\.We use the COMPAS dataset\([ProPublica 2016](https://arxiv.org/html/2609.02399#bib.bib34);[Angwin et al\. 2016](https://arxiv.org/html/2609.02399#bib.bib35)\), which is widely used in algorithmic fairness research\. The dataset contains defendants’ demographic and criminal\-history information and provides assessments of their recidivism risk\. After data cleaning and restricting the dataset toAfrican\-AmericanandCaucasianindividuals,5,2785,278instances remain\. Each instance is represented by eight input features: age, counts of juvenile felonies, misdemeanours, and other offences, number of prior offences, current charge degree, sex, and race\. Thescore\_textfield is used as the classification label and defines three risk categories:Low,Medium, andHigh\.

To train a classifier with a known bias, we modify the labels before splitting the dataset\. Specifically, we randomly relabel 70% of theAfrican\-Americaninstances in theLowcategory asMedium, and around 70% of those in theMediumcategory asHigh, while leaving the labels of allCaucasianinstances unchanged\. This controlled modification enables the classifier to learn a race\-related bias\.

Classifier Architecture and Training\.We train an MLP with one hidden layer containing 16 neurons on the modified dataset\. The training/validation/test split was 64/16/20 %\. The trained classifier achieves an accuracy of 64\.58% on the test set\. This is very low, but since we are only interested in explaining what a classifier learnt, it does not affect our experiments\.

![Refer to caption](https://arxiv.org/html/2609.02399v1/figures/NN_bias.png)Figure 4:Gradient\-based individual and contrastive AAEs for a defendant classified asMediumrisk\. Theraceargument is the least prominent in the individual explanation but becomes the most prominent when explaining why the defendant is classified asMediumrather thanLowrisk\.Explanations\.We select anAfrican\-Americaninstance from the test set that is correctly classified asMedium\. When its race is changed toCaucasianwhile all other input features are held fixed, the MLP’s prediction changes fromMediumtoLow\. Following the correspondence established in\([Potyka 2021](https://arxiv.org/html/2609.02399#bib.bib31)\), we represent the MLP as a QBAF in which neurons are represented as arguments and weighted neural connections as support or attack relations \. In particular, the eight input neurons correspond to eight input arguments, while the three output neurons correspond to the output argumentsLow,Medium, andHigh\. We then apply gradient\-based AAEs to quantify the attribution of each input argument to the strength of the output argumentMedium\.

As shown in Figure[4](https://arxiv.org/html/2609.02399#S9.F4)\(a\), the gradient\-based attribution from the input argumentraceto the output argumentMediumis only−0\.0175\-0\.0175, the smallest in absolute magnitude among the eight input arguments\. In contrast, when explaining why the prediction isMediumrather thanLow, the contrastive gradient\-based attribution ofraceis0\.27640\.2764and becomes the largest among all input arguments, as shown in Figure[4](https://arxiv.org/html/2609.02399#S9.F4)\(b\)\. This indicates thatracesubstantially increases the model’s preference forMediumoverLow, although this race\-related bias is not prominent in the individual AAEs\. Thus, in this example, contrastive explanations reveal the controlled race\-related bias in the MLP’s prediction\.

To investigate whether this observation extends beyond our illustrative instance, we analysed141141instances that are correctly classified asMediumin the test set, but whose predictions change toLowwhen race alone is changed fromAfrican\-AmericantoCaucasian\. For each instance, the eight input arguments are ranked according to the absolute values of their gradient\-based attributions\. The mean importance rank of theraceargument improves from5\.165\.16under the individual AAE to2\.622\.62under the contrastive AAE \(with mean absolute attributions of0\.14240\.1424and0\.44780\.4478, respectively\)\. Note that a rank closer to11indicates a larger absolute attribution and hence greater importance\. Hence, contrastive CAFs make the influence of theraceargument more prominent than AFs\.

## 10Related Work

Our work is related to\([Kampik et al\. 2024a](https://arxiv.org/html/2609.02399#bib.bib8)\), which considers a QBAF𝒬\\mathcal\{Q\}and its updated version𝒬′\\mathcal\{Q\}^\{\\prime\}, where the update may involve adding arguments, support/attack relations, or changing base scores of arguments\. For any two topic argumentsα\\alphaandβ\\betain both𝒬\\mathcal\{Q\}and𝒬′\\mathcal\{Q\}^\{\\prime\}, a*strength inconsistency*occurs when the relative ordering of their strengths is not preserved after the update \(e\.g\.,σ𝒬​\(α\)\>σ𝒬​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\>\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)butσ𝒬′​\(α\)<σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)<\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\)\. They define three types of argument\-set\-based explanations:*sufficient explanations*, whose changes alone can cause the inconsistency;*counterfactual explanations*, which are sufficient explanations whose reversal restores strength consistency; and*necessary explanations*, which intersect all sufficient explanations and therefore capture sets of changes that cannot all be avoided\. In contrast, our work focuses on explaining the relative strength difference between two topic arguments within a single QBAF, rather than explaining strength inconsistency induced by an update from one QBAF to another\.

\([Kampik et al\. 2026](https://arxiv.org/html/2609.02399#bib.bib25)\)consider a QBAF with multiple topic arguments and studies how to obtain a different, often desirable, ordering over their strengths by modifying the base scores of arguments in the QBAF\. This is different from our work in that we focus on explaining the strength difference between two topic arguments, rather than obtaining a different strength ordering\.

Our work is related to a growing line of research on explaining the strength of an individual topic argumentα∈𝒜\\alpha\\in\\mathcal\{A\}in QBAFs and their edge\-weighted variants\. Existing methods can be broadly divided into attribution\-based and counterfactual approaches\. Attribution\-based methods aim to explainσ𝒬​\(α\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)by measuring the influence of other arguments onα\\alpha\([Delobelle and Villata 2019](https://arxiv.org/html/2609.02399#bib.bib20);[Čyras et al\. 2022](https://arxiv.org/html/2609.02399#bib.bib21);[Kampik et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib23);[Yin et al\. 2023](https://arxiv.org/html/2609.02399#bib.bib9);[Anaissy et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib15);[Naudot et al\. 2026](https://arxiv.org/html/2609.02399#bib.bib19)\), or the influence of relations onα\\alpha\([Amgoud et al\. 2017](https://arxiv.org/html/2609.02399#bib.bib22);[Yin et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib10)\)\. By contrast, counterfactual explanations study how a desired strength value forα\\alphacan be obtained, for instance by modifying argument base scores\([Yin et al\. 2024a](https://arxiv.org/html/2609.02399#bib.bib11)\)or, in edge\-weighted QBAFs, by modifying edge weights\([Yin et al\. 2026b](https://arxiv.org/html/2609.02399#bib.bib14)\)\. These works focus on explaining a single topic argument, whereas we focus on explaining the difference between two topic arguments\.

## 11Conclusions

We introduced contrastive explanations for QBAFs to explain differences in the strengths of two topic arguments\. As we saw, anti\-symmetric CAFs that are calibrated with respect to an AF can be derived from the AF via subtraction\. Furthermore, we showed that if a CAF satisfies anti\-symmetry, additivity and is calibrated with respect to a plausible AF, it must be equivalent to the CAF derived from the AF via subtraction\. Additivity also allowed us to design a dynamic programming algorithm that requires only a linear number of CAF calls when computing attribution values for all combinations of topic arguments\. We considered three instanstiations based on removal, gradients and Shapley values and separated them by their distinctive properties\. A case study in healthcare illustrates the practical applicability of our approach, and experiments demonstrate that contrastive explanations can help uncover bias in classifiers\.

As future work, we plan to extend our approach to explain rankings among topic arguments\. We also intend to conduct user studies to assess whether contrastive explanations improve users’ understanding of the mechanisms underlying QBAFs\. Finally, we aim to explore further application domains in which contrastive explanations may be beneficial, e\.g\. in product recommendation\([Rago et al\. 2025](https://arxiv.org/html/2609.02399#bib.bib7)\)where contrastive explanations seem like a natural fit\.

## References

- Amgoudet al\.\(2017\)L\. Amgoud, J\. Ben\-Naim, and S\. VesicMeasuring the intensity of attacks in argumentation graphs with shapley value\.InInternational Joint Conference on Artificial Intelligence \(IJCAI\),pp\. 63–69\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Amgoud and Ben\-Naim \(2018\)L\. Amgoud and J\. Ben\-NaimEvaluation of arguments in weighted bipolar graphs\.International Journal of Approximate Reasoning99,pp\. 39–55\.Cited by:[§2](https://arxiv.org/html/2609.02399#S2.p3.1)\.
- Anaissyet al\.\(2025\)C\. A\. Anaissy, J\. Delobelle, S\. Vesic, and B\. YunImpact measures for gradual argumentation semantics\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Angwinet al\.\(2016\)J\. Angwin, J\. Larson, S\. Mattu, and L\. KirchnerMachine bias: there’s software used across the country to predict future criminals\. and it’s biased against blacks\.Note:ProPublicaAccessed: 2025\-05\-06External Links:[Link](https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing)Cited by:[§9](https://arxiv.org/html/2609.02399#S9.p2.1)\.
- Ayoobiet al\.\(2025\)H\. Ayoobi, N\. Potyka, and F\. ToniProtoArgNet: interpretable image classification with super\-prototypes and argumentation\.InAAAI Conference on Artificial Intelligence \(AAAI\),pp\. 1791–1799\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.
- Baroniet al\.\(2015\)P\. Baroni, M\. Romano, F\. Toni, M\. Aurisicchio, and G\. BertanzaAutomatic evaluation of design alternatives with quantitative argumentation\.Argument & Computation6,pp\. 24–49\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1),[§2](https://arxiv.org/html/2609.02399#S2.p1.1)\.
- Chin\-Parker and Cantelon \(2017\)S\. Chin\-Parker and J\. CantelonContrastive constraints guide explanation\-based category learning\.Cognitive science41\(6\),pp\. 1645–1655\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Cocarascuet al\.\(2019\)O\. Cocarascu, A\. Rago, and F\. ToniExtracting dialogical explanations for review aggregations with argumentative dialogical agents\.In18th International Conference on Autonomous Agents and MultiAgent Systems \(AAMAS\),Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.
- Čyraset al\.\(2022\)K\. Čyras, T\. Kampik, and Q\. WengDispute trees as explanations in quantitative \(bipolar\) argumentation\.InArgXAI 2022, 1st International Workshop on Argumentation for eXplainable AI, Cardiff, Wales, September 12, 2022,Vol\.3209\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Delobelle and Villata \(2019\)J\. Delobelle and S\. VillataInterpretability of gradual semantics in abstract argumentation\.InEuropean Conference on Symbolic and Quantitative Approaches with Uncertainty,pp\. 27–38\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Freedmanet al\.\(2025\)G\. Freedman, A\. Dejl, D\. Gorur, X\. Yin, A\. Rago, and F\. ToniArgumentative large language models for explainable and contestable claim verification\.InAAAI Conference on Artificial Intelligence \(AAAI\),T\. Walsh, J\. Shah, and Z\. Kolter \(Eds\.\),pp\. 14930–14939\.External Links:[Link](https://doi.org/10.1609/aaai.v39i14.33637),[Document](https://dx.doi.org/10.1609/AAAI.V39I14.33637)Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.
- Hilton \(1990\)D\. J\. HiltonConversational processes and causal explanation\.\.Psychological Bulletin107\(1\),pp\. 65\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Jacoviet al\.\(2021\)A\. Jacovi, S\. Swayamdipta, S\. Ravfogel, Y\. Elazar, Y\. Choi, and Y\. GoldbergContrastive explanations for model interpretability\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 1597–1611\.Cited by:[§9](https://arxiv.org/html/2609.02399#S9.p1.1)\.
- Kampiket al\.\(2024a\)T\. Kampik, K\. Čyras, and J\. R\. AlarcónChange in quantitative bipolar argumentation: sufficient, necessary, and counterfactual explanations\.International Journal of Approximate Reasoning164,pp\. 109066\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p1.1)\.
- Kampiket al\.\(2024b\)T\. Kampik, N\. Potyka, X\. Yin, K\. Čyras, and F\. ToniContribution functions for quantitative bipolar argumentation graphs: a principle\-based analysis\.International Journal of Approximate Reasoning173,pp\. 109255\.Cited by:[Appendix A](https://arxiv.org/html/2609.02399#A1.p30.1.1),[§10](https://arxiv.org/html/2609.02399#S10.p3.1),[§2](https://arxiv.org/html/2609.02399#S2.p2.1),[§2](https://arxiv.org/html/2609.02399#S2.p6.1),[§7](https://arxiv.org/html/2609.02399#S7.p1.1)\.
- Kampiket al\.\(2026\)T\. Kampik, X\. Yin, N\. Potyka, and F\. ToniStrength change explanations in quantitative argumentation\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p2.1)\.
- Koriet al\.\(2025\)A\. Kori, A\. Rago, and F\. ToniFree argumentative exchanges for explaining image classifiers\.InInternational Conference on Autonomous Agents and Multiagent Systems \(AAMAS\),S\. Das, A\. Nowé, and Y\. Vorobeychik \(Eds\.\),pp\. 1172–1180\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.
- Lipton \(1990\)P\. LiptonContrastive explanation\.Royal Institute of Philosophy Supplements27,pp\. 247–266\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Miller \(2019\)T\. MillerExplanation in artificial intelligence: insights from the social sciences\.Artificial intelligence267,pp\. 1–38\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Miller \(2021\)T\. MillerContrastive explanation: a structural\-model approach\.The Knowledge Engineering Review36,pp\. e14\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Mossakowski and Neuhaus \(2018\)T\. Mossakowski and F\. NeuhausModular semantics and characteristics for bipolar weighted argumentation graphs\.CoRRabs/1807\.06685\.External Links:[Link](http://arxiv.org/abs/1807.06685),1807\.06685Cited by:[§2](https://arxiv.org/html/2609.02399#S2.p4.1),[§2](https://arxiv.org/html/2609.02399#S2.p5.1)\.
- Naudotet al\.\(2026\)F\. Naudot, A\. Brännström, V\. Torra, and T\. KampikSet contribution functions for quantitative bipolar argumentation and their principles\.International Journal of Approximate Reasoning194,pp\. 109673\.External Links:ISSN 0888\-613X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.ijar.2026.109673),[Link](https://www.sciencedirect.com/science/article/pii/S0888613X26000496)Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Potyka and Booth \(2024a\)N\. Potyka and R\. BoothAn empirical study of quantitative bipolar argumentation frameworks for truth discovery\.InComputational Models of Argument \(COMMA\),C\. Reed, M\. Thimm, and T\. Rienstra \(Eds\.\),Frontiers in Artificial Intelligence and Applications, Vol\.388,pp\. 205–216\.External Links:[Link](https://doi.org/10.3233/FAIA240322),[Document](https://dx.doi.org/10.3233/FAIA240322)Cited by:[§2](https://arxiv.org/html/2609.02399#S2.p5.1)\.
- Potyka and Booth \(2024b\)N\. Potyka and R\. BoothBalancing open\-mindedness and conservativeness in quantitative bipolar argumentation \(and how to prove semantical from functional properties\)\.InInternational Conference on Principles of Knowledge Representation and Reasoning \(KR\),P\. Marquis, M\. Ortiz, and M\. Pagnucco \(Eds\.\),External Links:[Document](https://dx.doi.org/10.24963/KR.2024/56)Cited by:[Appendix A](https://arxiv.org/html/2609.02399#A1.p11.1.1),[Appendix A](https://arxiv.org/html/2609.02399#A1.p9.1.1),[§5](https://arxiv.org/html/2609.02399#S5.p4.1)\.
- Potyka \(2018\)N\. PotykaContinuous dynamical systems for weighted bipolar argumentation\.In16th International Conference on Principles of Knowledge Representation and Reasoning \(KR\),Cited by:[§2](https://arxiv.org/html/2609.02399#S2.p3.1),[§8](https://arxiv.org/html/2609.02399#S8.p3.1)\.
- Potyka \(2019\)N\. PotykaExtending modular semantics for bipolar weighted argumentation\.InInternational Conference on Autonomous Agents and MultiAgent Systems \(AAMAS\),E\. Elkind, M\. Veloso, N\. Agmon, and M\. E\. Taylor \(Eds\.\),pp\. 1722–1730\.Cited by:[§2](https://arxiv.org/html/2609.02399#S2.p5.1)\.
- Potyka \(2021\)N\. PotykaInterpreting neural networks as quantitative argumentation frameworks\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.35,pp\. 6463–6470\.Cited by:[§9](https://arxiv.org/html/2609.02399#S9.p5.1)\.
- ProPublica \(2016\)ProPublicaCOMPAS Recidivism Racial Bias Dataset\.Note:https://www\.kaggle\.com/datasets/danofer/compassAccessed: 2025\-05\-06Cited by:[§9](https://arxiv.org/html/2609.02399#S9.p2.1)\.
- Ragoet al\.\(2021\)A\. Rago, O\. Cocarascu, C\. Bechlivanidis, D\. A\. Lagnado, and F\. ToniArgumentative explanations for interactive recommendations\.Artif\. Intell\.296,pp\. 103506\.External Links:[Link](https://doi.org/10.1016/j.artint.2021.103506),[Document](https://dx.doi.org/10.1016/J.ARTINT.2021.103506)Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.
- Ragoet al\.\(2025\)A\. Rago, O\. Cocarascu, J\. Oksanen, and F\. ToniArgumentative review aggregation and dialogical explanations\.Artif\. Intell\.340,pp\. 104291\.External Links:[Link](https://doi.org/10.1016/j.artint.2025.104291),[Document](https://dx.doi.org/10.1016/J.ARTINT.2025.104291)Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1),[§11](https://arxiv.org/html/2609.02399#S11.p2.1)\.
- Ragoet al\.\(2016\)A\. Rago, F\. Toni, M\. Aurisicchio, and P\. BaroniDiscontinuity\-free decision support with quantitative argumentation debates\.In15th International Conference on the Principles of Knowledge Representation and Reasoning \(KR\),Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1),[§2](https://arxiv.org/html/2609.02399#S2.p3.1)\.
- Shapley \(1951\)L\. S\. ShapleyNotes on the n\-person game\.Cited by:[§7](https://arxiv.org/html/2609.02399#S7.p4.1)\.
- Sundararajanet al\.\(2017\)M\. Sundararajan, A\. Taly, and Q\. YanAxiomatic attribution for deep networks\.InInternational conference on machine learning,pp\. 3319–3328\.Cited by:[§9](https://arxiv.org/html/2609.02399#S9.p1.1)\.
- Van Bouwel and Weber \(2002\)J\. Van Bouwel and E\. WeberRemote causes, bad explanations?\.Journal for the Theory of Social Behaviour32\(4\)\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Yinet al\.\(2026a\)X\. Yin, T\. Miller, N\. Potyka, A\. Rago, and F\. ToniTowards an argumentative foundation for evaluative ai\.InWorkshop on Explainable Artificial Intelligence \(XAI\) at IJCAI \(To appear\),Cited by:[Figure 2](https://arxiv.org/html/2609.02399#S8.F2),[§8](https://arxiv.org/html/2609.02399#S8.p1.1),[§8](https://arxiv.org/html/2609.02399#S8.p3.1)\.
- Yinet al\.\(2026b\)X\. Yin, N\. Potyka, A\. Rago, T\. Kampik, and F\. ToniContestability in Edge\-Weighted Quantitative Bipolar Argumentation Frameworks\.InProceedings of the 23rd International Conference on Principles of Knowledge Representation and Reasoning,pp\. 676–687\.External Links:[Document](https://dx.doi.org/10.24963/kr.2026/64),[Link](https://doi.org/10.24963/kr.2026/64)Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Yinet al\.\(2023\)X\. Yin, N\. Potyka, and F\. ToniArgument attribution explanations in quantitative bipolar argumentation frameworks\.InECAI 2023,pp\. 2898–2905\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p3.1),[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Yinet al\.\(2024a\)X\. Yin, N\. Potyka, and F\. ToniCE\-qarg: counterfactual explanations for quantitative bipolar argumentation frameworks\.InProceedings of the International Conference on Principles of Knowledge Representation and Reasoning,Vol\.21,pp\. 697–707\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p3.1),[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Yinet al\.\(2024b\)X\. Yin, N\. Potyka, and F\. ToniExplaining arguments’ strength: unveiling the role of attacks and supports\.InProceedings of the Thirty\-Third International Joint Conference on Artificial Intelligence,pp\. 3622–3630\.Cited by:[§10](https://arxiv.org/html/2609.02399#S10.p3.1)\.
- Ylikoski \(2007\)P\. YlikoskiThe idea of contrastive explanandum\.InRethinking explanation,pp\. 27–42\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p2.1)\.
- Zhuet al\.\(2025\)Y\. Zhu, N\. Potyka, D\. Hernández, Y\. He, Z\. Ding, B\. Xiong, D\. Zhou, E\. Kharlamov, and S\. StaabArgRAG: explainable retrieval augmented generation using quantitative bipolar argumentation\.InInternational Conference on Neurosymbolic Learning and Reasoning \(NeSy\),Proceedings of Machine Learning Research, Vol\.284,pp\. 697–718\.Cited by:[§1](https://arxiv.org/html/2609.02399#S1.p1.1)\.

## Supplementary Material for “Contrastive Explanations in Quantitative Bipolar Argumentation Frameworks”

## Appendix AProofs

###### Property 1\(Antisymmetry\)\.

For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\},Φα⪰β​\(γ\)=−Φβ⪰α​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)\.

###### Proposition 1\.

IfΦ\\Phisatisfies antisymmetry, thenΦα⪰α​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\)\\\!\\\!=\\\!\\\!0\.

###### Proof\.

Antisymmetry impliesΦα⪰α​\(γ\)=−Φα⪰α​\(γ\)\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\)=\-\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\), so2​Φα⪰α​\(γ\)=02\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\)=0\. Hence,Φα⪰α​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\alpha\}\(\\gamma\)=0\. ∎

###### Property 2\(ϕ\\phi\-Calibration\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, ifϕα​\(γ\)=a\\phi\_\{\\alpha\}\(\\gamma\)=aandϕβ​\(γ\)=0\\phi\_\{\\beta\}\(\\gamma\)=0, thenΦα⪰β​\(γ\)=a\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=a\.

###### Proposition 2\(Inverseϕ\\phi\-Calibration\)\.

IfΦ\\Phisatisfies antisymmetry andϕ\\phi\-Calibration, thenϕα​\(γ\)=0\\phi\_\{\\alpha\}\(\\gamma\)=0andϕβ​\(γ\)=b\\phi\_\{\\beta\}\(\\gamma\)=bimpliesΦα⪰β​\(γ\)=−b\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-b\.

###### Proof\.

We haveΦα⪰β​\(γ\)=−Φβ⪰α​\(γ\)=−b,\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)=\-b,where we used antisymmetry for the first equality andϕ\\phi\-Calibration with the assumptionϕβ​\(γ\)=b\\phi\_\{\\beta\}\(\\gamma\)=bandϕα​\(γ\)=0\\phi\_\{\\alpha\}\(\\gamma\)=0for the second equality\. ∎

###### Property 3\(ϕ\\phi\-Neutrality\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}ifϕα​\(γ\)=ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)=\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)=0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=0\.

###### Property 4\(ϕ\\phi\-Monotonicity\)\.

Letϕ\\phibe an attribution function\. For anyα,β∈𝒜\\alpha,\\beta\\in\\mathcal\{A\}andγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, ifϕα​\(γ\)\>ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)\>\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0\.

###### Proposition 3\.

IfΦ\\Phisatisfies antisymmetry andϕ\\phi\-Monotonicity, thenϕα​\(γ\)<ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)<\\phi\_\{\\beta\}\(\\gamma\)impliesΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0\.

###### Proof\.

We haveΦα⪰β​\(γ\)=−Φβ⪰α​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\-\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)\.ϕ\\phi\-Monotonicity and the assumptionϕα​\(γ\)<ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)<\\phi\_\{\\beta\}\(\\gamma\)impliesΦβ⪰α​\(γ\)\>0\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)\>0, thusΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0\. ∎

###### Proposition 4\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived from an attribution functionϕ\\phiusing subtraction, thenΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies Antisymmetry,ϕ\\phi\-Calibration,ϕ\\phi\-Neutrality andϕ\\phi\-Monotonicity\.

###### Proof\.

Antisymmetry: IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived using subtraction, thenΦα⪰β​\(γ\)=ϕα​\(γ\)−ϕβ​\(γ\)=−\(ϕβ​\(γ\)−ϕα​\(γ\)\)=−Φβ⪰α​\(γ\)\.\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)=\-\(\\phi\_\{\\beta\}\(\\gamma\)\-\\phi\_\{\\alpha\}\(\\gamma\)\)=\-\\Phi\_\{\\beta\\succeq\\alpha\}\(\\gamma\)\.

ϕ\\phi\-Calibration: Ifϕα​\(γ\)=a\\phi\_\{\\alpha\}\(\\gamma\)=aandϕβ​\(γ\)=0\\phi\_\{\\beta\}\(\\gamma\)=0, thenΦα⪰β​\(γ\)=ϕα​\(γ\)−ϕβ​\(γ\)=a−0=a\.\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)=a\-0=a\.

ϕ\\phi\-Neutrality: Ifϕα​\(γ\)=ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)=\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)=ϕα​\(γ\)−ϕβ​\(γ\)=0\.\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)=0\.

ϕ\\phi\-Monotonicity: Ifϕα​\(γ\)\>ϕβ​\(γ\)\\phi\_\{\\alpha\}\(\\gamma\)\>\\phi\_\{\\beta\}\(\\gamma\), thenΦα⪰β​\(γ\)=ϕα​\(γ\)−ϕβ​\(γ\)\>0\.\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)\>0\.∎

###### Property 5\(Additivity\)\.

For anyα,β,η∈𝒜\\alpha,\\beta,\\eta\\in\\mathcal\{A\}and anyγ∈𝒜∖\{α,β,η\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\eta\\\},Φα⪰β​\(γ\)=Φα⪰η​\(γ\)\+Φη⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\Phi\_\{\\alpha\\succeq\\eta\}\(\\gamma\)\+\\Phi\_\{\\eta\\succeq\\beta\}\(\\gamma\)\.

###### Proposition 5\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived using subtraction, thenΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies additivity\.

###### Proof\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}is derived using subtraction, thenΦα⪰η​\(γ\)\+Φη⪰β​\(γ\)=\(ϕα​\(γ\)−ϕη​\(γ\)\)\+\(ϕη​\(γ\)−ϕβ​\(γ\)\)=ϕα​\(γ\)−ϕβ​\(γ\)=Φα⪰β​\(γ\)\.\\Phi\_\{\\alpha\\succeq\\eta\}\(\\gamma\)\+\\Phi\_\{\\eta\\succeq\\beta\}\(\\gamma\)=\\big\(\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\eta\}\(\\gamma\)\\big\)\+\\big\(\\phi\_\{\\eta\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)\\big\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\)=\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\.∎

###### Proposition 6\.

ϕαR,ϕαG\\phi\_\{\\alpha\}^\{R\},\\phi\_\{\\alpha\}^\{G\}andϕαS\\phi\_\{\\alpha\}^\{S\}are plausible AFs under all modular semantics\.

###### Proof\.

As shown in\([Potyka and Booth 2024b](https://arxiv.org/html/2609.02399#bib.bib41)\), Theorem 20, all modular semantics satisfy the*independence*property\. Hence, when adding a single argumentη\\etawithout connecting it to any other arguments, the strength values of existing arguments remain unchanged\. This implies immediately that the second condition holds forϕαR,ϕαG\\phi\_\{\\alpha\}^\{R\},\\phi\_\{\\alpha\}^\{G\}andϕαS\\phi\_\{\\alpha\}^\{S\}and that the first condition holds forϕαR\\phi\_\{\\alpha\}^\{R\}andϕαG\\phi\_\{\\alpha\}^\{G\}because the strength values in the differences in the definition ofϕαR\\phi\_\{\\alpha\}^\{R\}andϕαG\\phi\_\{\\alpha\}^\{G\}remain unchanged\.

ForϕαS\\phi\_\{\\alpha\}^\{S\}, satisfaction of the first condition is not obvious because the sum∑B⊆𝒜α¯∖\{β\}w⁡\(B\)⋅cβ​\(B\)\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}w\(B\)\\cdot c\_\{\\beta\}\(B\)will now range over a large number of subsets with different weights\. Note that we can partition the subsets for the extended graph into those containingη\\etaand those that do not containη\\etaand note that their number is equal \(we have\|\{B∣B⊆𝒜α¯∖\{β\}\}\|=\|\{B∪\{η\}∣B⊆𝒜α¯∖\{β\}\}\|\|\\\{B\\mid B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\\\}\|=\|\\\{B\\cup\\\{\\eta\\\}\\mid B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\\\}\|\)\. By independence, we havecβ​\(B\)=cβ​\(B∪\{η\}\)c\_\{\\beta\}\(B\)=c\_\{\\beta\}\(B\\cup\\\{\\eta\\\}\)\. Hence, the Shapley value for the extended graph is

∑B⊆𝒜α¯∪\{η\}∖\{β\}w′​\(B\)⋅cβ​\(B\)\\displaystyle\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\cup\\\{\\eta\\\}\\setminus\\\{\\beta\\\}\}w^\{\\prime\}\(B\)\\cdot c\_\{\\beta\}\(B\)=∑B⊆𝒜α¯∖\{β\}w′​\(B\)⋅cβ​\(B\)\+∑B⊆𝒜α¯∖\{β\}w′​\(B∪\{η\}\)⋅cβ​\(B∪\{η\}\)\\displaystyle=\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}w^\{\\prime\}\(B\)\\cdot c\_\{\\beta\}\(B\)\+\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}w^\{\\prime\}\(B\\cup\\\{\\eta\\\}\)\\cdot c\_\{\\beta\}\(B\\cup\\\{\\eta\\\}\)=∑B⊆𝒜α¯∖\{β\}\(w′​\(B\)\+w′​\(B∪\{η\}\)\)⋅cβ​\(B\),\\displaystyle=\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}\(w^\{\\prime\}\(B\)\+w^\{\\prime\}\(B\\cup\\\{\\eta\\\}\)\)\\cdot c\_\{\\beta\}\(B\),wherew′​\(X\)=\|X\|\!⋅\(\|𝒜α¯\|−\|X\|\)\!\(\|𝒜α¯\|\+1\)\!w^\{\\prime\}\(X\)=\\frac\{\|X\|\!\\cdot\\left\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|X\|\\right\)\!\}\{\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\+1\)\!\}\. We have

w′​\(B\)\+w′​\(B∪\{η\}\)\\displaystyle w^\{\\prime\}\(B\)\+w^\{\\prime\}\(B\\cup\\\{\\eta\\\}\)=\|B\|\!⋅\(\|𝒜α¯\|−\|B\|\)\!\(\|𝒜α¯\|\+1\)\!\+\(\|B\|\+1\)\!⋅\(\|𝒜α¯\|−\|B\|−1\)\!\(\|𝒜α¯\|\+1\)\!\\displaystyle=\\frac\{\|B\|\!\\cdot\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\)\!\}\{\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\+1\)\!\}\+\\frac\{\(\|B\|\+1\)\!\\cdot\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\-1\)\!\}\{\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\+1\)\!\}=\|B\|\!⋅\(\|𝒜α¯\|−\|B\|−1\)\!⋅\(\(\|𝒜α¯\|−\|B\|\)\+\(\|B\|\+1\)\)\(\|𝒜α¯\|\+1\)\!\\displaystyle=\\frac\{\|B\|\!\\cdot\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\-1\)\!\\cdot\(\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\)\+\(\|B\|\+1\)\)\}\{\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\+1\)\!\}=\|B\|\!⋅\(\|𝒜α¯\|−\|B\|−1\)\!\|𝒜α¯\|\!\\displaystyle=\\frac\{\|B\|\!\\cdot\(\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\-\|B\|\-1\)\!\}\{\|\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\|\!\}=w⁡\(B\)\.\\displaystyle=w\(B\)\.Hence,∑B⊆𝒜α¯∪\{η\}∖\{β\}w′​\(B\)⋅cβ​\(B\)=∑B⊆𝒜α¯∖\{β\}w⁡\(B\)⋅cβ​\(B\)=ϕαS​\(β\),\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\cup\\\{\\eta\\\}\\setminus\\\{\\beta\\\}\}w^\{\\prime\}\(B\)\\cdot c\_\{\\beta\}\(B\)=\\sum\_\{B\\subseteq\\mathcal\{A\}\_\{\\overline\{\\alpha\}\}\\setminus\\\{\\beta\\\}\}w\(B\)\\cdot c\_\{\\beta\}\(B\)=\\phi\_\{\\alpha\}^\{S\}\(\\beta\),which completes the proof\. ∎

###### Proposition 7\.

IfΦ\\Phiis a CAF derived from a plausible attribution functionϕ\\phi, andΦ\\Phisatisfies antisymmetry,ϕ\\phi\-calibration and additivity, then, under all modular gradual semantics,Φ\\Phiis equal to the CAF derived fromϕ\\phiusing subtraction\.

###### Proof\.

Consider an arbitrary QBAF𝒬\\mathcal\{Q\}evaluated under a modular gradual semantics, and the QBAF𝒬′\\mathcal\{Q\}^\{\\prime\}resulting from𝒬\\mathcal\{Q\}by adding an isolated argumentη\\eta\. For clarity, Condition \(2\) of Plausibility is intended symmetrically: in𝒬′\\mathcal\{Q\}^\{\\prime\}, bothϕα​\(η\)=0\\phi\_\{\\alpha\}\(\\eta\)=0andϕη​\(α\)=0\\phi\_\{\\eta\}\(\\alpha\)=0hold for every argumentα\\alphafrom𝒬\\mathcal\{Q\}\. First note that modularity of the gradual semantics implies that it satisfies independence\([Potyka and Booth 2024b](https://arxiv.org/html/2609.02399#bib.bib41)\)\. Hence, addingη\\etawill not change the strength values of arguments in𝒬\\mathcal\{Q\}and by plausibility ofϕ\\phi, the attribution values underϕ\\phiwill remain unchanged andϕα​\(η\)=ϕη​\(α\)=0\\phi\_\{\\alpha\}\(\\eta\)=\\phi\_\{\\eta\}\(\\alpha\)=0for all argumentsα\\alphafrom𝒬\\mathcal\{Q\}\.

To distinguish CAF values under𝒬\\mathcal\{Q\}and𝒬′\\mathcal\{Q\}^\{\\prime\}, we writeΦ\\PhiandΦ′\\Phi^\{\\prime\}, respectively\. For all argumentsα,β,γ\\alpha,\\beta,\\gammafrom𝒬\\mathcal\{Q\}, we haveΦα⪰β​\(γ\)=f⁡\(ϕα​\(γ\),ϕβ​\(γ\)\)=Φα⪰β′​\(γ\)=Φα⪰η′​\(γ\)\+Φη⪰β′​\(γ\)=ϕα​\(γ\)−ϕβ​\(γ\),\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=f\(\\phi\_\{\\alpha\}\(\\gamma\),\\phi\_\{\\beta\}\(\\gamma\)\)=\\Phi^\{\\prime\}\_\{\\alpha\\succeq\\beta\}\(\\gamma\)=\\Phi^\{\\prime\}\_\{\\alpha\\succeq\\eta\}\(\\gamma\)\+\\Phi^\{\\prime\}\_\{\\eta\\succeq\\beta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\)\-\\phi\_\{\\beta\}\(\\gamma\),where we used the definition of derived CAFs and plausibility for the first and second equality, additivity for the third, and antisymmetry,ϕ\\phi\-calibration and Proposition[2](https://arxiv.org/html/2609.02399#Thmproposition2a)for the fourth \(sinceϕη​\(γ\)=0\\phi\_\{\\eta\}\(\\gamma\)=0,ϕ\\phi\-calibration impliesΦα⪰η′​\(γ\)=ϕα​\(γ\)\\Phi^\{\\prime\}\_\{\\alpha\\succeq\\eta\}\(\\gamma\)=\\phi\_\{\\alpha\}\(\\gamma\), and Proposition[2](https://arxiv.org/html/2609.02399#Thmproposition2a)impliesΦη⪰β′​\(γ\)=−ϕβ​\(γ\)\\Phi^\{\\prime\}\_\{\\eta\\succeq\\beta\}\(\\gamma\)=\-\\phi\_\{\\beta\}\(\\gamma\)\)\. ∎

###### Proposition 8\.

The CAFsΦα⪰βR,Φα⪰βG,Φα⪰βS\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\},\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\},\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}derived from the removal\-based, gradient\-based and Shapley\-based AFs using subtraction are defined as follows:

Φα⪰βR\(γ\)=\(σ𝒬\(α\)−σ𝒬\(β\)\)−\(σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\)\.\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)=\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)\.Φα⪰βG​\(γ\)=limϵ→0\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)ϵ\.\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{\\epsilon\}\.Φα⪰βS​\(γ\)=∑B⊆𝒜∖\{α,β,γ\}\(w⁡\(B\)⋅\(cγ→α​\(B\)−cγ→β​\(B\)\)\+w⁡\(B∪\{β\}\)⋅\(cγ→α​\(B∪\{β\}\)−cγ→β​\(B∪\{α\}\)\)\),\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\gamma\\\}\}\\big\(w\(B\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\)\)\+w\(B\\cup\\\{\\beta\\\}\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\\cup\\\{\\beta\\\}\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\\cup\\\{\\alpha\\\}\)\)\\big\),wherecγ→x\(X\)=σ𝒬↓X∪\{x,γ\}\(x\)−σ𝒬↓X∪\{x\}\(x\)c\_\{\\gamma\\rightarrow x\}\(X\)=\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{X\\cup\\\{x,\\gamma\\\}\}\}\}\(x\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{X\\cup\\\{x\\\}\}\}\}\(x\)\.

###### Proof\.

1\.Φα⪰βR\(γ\)=\(σ𝒬\(α\)−σ𝒬\(β\)\)−\(σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\)=\(σ𝒬\(α\)−σ𝒬↓𝒜′\(α\)\)−\(σ𝒬\(β\)−σ𝒬↓𝒜′\(β\)\)=ϕαR\(γ\)−ϕβR\(γ\)\.\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)=\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)=\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)=\\phi\_\{\\alpha\}^\{R\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{R\}\(\\gamma\)\.

2\. Here,𝒬′\\mathcal\{Q\}^\{\\prime\}denotes the perturbed QBAF𝒬ϵ\\mathcal\{Q\}\_\{\\epsilon\}defined in Definition[5](https://arxiv.org/html/2609.02399#Thmdefinition5)\.

Φα⪰βG​\(γ\)=limϵ→0\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)ϵ=limϵ→0σ𝒬′​\(α\)−σ𝒬​\(α\)ϵ−limϵ→0σ𝒬′​\(β\)−σ𝒬​\(β\)ϵ=ϕαG​\(γ\)−ϕβG​\(γ\),\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{\\epsilon\}=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\}\{\\epsilon\}\-\\lim\_\{\\epsilon\\to 0\}\\frac\{\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\}\{\\epsilon\}=\\phi\_\{\\alpha\}^\{G\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{G\}\(\\gamma\),where the second equality follows from linearity of limits\.

3\. We have

Φα⪰βS​\(γ\)\\displaystyle\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=ϕαS​\(γ\)−ϕβS​\(γ\)\\displaystyle=\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{S\}\(\\gamma\)=∑B⊆𝒜∖\{α,γ\}w⁡\(B\)⋅cγ→α​\(B\)−∑B⊆𝒜∖\{β,γ\}w⁡\(B\)⋅cγ→β​\(B\)\\displaystyle=\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\gamma\\\}\}w\(B\)\\cdot c\_\{\\gamma\\rightarrow\\alpha\}\(B\)\-\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\beta,\\gamma\\\}\}w\(B\)\\cdot c\_\{\\gamma\\rightarrow\\beta\}\(B\)=∑B⊆𝒜∖\{α,β,γ\}\(w⁡\(B\)⋅cγ→α​\(B\)\+w⁡\(B∪\{β\}\)⋅cγ→α​\(B∪\{β\}\)\)\\displaystyle=\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\gamma\\\}\}\\big\(w\(B\)\\cdot c\_\{\\gamma\\rightarrow\\alpha\}\(B\)\+w\(B\\cup\\\{\\beta\\\}\)\\cdot c\_\{\\gamma\\rightarrow\\alpha\}\(B\\cup\\\{\\beta\\\}\)\\big\)−∑B⊆𝒜∖\{α,β,γ\}\(w\(B\)⋅cγ→β\(B\)\+w\(B∪\{α\}\)⋅cγ→β\(B∪\{α\}\)\)\\displaystyle\\ \-\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\gamma\\\}\}\\big\(w\(B\)\\cdot c\_\{\\gamma\\rightarrow\\beta\}\(B\)\+w\(B\\cup\\\{\\alpha\\\}\)\\cdot c\_\{\\gamma\\rightarrow\\beta\}\(B\\cup\\\{\\alpha\\\}\)\\big\)=∑B⊆𝒜∖\{α,β,γ\}\(w⁡\(B\)⋅\(cγ→α​\(B\)−cγ→β​\(B\)\)\+w⁡\(B∪\{β\}\)⋅\(cγ→α​\(B∪\{β\}\)−cγ→β​\(B∪\{α\}\)\)\),\\displaystyle=\\sum\_\{B\\subseteq\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta,\\gamma\\\}\}\\big\(w\(B\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\)\)\+w\(B\\cup\\\{\\beta\\\}\)\\cdot\(c\_\{\\gamma\\rightarrow\\alpha\}\(B\\cup\\\{\\beta\\\}\)\-c\_\{\\gamma\\rightarrow\\beta\}\(B\\cup\\\{\\alpha\\\}\)\)\\big\),where, for the last equality, we used the fact that the value ofw⁡\(X\)w\(X\)depends only on the size ofXXand thereforew⁡\(B∪\{β\}\)=w⁡\(B∪\{α\}\)w\(B\\cup\\\{\\beta\\\}\)=w\(B\\cup\\\{\\alpha\\\}\)\. ∎

###### Proposition 9\.

IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}satisfies additivity, Algorithm[1](https://arxiv.org/html/2609.02399#alg1)computesΦti⪰tj​\(γ\)\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)for all1≤i<j≤T1\\leq i<j\\leq TwithO⁡\(T\)O\(T\)Φα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}\-calls\. IfΦα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}can be computed in timeO⁡\(C\)O\(C\), the overall time complexity of the algorithm isO⁡\(T⋅C\+T2\)O\(T\\cdot C\+T^\{2\}\)\.

###### Proof\.

For the number of CAF calls, note that the algorithm usesT−1=O⁡\(T\)T\-1=O\(T\)Φα⪰β\\Phi\_\{\\alpha\\succeq\\beta\}\-calls in the for\-loop from line 2 to 4 and at no other place\.

To see that all values are computed correctly, think of the values as organised in aT×TT\\times Tmatrix\. In lines 2 to 4, we initialise the super diagonal\. We have to assure that the matrix values are correctly initialised in lines 5 to 9\. We letM⁡\[i,j\]=M⁡\[i,j−1\]\+M⁡\[j−1,j\]M\[i,j\]=M\[i,j\-1\]\+M\[j\-1,j\]\. Hence, ifM⁡\[i,j−1\]=Φti⪰tj−1​\(γ\)M\[i,j\-1\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{j\-1\}\}\(\\gamma\)andM⁡\[j−1,j\]=Φtj−1⪰tj​\(γ\)M\[j\-1,j\]=\\Phi\_\{t\_\{j\-1\}\\succeq t\_\{j\}\}\(\\gamma\), Additivity implies thatM⁡\[i,j\]=Φti⪰tj​\(γ\)M\[i,j\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)as desired\.M⁡\[j−1,j\]=Φtj−1⪰tj​\(γ\)M\[j\-1,j\]=\\Phi\_\{t\_\{j\-1\}\\succeq t\_\{j\}\}\(\\gamma\)follows immediately from lines 2 to 4, hence it remains to check thatM⁡\[i,j−1\]=Φti⪰tj−1​\(γ\)M\[i,j\-1\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{j\-1\}\}\(\\gamma\)\. For every1≤i≤T−21\\leq i\\leq T\-2, we prove by induction onjjthatM⁡\[i,j\]=Φti⪰tj​\(γ\)M\[i,j\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{j\}\}\(\\gamma\)\. For the induction basej=i\+2j=i\+2, the entriesM⁡\[i,i\+1\]M\[i,i\+1\]andM⁡\[i\+1,i\+2\]M\[i\+1,i\+2\]are correctly initialised in lines 2 to 4\. Hence, by Additivity,M⁡\[i,i\+2\]=M⁡\[i,i\+1\]\+M⁡\[i\+1,i\+2\]=Φti⪰ti\+2​\(γ\)M\[i,i\+2\]=M\[i,i\+1\]\+M\[i\+1,i\+2\]=\\Phi\_\{t\_\{i\}\\succeq t\_\{i\+2\}\}\(\\gamma\)\. For the induction step, assume that the claim holds for alli\+2≤j≤Ni\+2\\leq j\\leq N, whereN<TN<T\. Then, forj=N\+1j=N\+1, we computeM⁡\[i,N\+1\]=M⁡\[i,N\]\+M⁡\[N,N\+1\]M\[i,N\+1\]=M\[i,N\]\+M\[N,N\+1\]\. We already established thatM⁡\[N,N\+1\]M\[N,N\+1\]is correctly initialised in lines 2 to 4, and the induction assumption implies thatM⁡\[i,N\]M\[i,N\]is computed correctly\. Hence, Additivity implies thatM⁡\[i,N\+1\]M\[i,N\+1\]is computed correctly, which completes the correctness proof\.

For the time complexity, lines 2 to 4 run in timeO⁡\(T⋅C\)O\(T\\cdot C\)\. In lines 5 to 9, we fillT−2T\-2entries for the first row,T−3T\-3for the second and so on\. Hence, overall, we have to fill∑k=1T−2k=\(T−2\)⋅\(T−1\)2\\sum\_\{k=1\}^\{T\-2\}k=\\frac\{\(T\-2\)\\cdot\(T\-1\)\}\{2\}entries\. Each entry requires one addition, hence the overall number of operations isO⁡\(T2\)O\(T^\{2\}\)\. ∎

###### Property 6\(Counterfactuality\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Counterfactuality*iff, for anyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, letting𝒜′=𝒜∖\{γ\}\\mathcal\{A\}^\{\\prime\}=\\mathcal\{A\}\\setminus\\\{\\gamma\\\}, the following statements hold: 1\. IfΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0, thenσ𝒬\(α\)−σ𝒬\(β\)<σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)<\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\); 2\. IfΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0, thenσ𝒬\(α\)−σ𝒬\(β\)\>σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\>\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\.

###### Proposition 10\.

Φα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)satisfies Counterfactuality, whileΦα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)andΦα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)can violate Counterfactuality\.

###### Proof\.

IfΦα⪰βR​\(γ\)<0\\Phi^\{R\}\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0, then\(σ𝒬\(α\)−σ𝒬\(β\)\)−\(σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\)<0\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)<0, that is,\(σ𝒬\(α\)−σ𝒬\(β\)\)<\(σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)\)\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)<\\big\(\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)\\big\)\. The case whereΦα⪰βR​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)\>0follows analogously\.

To show that the gradient\- and Shapley\-based CAFs may violate Counterfactuality, consider the QBAF𝒬=⟨𝒜,ℛ−,ℛ\+,τ⟩\\mathcal\{Q\}=\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\\rangle, where𝒜=\{α,β,γ,η\}\\mathcal\{A\}=\\\{\\alpha,\\beta,\\gamma,\\eta\\\},ℛ\+=\{\(η,α\)\}\\mathcal\{R\}^\{\+\}=\\\{\(\\eta,\\alpha\)\\\},ℛ−=\{\(γ,η\),\(γ,β\),\(η,β\),\(β,α\)\}\\mathcal\{R\}^\{\-\}=\\\{\(\\gamma,\\eta\),\(\\gamma,\\beta\),\(\\eta,\\beta\),\(\\beta,\\alpha\)\\\}, andτ⁡\(x\)=1/2\\tau\(x\)=1/2for everyx∈𝒜x\\in\\mathcal\{A\}\. Under DF\-QuAD semantics, we obtainσ𝒬​\(α\)=17/32\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)=17/32andσ𝒬​\(β\)=3/16\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)=3/16\. Therefore,σ𝒬​\(α\)−σ𝒬​\(β\)=11/32\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)=11/32\.

Let𝒜′=𝒜∖\{γ\}\\mathcal\{A\}^\{\\prime\}=\\mathcal\{A\}\\setminus\\\{\\gamma\\\}\. After removingγ\\gamma, we obtainσ𝒬↓𝒜′\(α\)=5/8\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)=5/8andσ𝒬↓𝒜′\(β\)=1/4\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)=1/4\. Hence,σ𝒬↓𝒜′\(α\)−σ𝒬↓𝒜′\(β\)=3/8\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)=3/8\.

A direct calculation givesΦα⪰βG​\(γ\)=1/8\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=1/8\>0\. For the Shapley\-based CAF, we obtainϕαS\(γ\)=−1/32\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)=\-1/32andϕβS\(γ\)=−5/32\\phi\_\{\\beta\}^\{S\}\(\\gamma\)=\-5/32\. Therefore,Φα⪰βS​\(γ\)=ϕαS​\(γ\)−ϕβS​\(γ\)=1/8\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{S\}\(\\gamma\)=1/8\>0\. However,11/32<3/811/32<3/8\. Thus, both attribution scores are positive even though removingγ\\gammaincreases, rather than decreases, the strength difference betweenα\\alphaandβ\\beta\. Therefore,Φα⪰βG\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}andΦα⪰βS\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}may violate Counterfactuality\. ∎

###### Property 7\(Local Faithfulness\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Local Faithfulness*wrt\.σ\\sigmaiff, for anyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}, there existsδ\>0\\delta\>0such that, for alle∈\[τ⁡\(γ\)−δ,τ⁡\(γ\)\+δ\]∩\[0,1\]e\\in\[\\tau\(\\gamma\)\-\\delta,\\tau\(\\gamma\)\+\\delta\]\\cap\[0,1\], letting𝒬′=⟨𝒜,ℛ−,ℛ\+,τ′⟩\\mathcal\{Q\}^\{\\prime\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau^\{\\prime\}\\right\\ranglebe the QBAF such thatτ′​\(γ\)=e\\tau^\{\\prime\}\(\\gamma\)=eandτ′​\(η\)=τ​\(η\)\\tau^\{\\prime\}\(\\eta\)=\\tau\(\\eta\)for allη∈𝒜∖\{γ\}\\eta\\in\\mathcal\{A\}\\setminus\\\{\\gamma\\\}\. the following statements hold: 1\. IfΦα⪰β​\(γ\)<0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)<0, thenσ𝒬​\(α\)−σ𝒬​\(β\)≤σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\leq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere<τ⁡\(γ\)e<\\tau\(\\gamma\), andσ𝒬​\(α\)−σ𝒬​\(β\)≥σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\geq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere\>τ⁡\(γ\)e\>\\tau\(\\gamma\); 2\. IfΦα⪰β​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\>0, thenσ𝒬​\(α\)−σ𝒬​\(β\)≥σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\geq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere<τ⁡\(γ\)e<\\tau\(\\gamma\), andσ𝒬​\(α\)−σ𝒬​\(β\)≤σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\leq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)whenevere\>τ⁡\(γ\)e\>\\tau\(\\gamma\)\.

###### Proposition 11\.

Φα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)satisfies Local Faithfulness, whileΦα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)andΦα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)can violate Local Faithfulness\.

###### Proof\.

IfΦα⪰βG​\(γ\)=limϵ→0\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)ϵ<0\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=\\lim\_\{\\epsilon\\to 0\}\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{\\epsilon\}<0, then there existsδ\>0\\delta\>0such that for any0<\|ϵ\|<δ0<\|\\epsilon\|<\\delta, we have\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)ϵ=\(σ𝒬′​\(α\)−σ𝒬′​\(β\)\)−\(σ𝒬​\(α\)−σ𝒬​\(β\)\)e−τ⁡\(γ\)<0\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{\\epsilon\}=\\frac\{\\big\(\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\\big\)\-\\big\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\big\)\}\{e\-\\tau\(\\gamma\)\}<0\. Ife<τ⁡\(γ\)e<\\tau\(\\gamma\), thenσ𝒬​\(α\)−σ𝒬​\(β\)≤σ𝒬′​\(α\)−σ𝒬′​\(β\)\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\leq\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}^\{\\prime\}\}\(\\beta\)\. The remaining three cases follow analogously\.

To show that the removal\- and Shapley\-based CAFs may violate Local Faithfulness, consider the QBAF𝒬=⟨𝒜,ℛ−,ℛ\+,τ⟩\\mathcal\{Q\}=\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\\rangle, where𝒜=\{α,β,γ,η\}\\mathcal\{A\}=\\\{\\alpha,\\beta,\\gamma,\\eta\\\},ℛ\+=\{\(γ,η\)\}\\mathcal\{R\}^\{\+\}=\\\{\(\\gamma,\\eta\)\\\},ℛ−=\{\(γ,α\),\(γ,β\),\(η,β\),\(β,α\)\}\\mathcal\{R\}^\{\-\}=\\\{\(\\gamma,\\alpha\),\(\\gamma,\\beta\),\(\\eta,\\beta\),\(\\beta,\\alpha\)\\\}, andτ⁡\(x\)=1/2\\tau\(x\)=1/2for everyx∈𝒜x\\in\\mathcal\{A\}\. Under DF\-QuAD semantics, we obtainσ𝒬​\(α\)=15/64\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)=15/64andσ𝒬​\(β\)=1/16\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)=1/16, and henceσ𝒬​\(α\)−σ𝒬​\(β\)=11/64\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)=11/64\.

Let𝒜′=𝒜∖\{γ\}\\mathcal\{A\}^\{\\prime\}=\\mathcal\{A\}\\setminus\\\{\\gamma\\\}\. After removingγ\\gamma, we obtainσ𝒬↓𝒜′\(α\)=3/8\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\alpha\)=3/8andσ𝒬↓𝒜′\(β\)=1/4\\sigma\_\{\\mathcal\{Q\}\_\{\\downarrow\_\{\\mathcal\{A\}^\{\\prime\}\}\}\}\(\\beta\)=1/4\. Therefore,Φα⪰βR​\(γ\)=11/64−\(3/8−1/4\)=3/64\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)=11/64\-\(3/8\-1/4\)=3/64\>0\.

However, a direct calculation givesΦα⪰βG\(γ\)=−5/32<0\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)=\-5/32<0\. Hence, for every sufficiently small increasee\>τ⁡\(γ\)e\>\\tau\(\\gamma\), the strength difference betweenα\\alphaandβ\\betadecreases\. This contradicts Local Faithfulness for the positive removal\-based attributionΦα⪰βR​\(γ\)\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)\>0\.

For the Shapley\-based CAF, we obtainϕαS\(γ\)=−35/192\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)=\-35/192andϕβS\(γ\)=−7/32\\phi\_\{\\beta\}^\{S\}\(\\gamma\)=\-7/32\. Hence,Φα⪰βS​\(γ\)=ϕαS​\(γ\)−ϕβS​\(γ\)=7/192\>0\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{S\}\(\\gamma\)=7/192\>0\. Since the strength difference decreases for arbitrarily small increases in the base score ofγ\\gamma, this positive attribution also violates Local Faithfulness\. Therefore,Φα⪰βR\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}andΦα⪰βS\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}may violate Local Faithfulness\.

∎

###### Property 8\(Cross\-topic\-adjusted Efficiency\)\.

Φα⪰β​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)satisfies*Cross\-topic\-adjusted Efficiency*iff∑γ∈𝒜∖\{α,β\}Φα⪰β​\(γ\)\+\(ϕα​\(β\)−ϕβ​\(α\)\)=\(σ𝒬​\(α\)−σ𝒬​\(β\)\)−\(τ⁡\(α\)−τ⁡\(β\)\)\\sum\_\{\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}\}\\Phi\_\{\\alpha\\succeq\\beta\}\(\\gamma\)\+\\bigl\(\\phi\_\{\\alpha\}\(\\beta\)\-\\phi\_\{\\beta\}\(\\alpha\)\\bigr\)=\\bigl\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\bigr\)\-\\bigl\(\\tau\(\\alpha\)\-\\tau\(\\beta\)\\bigr\)\.

###### Proposition 12\.

Φα⪰βS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)satisfies Cross\-topic\-adjusted Efficiency, whileΦα⪰βR​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}\(\\gamma\)andΦα⪰βG​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}\(\\gamma\)can violate it\.

###### Proof\.

Since the Shapley\-based CAF is derived using subtraction, we haveΦα⪰βS​\(γ\)=ϕαS​\(γ\)−ϕβS​\(γ\)\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)=\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)\-\\phi\_\{\\beta\}^\{S\}\(\\gamma\)for everyγ∈𝒜∖\{α,β\}\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}\. Therefore,∑γ∈𝒜∖\{α,β\}Φα⪰βS​\(γ\)\+\(ϕαS​\(β\)−ϕβS​\(α\)\)=∑γ∈𝒜∖\{α\}ϕαS​\(γ\)−∑γ∈𝒜∖\{β\}ϕβS​\(γ\)\\sum\_\{\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha,\\beta\\\}\}\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}\(\\gamma\)\+\\bigl\(\\phi\_\{\\alpha\}^\{S\}\(\\beta\)\-\\phi\_\{\\beta\}^\{S\}\(\\alpha\)\\bigr\)=\\sum\_\{\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\alpha\\\}\}\\phi\_\{\\alpha\}^\{S\}\(\\gamma\)\-\\sum\_\{\\gamma\\in\\mathcal\{A\}\\setminus\\\{\\beta\\\}\}\\phi\_\{\\beta\}^\{S\}\(\\gamma\)\. By the Efficiency property of individual Shapley attributions, established in\([Kampik et al\. 2024b](https://arxiv.org/html/2609.02399#bib.bib23)\), this is equal to\(σ𝒬​\(α\)−τ⁡\(α\)\)−\(σ𝒬​\(β\)−τ⁡\(β\)\)\\bigl\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\tau\(\\alpha\)\\bigr\)\-\\bigl\(\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\-\\tau\(\\beta\)\\bigr\), which is equivalent to\(σ𝒬​\(α\)−σ𝒬​\(β\)\)−\(τ⁡\(α\)−τ⁡\(β\)\)\\bigl\(\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)\-\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)\\bigr\)\-\\bigl\(\\tau\(\\alpha\)\-\\tau\(\\beta\)\\bigr\)\. Hence,Φα⪰βS\\Phi\_\{\\alpha\\succeq\\beta\}^\{S\}satisfies Cross\-topic\-adjusted Efficiency\.

To show that the removal\- and gradient\-based CAFs may violate Cross\-topic\-adjusted Efficiency, consider the QBAF𝒬=⟨𝒜,ℛ−,ℛ\+,τ⟩\\mathcal\{Q\}=\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau\\rangle, where𝒜=\{α,β,γ,η\}\\mathcal\{A\}=\\\{\\alpha,\\beta,\\gamma,\\eta\\\},ℛ−=∅\\mathcal\{R\}^\{\-\}=\\emptyset,ℛ\+=\{\(γ,α\),\(η,α\)\}\\mathcal\{R\}^\{\+\}=\\\{\(\\gamma,\\alpha\),\(\\eta,\\alpha\)\\\}, andτ⁡\(x\)=1/2\\tau\(x\)=1/2for everyx∈𝒜x\\in\\mathcal\{A\}\. Under DF\-QuAD semantics,σ𝒬​\(α\)=7/8\\sigma\_\{\\mathcal\{Q\}\}\(\\alpha\)=7/8andσ𝒬​\(β\)=1/2\\sigma\_\{\\mathcal\{Q\}\}\(\\beta\)=1/2\. Hence, the difference between the final\-strength gap and the base\-score gap is3/83/8\.

For the removal\-based CAF,ϕαR​\(γ\)=ϕαR​\(η\)=1/8\\phi\_\{\\alpha\}^\{R\}\(\\gamma\)=\\phi\_\{\\alpha\}^\{R\}\(\\eta\)=1/8, while all corresponding attributions toβ\\betaand both cross\-topic attributions are00\. Therefore, the left\-hand side of Cross\-topic\-adjusted Efficiency is1/8\+1/8=1/41/8\+1/8=1/4, which is different from3/83/8\.

For the gradient\-based CAF,ϕαG​\(γ\)=ϕαG​\(η\)=1/4\\phi\_\{\\alpha\}^\{G\}\(\\gamma\)=\\phi\_\{\\alpha\}^\{G\}\(\\eta\)=1/4, while all corresponding attributions toβ\\betaand both cross\-topic attributions are00\. Therefore, the left\-hand side of Cross\-topic\-adjusted Efficiency is1/4\+1/4=1/21/4\+1/4=1/2, which is also different from3/83/8\. Hence,Φα⪰βR\\Phi\_\{\\alpha\\succeq\\beta\}^\{R\}andΦα⪰βG\\Phi\_\{\\alpha\\succeq\\beta\}^\{G\}may violate Cross\-topic\-adjusted Efficiency\. ∎

## Appendix BAdditional Contrastive Explanations in Section[8](https://arxiv.org/html/2609.02399#S8)

![Refer to caption](https://arxiv.org/html/2609.02399v1/figures/Gradient-based_explanations.png)Figure 5:Gradient\-based contrastive and individual explanations for treatment selection\. Green bars indicate positive influence, while red bars indicate negative influence\.
## Appendix CComputing Environment

All experiments were conducted locally on a personal laptop running Microsoft Windows 11 Home, equipped with an Intel Core Ultra 5 225H CPU \(14 cores\) and 32 GB of RAM\. All computations were performed on the CPU\. The software environment consisted of Python 3\.11\.9, PyTorch 2\.13\.0, NumPy 2\.4\.4, pandas 3\.0\.3, scikit\-learn 1\.9\.0, and Matplotlib 3\.10\.8\.

Similar Articles