Towards an Argumentative Foundation for Evaluative AI
摘要
This position paper advocates computational argumentation as a formal foundation for Evaluative AI, which supports human decision-making by presenting competing hypotheses with evidence for and against, rather than single recommendations.
arXiv:2608.07473v1 Announce Type: new
Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each. In this position paper, we advocate (computational) argumentation as a particularly suitable paradigm to provide a formal, computable foundation for forms of EAI that are explainable and contestable, setting the ground for a long-term research agenda towards distributed and human-centred EAI systems.
查看缓存全文
缓存时间: 2026/08/11 08:01
# Towards an Argumentative Foundation for Evaluative AI
Source: [https://arxiv.org/html/2608.07473](https://arxiv.org/html/2608.07473)
\\settopmatter
printacmref=false\\setcopyrightnone\\acmConference\[arXiv\]arXiv preprint\\acmDOI\\acmPrice\\acmISBN\\acmSubmissionID¡¡submission id¿¿\\affiliation\\institutionImperial College London\\cityLondon\\countryUnited Kingdom\\affiliation\\institutionThe University of Queensland\\cityBrisbane\\countryAustralia\\affiliation\\institutionCardiff University\\cityCardiff\\countryUnited Kingdom\\affiliation\\institutionKing’s College London\\cityLondon\\countryUnited Kingdom\\affiliation\\institutionImperial College London\\cityLondon\\countryUnited Kingdom
###### Abstract\.
Evaluative AI \(EAI\) has been recently proposed as a way to support human decision\-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each\. In this position paper, we advocate \(computational\) argumentation as a most suitable paradigm to provide a formal, computable foundation for forms of EAI that are explainable and contestable\. Argumentation can also naturally pave the way to a multi\-agent vision for EAI in which diverse evaluative models, derived from different sources and reasoning styles, can interact and jointly deliberate on hypotheses and evidence\. Overall, this position paper sets the ground for a long\-term research agenda towards distributed and human\-centred EAI systems\.
###### Key words and phrases:
Evaluative AI, Argumentation, Explainability, Contestability
## 1\.Introduction
Explainable AI \(XAI\) aims to enhance the transparency and trustworthiness of AI systems by elucidating their decision\-making processesadadi2018peeking, a capability that is essential in high\-stakes domains such as healthcare, the judiciary, and finance\.*Evaluative AI \(EAI\)*EAI\_Tim\_Millerhas been recently proposed as a new form of XAI to support human decision\-making in a*hypothesis\-driven*rather than*recommendation\-driven*fashion\. The latter is a widely used paradigm in XAI: an AI model first produces an output, and XAI methods subsequently explain or justify it \(e\.g\.,lundberg2017unified;ribeiro2016should;wachter2017counterfactual\)\. Yet, recent studies show that such “recommend\-and\-explain” paradigm can induce cognitive fixation, leading users to either over\-rely on or dismiss AI recommendations without sufficient deliberationbuccinca2021trust;gajos2022people;sivaraman2023ignore\. To address these issues, instead of directly presenting a recommendation, EAI offers multiple plausible*hypotheses*together with structured*evidence*for and against each\. Thus, by ensuring that the human actively evaluates and deliberates between competing hypotheses, EAI preserves human agency in decision\-making\. This, in turn, strengthens the human\-in\-the\-loop paradigm and repositions AI systems from decision\-makers to facilitators\.
In this position paper we set the ground for a long\-term research agenda towards distributed and human\-centred EAI systems with clear formal and algorithmic backings\. Specifically, we adopt a formal understanding of the EAI problem as a*ranking\-based*problem over all potential hypotheses, given the pro and con evidence associated with each, as was suggested as future work byEAI\_Tim\_Miller\. For example, in healthcare, given a patient’s symptoms \(evidence\), an EAI\-driven system may generate a ranked list of possible diagnoses \(hypotheses\) to assist practitioners in their decision\-making process\. We see a ranking over the hypotheses not as a final decision but, rather, as reflecting the “preferences” of the EAI\-driven system proposing it, given the evidence\. Then, our main position is that \(computational\) argumentation, particularly through the use of weighted Quantitative Bipolar Argumentation Frameworks \(wQBAFs\)mossakowski2018modular;chi2021optimized;potyka2021interpreting, is a most suitable paradigm to provide a formal and computational foundation for ranking\-based EAI, with advantages over the originally proposed Weight of Evidence \(WoE\) methodTim\_WoEin terms of reasoning with structured hypotheses and evidence, explainabilityvcyras2021argumentative;vassiliades2021argumentation, human engagement \(in particular as concerns the possibility for humans to contestleofante2024contestablespecifically evidence, hypotheses, and ultimately rankings\), and finally distributed variants of EAI where diverseagents, derived from different sources and reasoning styles, can interact and jointly deliberate on hypotheses and evidence\.
Figure 1\.Skeleton of an argumentative solution to an example ranking\-based EAI problem in healthcare\. Blue nodes denote arguments; green/red edges indicate, resp\., support/attack relations\. If, in the corresponding wQBAF, all arguments have an initial weight of 0\.5 \(the neutral value in \[0,1\]\) and the edges all have the same edge weight of 1 \(the top value in \[0,1\]\) , then the O\-QuAD semanticschi2021optimizedyields the rankingAnti\-viral=Antibiotics≻Bronchodilator\\texttt\{Anti\-viral\}=\\texttt\{Antibiotics\}\\succ\\texttt\{Bronchodilator\}with corresponding strengths 0\.40, 0\.40, 0\.31\.### Motivating Illustration\.
Figure[1](https://arxiv.org/html/2608.07473#S1.F1)depicts the skeleton of an argumentative solution to the ranking\-based EAI problem in a healthcare scenario111This is a simplified example used solely for illustration and does not reflect clinically validated relations\.\. This amounts to*arguments*, organised hierarchically across three layers via*attack and support relations*\. The arguments in the ‘symptom layer’ represent observable evidence \(e\.g\., high fever, cough\), in the ‘treatment layer’ represent possible clinical decisions \(e\.g\., anti\-viral therapy, antibiotics\), and in the intermediate ‘diagnosis layer’ \(e\.g\., viral/bacterial pneumonia, asthma\) play a dual role: as hypotheses with respect to the evidence below, but also as evidence for the hypotheses above\.222Note that our argumentative solution is not restricted to 3\-layer structure, and can accommodate any structure\.The support and attack relations reflect dependencies between hypotheses and evidence\. A wQBAF then augments this skeleton by associating to each argument an*initial weight*, which may represent their importance or the confidence level in its validity, and to each edge a*relation weight*, which may capture the influence of the edge towards the affected argument\. Further, quantitative*evaluation methods*\(e\.g\.,chi2021optimized;potyka2021interpretingcan be used to determine the*\(final\) strength*of each argument, recursively based on its own initial weight and the combined strengths of its supporters and attackers, enabling conflict resolution with the available informationvcyras2021argumentative;potyka2021interpreting;ayoobi2023sparx;potyka2023explaining\. Finally, from the arguments’ strength, we can obtain a ranked list of the hypotheses in the treatment layer\. The wQBAF and associated strengths are interpretable and can be used for explanation of the ranking and/or the validity of any hypothesis or evidence, in the spirit ofvcyras2021argumentative;vassiliades2021argumentation, in a variety of formats, e\.g\.,kampik2022explaining;AAE\_ECAI;amgoud2017measuring;YIN\_RAE\_IJCAI;AAEsRAEs\-JAIR24;Caren\_2025impactmeasure;Tim\_set\_contributionMoreover, should a human disagree with any component of the QBAF and/or the ranking, they can contest it, in the spirit ofyin2025contestability, e\.g\. to disagree with the initial weight of some evidence and/or add additional pro and con evidence or hypotheses\. This ability to contest promotes human\-in\-the\-loop decision\-making, which lies at the heart of EAI\. Finally, the wQBAF may present the opinion of an agent and could be integrated with opinions by other agents, e\.g\., with argumentative exchanges as inrago2023interactive, or result from the aggregation of QBAFs from different agents, e\.g\., as inmerging\_argumentation\.
### Contributions\.
The vision of this position paper is threefold\.
- •EAI can be formally understood as a ranking\-based problem \(Section[3](https://arxiv.org/html/2608.07473#S3)\)\.
- •wQBAFs from the field of \(computational\) argumentation provide a principled paradigm for realising ranking\-based EAI in an explainable and contestable manner suitable for human\-centred EAI systems \(Section[4](https://arxiv.org/html/2608.07473#S4)\)\.
- •wQBAFs can empower agent\-level EAI within a multi\-agent vision for EAI, in which diverse evaluative models interact and jointly deliberate \(Section[5](https://arxiv.org/html/2608.07473#S5)\)\.
## 2\.Related Work
### Evaluative AI \(EAI\)
EAI is a recent paradigmEAI\_Tim\_Miller, and its practical realisation is still in the early stages of investigation\. Only a handful of approaches have been proposed so far to realise EAI, notablyTim\_WoE;ermellino2024approach:Tim\_WoEadopt Weight of Evidence \(WoE\)melis2021humanto quantify both the direction \(supporting or opposing\) and the magnitude of the impact that each piece of evidence has on candidate hypotheses; andermellino2024approachutilise large language models \(LLMs\) to generate pro and con evidence in a conversational style\. In contrast to these approaches, our envisaged ranking\-based argumentative approach to EAI allows richer structures of dependencies between hypotheses and evidence, e\.g\., hypotheses may influence other hypotheses\. It also allows \(i\) to separate the stance of evidence, \(ii\) to determine the rating and ranking of hypotheses via existing evaluation methods from wQBAFs, benefiting from an extensive literature on their formal and practical propertiesbaroni2018many;yin2025contestabilityand \(iii\) to derive faithful explanations for the rankings from the evaluations, while also \(iv\) mitigating against hallucination risks when evidence and hypotheses are drawn from LLMs with the help of the contestable nature of argumentative solutionsargllms\.
### Argumentative XAI
Argumentation has emerged as an influential paradigm for supporting XAI with structured, transparent, and human\-aligned reasoning processes \(see recent surveysvcyras2021argumentative;vassiliades2021argumentation\)\. Our EAI vision can be seen as falling under the category of*intrinsic*argumentative explanationsvcyras2021argumentativein which the underlying model is already argumentative \(in our vision, a wQBAF\)\. We can then leverage on a variety of explanation formats for explaining rankings between hypotheses, e\.g\. based on assessing the most influential evidence for or against various ranked hypotheses, in the spirit of*argument/relation attribution explanations*kampik2022explaining;AAE\_ECAI;YIN\_RAE\_IJCAI;AAEsRAEs\-JAIR24;amgoud2017measuring;Caren\_2025impactmeasure;Tim\_set\_contribution\. For instance, the top\-ranked hypothesisAnti\-viral\(Figure[1](https://arxiv.org/html/2608.07473#S1.F1)\) attains a strength of 0\.40, whereHigh Fever,CoughandPositive PCR testare computed as the most influential symptoms following the approach ofAAE\_ECAI\. An important aspect of explanations drawn from intrinsically argumentative solutions is that they are by definition*faithful*argllms\. We envisage that this is an important endorsement for our vision for trustworthy EAI\.
### Contestability\.
The need for AI to be contestable in general is widely acknowledgedcontestableAI\-alfrink, and argumentation is advocated by several as a most suitable formalism to achieve principled forms of contestable AIleofante2024contestable;aamas25contestableAI\. Contestable argumentative solutions have been proposed, e\.g\., when QBAFs are generated by LLMsargllmsand with the help of*counterfactual explanations*yin2024qarg;kampik2024change;yin2025contestability\. The latter indicate how to modify the weights of arguments or relations to change the strength of a hypothesis to a desired one, which allow users to interact with and challenge EAI systems built argumentatively\. The contestability naturally afforded by argumentative solution is a strong motivation for our vision of ranking\-based argumentative EAI that is amenable to human consumption\.
### Multi\-Agent Argumentation\.
Several works envisage argumentation as the basis for multi\-agent interaction and deliberation, and can serve as a starting point for a multi\-agent vision of argumentative EAI\. Existing approaches are based on static aggregation of argumentation frameworks from various agents, e\.g\., abstract argumentation frameworks inmerging\_argumentationand bipolar argumentation frameworks inDBLP:journals/aamas/DickieLBRT25, or on dynamic, dialogue\-based combinations, e\.g\., of weighted abstract argumentation frameworks inDBLP:conf/atal/TarleBM22and of QBAFs inrago2023interactive\. The latter could be especially useful as a starting point to support our vision of wQBAF\-based EAI\.
## 3\.The Ranking\-based EAI Problem
Letℋ\\mathcal\{H\}be a finite, non\-empty set of*hypotheses*, andℰ\\mathcal\{E\}be a finite set of \(pieces of\)*evidence*\. Hypotheses and evidence can be associated with an*initial weight*that capture how plausible they are before considering any dependencies amongst them\.
###### Definition 0\(Initial Weight\)\.
The functionτ:\(ℰ∪ℋ\)→\[0,1\]\\tau:\(\\mathcal\{E\}\\cup\\mathcal\{H\}\)\\rightarrow\[0,1\]maps eachx∈\(ℰ∪ℋ\)x\\in\(\\mathcal\{E\}\\cup\\mathcal\{H\}\)to its*initial weight*τ\(x\)\\tau\(x\)\.
In the motivating illustration, the initial weight of a symptom \(e\.g\.,cough\) or of a medical condition \(e\.g\.,asthma\) may capture its perceived severity for a given patient\. For instance, ahigher fevercorresponds to a higher initial weight, reflecting the greater severity of the symptom\.
\(Positive and negative\) dependencies within hypotheses and evidence can be modelled in terms of \(pro and con, resp\.\) relations\.
###### Definition 0\(Pro/Con Relations\)\.
𝗉𝗋𝗈,𝖼𝗈𝗇⊆\(ℰ×\(ℰ∪ℋ\)\)∪\(ℋ×ℋ\)\\mathsf\{pro\},\\mathsf\{con\}\\subseteq\(\\mathcal\{E\}\\times\(\\mathcal\{E\}\\cup\\mathcal\{H\}\)\)\\cup\(\\mathcal\{H\}\\times\\mathcal\{H\}\)are disjoint binary relations, referred to, resp\., as the*pro*and*con relations*\. For allx,y∈\(ℰ∪ℋ\)x,y\\in\(\\mathcal\{E\}\\cup\\mathcal\{H\}\),xxis called*pro/con evidence or hypothesis*foryyiff\(x,y\)∈𝗉𝗋𝗈\(x,y\)\\in\\mathsf\{pro\}/\(x,y\)∈𝖼𝗈𝗇\(x,y\)\\in\\mathsf\{con\}, resp\.
Note that here we consider, besides direct relations from evidence to hypotheses \(ℰ×ℋ\\mathcal\{E\}\\times\\mathcal\{H\}\) as inEAI\_Tim\_Miller, also relations among evidence \(ℰ×ℰ\\mathcal\{E\}\\times\\mathcal\{E\}\) and among hypotheses \(ℋ×ℋ\\mathcal\{H\}\\times\\mathcal\{H\}\)\. In this way, evidence may also attack or support other evidence\. This extension allows multi\-step reasoning and evaluation\.
To quantify how strongly a piece of evidence affects the plausibility of another element \(hypothesis or evidence\), we use the notion of*relation weight*\.
###### Definition 0\(Relation Weight\)\.
The functionw:\(𝗉𝗋𝗈∪𝖼𝗈𝗇\)→\[0,1\]w:\(\\mathsf\{pro\}\\cup\\mathsf\{con\}\)\\rightarrow\[0,1\]maps eachr∈\(𝗉𝗋𝗈∪𝖼𝗈𝗇\)r\\in\(\\mathsf\{pro\}\\cup\\mathsf\{con\}\)to its*relation weight*w\(r\)w\(r\)\.
In the motivating illustration, the relation weight for an evidence\-hypothesis pair \(e\.g\.,\(cough,asthma\)\) represents the strength of the influence from the symptomcoughto the diagnosisasthma\. A higher weight indicates a stronger influence\.
We next formally define EAI as a*ranking\-based problem*\.
###### Definition 0\(The Ranking\-based EAI Problem\)\.
GivenE=⟨ℰ,ℋ,E=\\left\\langle\\mathcal\{E\},\\mathcal\{H\},\\right\.𝗉𝗋𝗈,𝖼𝗈𝗇,τ,w⟩\\left\.\\\!\\mathsf\{pro\},\\mathsf\{con\},\\tau,w\\right\\rangle, the*ranking\-based EAI problem*is to determine a total preorder≽E\\succcurlyeq\_\{E\}onℋ\\mathcal\{H\}\.333A total preorder is a reflexive, transitive and total relation\.
Intuitively,≽E\\succcurlyeq\_\{E\}represents the relative*plausibility*of hypotheses given the weighted pro and con relations, and the initial weights of evidence and hypotheses\. For anyh1,h2∈ℋh\_\{1\},h\_\{2\}\\in\\mathcal\{H\},h1≽Eh2h\_\{1\}\\succcurlyeq\_\{E\}h\_\{2\}implies thath2h\_\{2\}is not more plausible thanh1h\_\{1\}givenEE\.
Formalising EAI as a ranking\-based problem offers two key advantages\. First, ranking naturally accommodates multiple hypotheses rather than forcing a single recommendation, which aligns directly with the core philosophy of EAI\. Second, ranking offers a concrete computational framework, both for generating rankings \(e\.g\.,herbrich2000large;burges2005learning\) and for evaluating them \(e\.g\., Kendall’sτ\\taukendall1938new\), which transforms EAI from a conceptual vision into a practicable research agenda\. We discuss next how to use argumentation to identify solutions to the ranking\-based EAI problem\.
## 4\.An Argumentative Solution
In this section, we advocate that argumentation can lead to the identification of solutions to the ranking\-based EAI problem with desirable properties\. Specifically, we envisage the use of*weighted Quantitative Bipolar Argumentation Frameworks \(wQBAFs\)*mossakowski2018modular;chi2021optimized;potyka2021interpreting, which are quintuples𝒬=⟨𝒜,ℛ−,ℛ\+,τ,w⟩\\mathcal\{Q\}=\\left\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\},\\tau,w\\right\\ranglewhere𝒜\\mathcal\{A\}is a finite set of*arguments*,ℛ−,ℛ\+⊆𝒜×𝒜\\mathcal\{R\}^\{\-\},\\mathcal\{R\}^\{\+\}\\subseteq\\mathcal\{A\}\\times\\mathcal\{A\}are disjoint binary relations called*attack*and*support*,τ:𝒜→\[0,1\]\\tau:\\mathcal\{A\}\\rightarrow\[0,1\]is a*base score function*, andw:ℛ−∪ℛ\+→\[0,1\]w:\\mathcal\{\{R\}^\{\-\}\\cup\{R\}^\{\+\}\}\\rightarrow\[0,1\]is an*edge weight function*\. These wQBAFs can be used to represent EAI problems\.
###### Definition 0\(wQBAF Representation for EAI Problems\)\.
The*wQBAF representation of*E=⟨ℰ,ℋ,E=\\left\\langle\\mathcal\{E\},\\mathcal\{H\},\\right\.𝗉𝗋𝗈,𝖼𝗈𝗇,τ,w⟩\\left\.\\\!\\mathsf\{pro\},\\mathsf\{con\},\\tau,w\\right\\rangleis the wQBAF𝒬=⟨𝒜,ℛ\+,ℛ−,τ,w⟩\\mathcal\{Q\}=\\langle\\mathcal\{A\},\\mathcal\{R\}^\{\+\},\\mathcal\{R\}^\{\-\},\\tau,w\\ranglesuch that:
- •𝒜=ℰ∪ℋ\\mathcal\{A\}=\\mathcal\{E\}\\cup\\mathcal\{H\};
- •ℛ−=𝖼𝗈𝗇\\mathcal\{R\}^\{\-\}=\\mathsf\{con\};
- •ℛ\+=𝗉𝗋𝗈\\mathcal\{R\}^\{\+\}=\\mathsf\{pro\}\.
We define𝒜\\mathcal\{A\}as the setℋ∪ℰ\\mathcal\{H\}\\cup\\mathcal\{E\}, so that debate can take place about all evidence and hypotheses\. As for the other components, we use those of the EAI problem directly\.
Once we map the EAI problem into an argumentative representation, identifying solutions to the ranking\-based EAI problem requires two steps\. First, we compute the strengths of arguments \(both evidence and hypotheses\) using quantitative evaluation methods, which update the initial weights of arguments by taking support and attack relations into account \(see, e\.g\.,chi2021optimized;potyka2021interpreting\)\.444Existing ranking\-based semanticsamgoud2013rankingare defined for abstract argumentation frameworks and cannot be directly adapted to our quantitative setting\.Second, we derive a ranking over hypotheses by comparing their strengths, as illustrated in the Introduction\. The wQBAF solution to the ranking\-based EAI problem is naturally explainable: qualitatively, the graphical structure shows the relationship between different arguments and users can see the reasoning path from a piece of evidence to a hypothesis; quantitatively, the final strength of each hypothesis is explainable via many possible explanation method for QBAFs, such as attribution methodAAE\_ECAI;YIN\_RAE\_IJCAIand counterfactual explanationsyin2024qarg;yin2025contestability, affording also contestability by humans who can modify the weights of arguments or relations to change the strength of a hypothesis to a desired one\. Our envisaged argumentative solution also leads naturally to the satisfaction of a number of desirable principles for EAI, introduced\. These principles serve both as design guidelines for selecting or developing ranking methods and as criteria for evaluating their behaviour\.
### Point\-wise Principles
Point\-wise principles consider the ranking of individual hypotheses\.
Principle 1 \(Monotonicity\)\.Monotonicity is a fundamental guarantee to ensure that rankings remain aligned with rational intuition and expectation\. Monotonicity requires that the relative ranking of an individual hypothesis should change in a way consistent with the direction of change in the underlying factors, such as the introduction of new evidence or updates to the weights of existing evidence\. For instance, adding additional pro evidence for a hypothesis should never cause its rank to fall; and increasing the initial weight of its piece of pro evidence should not lower its position relative to others\.
Many widely used evaluation methods for wQBAFs \(e\.g\., O\-QuADchi2021optimized, MLP\-based semanticspotyka2021interpreting, and edge\-weighted variants of REBamgoud2018evaluationand QEPotyka18\) are known to satisfy Principle 1yin2025contestability\.
Principle 2 \(Balance\)\.Balance ensures that a ranking remains unchanged when the pro and con evidence are balanced\. In particular, any modification of the evidence that preserves the equality between the overall strength of pro and con evidence should not affect the relative ranking of hypotheses\. This includes, for instance, redistributing weights among existing pro and con evidence or introducing additional evidence, provided that the total pro and total con influence remain equal and therefore cancel each other out\.
The previously mentioned semantics also satisfy Principle 2\.
### Pair\-wise Principles
Pair\-wise principles consider the ranking comparison of two hypotheses\.
Principle 3 \(Equivalence\)\.Equivalence states that two hypotheses must be ranked equally if they share the same set of pro and con evidence with the same initial weight\. This principle is fundamental for ranking fairness, ensuring symmetric treatment for evidentially symmetric cases\. Beyond fairness, it is also a prerequisite for rational and intuitive ranking method behaviour\.
The previously mentioned semantics also satisfy Principle 3\.
Principle 4 \(Dominance\)\.The Dominance principle establishes a preference in the ranking under controlled conditions\. For example, a hypothesis must be ranked no lower than another if, given an equivalent set of con evidence against both, either \(1\) its set of pro evidence is a superset of the other’s, meaning that more pro evidence ranking higher; or \(2\) the aggregated weight of its pro evidence is greater\. This principle is fundamental because it ensures the ranking objectively rewards hypotheses with superior pro evidence \(either in number or in strengths\), providing a clear and fair rationale for comparative assessment\.
The previously mentioned semantics also satisfy Principle 4\.
### General Principles
While pointwise and pairwise principles ensure local rationality, general principles govern the global behaviour of rankings\. We envision two key principles at this level\.
Principle 5 \(Robustness\)\.Robustness requires that the overall ranking exhibits stability when minor perturbations to the underlying evidence or its associated weights occur\. Specifically, small variations in the evidence should not induce large\-scale or counter\-intuitive reversals in the resulting ranking order\. This principle is critical for ensuring the reliability and practical dependability of the EAI rankings, as users must be able to trust that its evaluations are not brittle or overly sensitive to noise\.
It is unclear whether existing evaluation methods for wQBAFs ensure robustness at the ranking level; if not, new robustness\-oriented ranking semantics may be needed\.
Principle 6 \(Explainability\)\.Explainability requires that the ranking is generated through a mechanism whose logic is transparent and whose outcomes can be clearly explained\. This entails providing both local explanations \(e\.g\., why one hypothesis is ranked above another\) and a global rationale for the overall ordering\. This principle is fundamental to transforming the ranking from an opaque output to an understandable result, thereby fostering user comprehension and trust\.
As we already argued earlier, a large array of argumentative explanations drawn from wQBAFs can support this principle\.
Principle 7 \(Contestability\)\.Contestability ensures that the EAI ranking can be challenged by users\. This includes the ability to question not only the veracity of individual evidence, or their initial weights assigned, but also the final ranking itself\. By explicitly supporting such challenges, this principle reaffirms the role of EAI as a deliberative partner in a human\-AI team, empowering users to engage with and refine the system’s reasoning process critically\.
Again, as already argued earlier, wQBAFs show promise to support this principle\.
## 5\.Multi\-agent EAI via Argumentation
Real\-world evaluative processes rarely stem from a single viewpoint: they integrate heterogeneous information sources \(from human experts to domain tools and even LLM\-generated evidence\), diverse ways of identifying argumentative relations, and different reasoning mechanisms induced by alternative gradual semantics\. In the illustration, different agents may contribute the parts of the wQBAF relative to the different hypotheses in the treatment layer, reflecting different medical expertise\. A multi\-agent perspective therefore not only mitigates individual bias but also enables richer, more complementary evidence and reasoning styles, opening the door to more robust and trustworthy evaluations\. Viewing each wQBAF as an autonomous argumentative agent suggests a broader research vision: agents may engage in argumentative communication, exchanging arguments or relations to surface new evidence, reconcile inconsistencies, or converge through principled exchange protocolsrago2023interactive\. Also, multiple wQBAFs may be fused into a group\-level evaluation via semantic alignment of arguments, clustering of evidence structuresgorur2025retrieval, or ensemble\-style aggregation over hypothesis rankingsganaie2022ensemble\. This multi\-agent outlook positions argumentation\-based EAI as the foundation for a distributed, resilient, and genuinely deliberative form of EAI, where interacting argumentative agents together construct evaluations that have the potential not only of being less biased and more robust, but also explainable and contestable\.
## 6\.Discussion
The primary impact of this position paper lies in transitioning EAI from a conceptual proposal to a tractable research agenda by establishing a formal, ranking\- and argumentation\-based foundation\. Furthermore, it inherently paves the way towards explainability, as the argumentative structure makes the entire evaluation process explicit and auditable\. Finally, this explicit reasoning process empowers contestability, transforming the AI from a static recommender into a dynamic deliberative aider that users can meaningfully challenge and refine, thereby fostering a new paradigm of collaborative and trustworthy human\-in\-the\-loop decision\-making\.
Some potential limitations may shape the scope of our vision\. First, our framework assumes that wQBAFs, including arguments, relations, and associated credibility score, are already correctly constructed, whereas in practice the reliable extraction of argumentative structure remains a core bottleneck; advances in LLM\-based argument mining may help alleviate this challengegorur2025can;gorur2025retrieval\. Building on this, even when a wQBAF is successfully constructed, the openness required for contestability may introduce a second limitation: users may strategically or unintentionally provide misleading or low\-quality evidence, creating risks of manipulation; mechanisms such as permissioned contribution and trust\-aware weighting may offer partial safeguards\. Finally, beyond individual agents, multi\-agent EAI systems may naturally exhibit deep and sometimes irreducible disagreementwu2025hidden\. Rather than enforcing consensus, this limitation suggests the need for managing such pluralism, supporting structured disagreementliang2024encouraging, and helping users navigate conflicting evaluations\.
## References相似文章
人工智能评估应与人类协作
这篇立场论文主张,人工智能评估应转向评估人机团队,而非超人类性能,以促进更好的社会成果。
人工智能与人类评判的批判性思维反论证
本研究探讨在教育情境下,学生针对AI生成内容撰写反论证以培养批判性思维,并发现前沿大语言模型能够以与人类评估者中等一致性的方式评估此类写作。
立场:我们需要实用的AI对齐方法来镜像人类推理
这篇立场论文认为,用于高风险决策的AI系统应以与其用户相似的方式进行推理,并忠实地传达这种推理,同时概述了实现这种“认知对齐AI”的研究议程。
立场:推理是一种可学习的基于规则的过程
这篇立场论文认为,AI推理缺乏清晰的操作性定义,削弱了评估的有效性,并提出将推理定义为一种可学习的基于规则的过程,同时提供研究最佳实践的检查清单。
AI认知风险:新兴机制与证据 [R]
一篇由30位专家合著的新论文探讨了来自人工智能的认知风险—即对我们形成准确信念和良好推理能力的威胁—包括说服、认知卸载和反馈循环等机制,并概述了减轻这些风险的方向。