Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
Summary
This position paper argues that AI systems used in high-stakes decision-making should reason similarly to their users and faithfully communicate that reasoning, and outlines a research agenda for achieving such 'cognitively-aligned AI'.
View Cached Full Text
Cached at: 08/14/26, 09:24 AM
# Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
Source: [https://arxiv.org/html/2608.12372](https://arxiv.org/html/2608.12372)
Breanna K\. NguyenCyrus CousinsVincent ConitzerWalter Sinnott\-ArmstrongJana Schaich Borg
###### Abstract
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision\-makers\. This position paper argues that in many settings, particularly high\-stakes decision\-making, we need accurate*cognitively\-aligned*AI systems that*reason similarly to their users*, and*faithfully communicate*their reasoning\. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment “essential” when an AI’s rationale for a judgment or action is important to them\. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps\. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely\.
Cognitive alignment, Trustworthy AI
## 1Introduction
If you were a literary editor considering using AI to help make initial content reviews, would you prefer \(A\) an AI that uses your personal editorial standards and style preferences, or \(B\) an AI that produces similar judgments to \(A\), but via opaque methods that you don’t recognize? What if you were a fiduciary managing long\-term assets on behalf of a family\. Which would you trust more: \(A\) an AI tool that explicitly and faithfully applies your investment philosophy and the family’s risk assessment approach, or \(B\) an AI tool that predicts your choices for the family accurately, but does so via opaque mechanisms? Finally, suppose you are a physician determining which dying patient will receive an available kidney transplant\. Your hospital requires you to use AI to make allocation decisions on your behalf\. Would you prefer \(A\) an AI that makes decisions in the way you would whilst calm and rested, or \(B\) an AI that makes similar decisions based on a mechanism you understand, but is foreign to you?
Presumably, accuracy matters here\. If you believe one AI tool is dramatically more accurate than another, that could be sufficient reason for you to prefer it\. But imagine that the AIs you have to choose between are comparably accurate overall\. Would you still prefer \(A\) over \(B\) in any of these scenarios? Our position is that:Many people prefer to use and delegate to AI that both reasons as they would given sufficient time and information, and can faithfully convey that reasoning, particularly in high\-stakes applications\. The machine learning field should have robust frameworks and methodologies for producing such*cognitively\-aligned AI*\.
Gonzalez and Heidari \([2025](https://arxiv.org/html/2608.12372#bib.bib327)\)argue that cognitively\-aligned AI would make for effective partners in dynamic human\-AI cooperative interactions\. Here we make a different argument:Cognitively\-aligned AI is*strategically important*to the field of AI alignment\.To clarify, our position is not that cognitively\-aligned AI is*always*best in high\-stakes situations\. Rather, our claim is that it may be the only kind of AI some individuals, organizations, or governments are willing to trust to stand in for them autonomously when stakes are high\. Less provocatively, some people may simply prefer AI that truly thinks like them, or at least reasons the way they try to reason, over AI that reasons in a foreign way, particularly if the AI is meant to serve as a delegate or surrogate for them\. ML researchers and practitioners do not need to be one of these people who personally benefit from or prefer cognitively\-aligned AI, but we argue that the ML field should be able to support those who do\.

Figure 1:Participants’ preference for*Human\-Reasoning AI*over*Machine\-Reasoning AI*,*Process\-Hidden AI*, and*No Preference*across task domains,±\\pm95% Bonferroni\-corrected bootstrap confidence interval\. Domains where at least 50% of participants chose Human\-Reasoning AI are green, those with under 50% are gray, and statistically indeteriminate settings are cyan\.In what follows, we present evidence that users prefer AI that thinks like them \([Section2](https://arxiv.org/html/2608.12372#S2)\), discuss why existing alignment methods fall short \([Section3](https://arxiv.org/html/2608.12372#S3)\), and outline a research agenda for advancing cognitive alignment \([Section4](https://arxiv.org/html/2608.12372#S4)\)\.
## 2Evidence Users Want Cognitive Alignment
There are many situations where people may*not care*whether AI thinks like them, or might have so little trust in their own reasoning — or any human reasoning — that they actually*prefer*AIs that think differently\. If a machine\-reasoning AI is as accurate as a human\-reasoning AI, people might be ambivalent about cognitive alignment in domains like weather prediction or debugging code\.
However, in high\-stakes domains like medical or military decision\-making, trustworthiness criteria change\. Here users report needing to understand the*rationale behind an AI’s decisions*, so they can better predict where it may fail\(Tonekaboniet al\.,[2019](https://arxiv.org/html/2608.12372#bib.bib330); Chamolaet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib318); Hammet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib319); de Brito Duarteet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib320)\)\. This largely motivates the field of explainable AI \(xAI\), or AI that can explain decisions in a way humans can understand\. The extent to which an AI’s decisions are explainable correlates with*user trust*in the AI and*willingness to use*or*follow recommendations from*it\.\(Shin,[2021](https://arxiv.org/html/2608.12372#bib.bib331); Lopezet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib333)\), particularly when the AI makes choices that differ from those of a human user\(Riveiro and Thill,[2022](https://arxiv.org/html/2608.12372#bib.bib341)\)\. In some domains, the*nature*of explanations matter, too, e\.g\., physicians want explanations of the critical steps an AI takes towards a decision to use language that naturally relates to their practice, and want rationales that reference familiar evidence, like test results and medical literature citations\(Cortiet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib362)\)\.
Importantly, at least in theory, an AI can explain its decisions without relying on the same reasoning processes as a human\. In some cases this may even be desirable, as AI’s alternative approaches can detect errors humans miss, improving human\-AI team performance overall\(Wilderet al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib395)\)\(also note evidence to the contraryVaccaroet al\.\([2024](https://arxiv.org/html/2608.12372#bib.bib394)\)\)\. However, trust becomes increasingly critical as decision stakes rise, and explanations that users cannot understand are unlikely to engender trust\. Despite this, under 1% of xAI methods have actually been evaluated for human understanding\(Siuet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib340)\)\. When tested, even experts often misunderstand AI explanations\(Hurleyet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib339); Ehsanet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib342)\), while non\-experts require additional support\(Schulze\-Weddige and Zylowski,[2021](https://arxiv.org/html/2608.12372#bib.bib343); Bobeket al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib344); Morandiniet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib345)\)\. Such failures can be consequential, as illustrated by Air Force pilots who worry AI rationales will be impossible to interpret and will introduce confusion and conflict within human–AI teams\(Lopezet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib333)\)\.
Cognitive faithfulness may improve explanation understanding by making AI reasoning easier for users to comprehend\. Humans are likely to find explanations grounded in familiar reasoning easier to understand than those invoking unfamiliar logic\(Grgić\-Hlačaet al\.,[2022](https://arxiv.org/html/2608.12372#bib.bib329); Sawet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib335)\)\. Although this hypothesis has not been tested directly, neuroscience research suggests that people default to interpreting others’ thinking by comparison to their own\(Bradfordet al\.,[2015](https://arxiv.org/html/2608.12372#bib.bib346)\)\. Evaluating unfamiliar reasoning frameworks thus requires inhibiting one’s own perspective, necessitating further cognitive effort\(Samuelet al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib347)\)\.Tonekaboniet al\.\([2019](https://arxiv.org/html/2608.12372#bib.bib330)\)report that some clinicians expect AI systems that reason using human\-like strategies to be more interpretable, andCortiet al\.\([2024](https://arxiv.org/html/2608.12372#bib.bib362)\)find that explanations that mirror physicians’ natural thought processes are viewed especially favorably\. Contrapositively,Bobeket al\.\([2025](https://arxiv.org/html/2608.12372#bib.bib344)\)find that domain experts struggle to interpret xAI outputs that lack alignment with disciplinary reasoning practices\. Consistent with these views,Chandaet al\.\([2024](https://arxiv.org/html/2608.12372#bib.bib328)\)find that dermatologists’ trust in an AI designed to align with clinicians’ diagnostic approach to melanoma correlates with the*degree of overlap*between a dermatologist’s explanations and the AI’s explanations\. Similarly, both crowdsourced participants and military medical triage experts are more likely to trust and delegate decisions to AIs they perceive as making medical decisions in the way they would\(Summervilleet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib323)\)\.
Cognitively aligned AI can also be desirable beyond explainability and comprehensibility\. It seems naturally desirable for AI “productivity twins” who are meant to serve as user surrogates\. Additionally, in ethical situations when moral ethical integrity is on the line, some users may want an AI to act for what they regard as the*right ethical reasons*, which is difficult to satisfy when an AI reasons in ways that feel fundamentally foreign\(Zhi\-Xuanet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib349); Gabriel,[2020](https://arxiv.org/html/2608.12372#bib.bib163)\)\. Consistent with this, users are more willing to forgive AI errors when they believe the system is guided by rules, principles, or logic they view as ethically sound and well\-intentioned\(Phillips and Malle,[2025](https://arxiv.org/html/2608.12372#bib.bib332)\)\. Further, in medical triage settings people are more willing to trust and delegate decisions to AIs whose prioritization of ethical considerations aligns with their own\(Summervilleet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib323); McVayet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib357)\)\.
Taken together, existing evidence suggests that people seek, prefer, and are more willing to delegate to cognitively aligned AI in many high\-impact domains\. However, we are unaware of prior work that directly asks users whether they want AI to think like them\. Next, we present a pilot study that addresses that gap\.
##### A proof\-of\-concept study to fill evidential gaps\.
We asked 150 Prolific participants to evaluate the desirability of cognitively aligned AI in real\-world domains \(see Appendix[A\.1](https://arxiv.org/html/2608.12372#A1.SS1)for methodological details\)\. Participants first viewed descriptions of three hypothetical AI systems, that each start with, “This type of AI gives you a recommendation or makes a decision\.” The descriptions continue as follows:
- Process\-Hidden AI:It is not possible for it to accurately tell you how it arrived at its recommendation or decision, or what reasoning it used to arrive at its output\. Machine\-Reasoning AI:It accurately communicates to you how it arrived at its output, but the reasoning process it uses can feel unfamiliar or foreign to human ways of thinking, and is very different from the one that you usually use to make similar recommendations or decisions\. Human\-Reasoning AI:It accurately communicates to you how it arrived at its recommendation or decision, and the reasoning process it uses is intentionally designed to mirror how a thoughtful, informed person would approach the problem\.
Participants were then asked, “If the AIs were equally accurate \(that is, they perform equally well overall at making correct recommendations or decisions, on average\), can you imagine any scenarios in which you would prefer Human\-Reasoning AI over Process\-Hidden AI or Machine\-Reasoning AI?” 86\.6% of participants responded “yes” to this question\. When asked to briefly describe those scenarios, responses included “theorem proving,” “student grading,” “Advice on politics, morals, or philosophy,” “High\-stakes decisions like healthcare or finances where understanding the reasoning matters,” and “situations requiring human feelings such as empathy, compassion, and romance” \(full list of responses provided in Appendix[A\.2](https://arxiv.org/html/2608.12372#A1.SS2)\)\.
Next, we asked which AI type participants would prefer across 16 domains, again specifying that the systems had equal accuracy \(survey questions provided in Appendix[B](https://arxiv.org/html/2608.12372#A2)\)\.[Figure1](https://arxiv.org/html/2608.12372#S1.F1)shows that participants often preferred Human\-Reasoning AI\. In fact, the statistically significant majority of participants preferred Human\-Reasoning AI in 5 of the 16 domains, including seeking advice for moral dilemmas, medical allocations, bail eligibility, and military targeting\. Domains where participants were least likely to have a preference for any type of AI system over another included predicting the weather and scheduling team meetings \([Fig\.A4](https://arxiv.org/html/2608.12372#A1.F4)\)\.
We collected more detailed responses for two current AI use cases: autonomous vehicles and kidney allocation\. Participants rated the desirability of the attributes presented in[Figure2](https://arxiv.org/html/2608.12372#S2.F2), which were chosen to distinguish reasons one might prefer one AI over another\. Most participants considered accuracy essential, but over 25% also rated “The AI makes decisions similarly to how you would with sufficient time and information” as*essential*, with another 25% rating it as*very desirable*\. Participants were significantly more likely to rate this cognitively\-aligned quality essential or very desirable in high\-stakes versus low\-stakes scenarios \(Wilcoxonp<0\.001p\{<\}0\.001for all qualities; CIs in Appendix[A\.2](https://arxiv.org/html/2608.12372#A1.SS2)\) — see[Figure3](https://arxiv.org/html/2608.12372#S3.F3)\. Notably, over half of participants also rated the following qualities as essential or very desirable: truthful explanations of the AI’s actual decision process, easily understandable explanations, mechanisms to adjust the AI’s reasoning, and correction for systematic errors or biases\.
For the autonomous vehicle and kidney allocation scenarios, participants additionally ranked variants of Human\-Reasoning, Machine\-Reasoning, and Process\-Hidden AIs that differed in understandability and modifiability, but not in accuracy\. Their top three choices, in order, were consistently Human\-Reasoning AI that was easy to understand and modifiable; Human\-Reasoning AI that was easy to understand but unmodifiable; then Machine\-Reasoning AI that was easy to understand and modifiable \(see[FigureA5](https://arxiv.org/html/2608.12372#A1.F5)\)\.
Our participant sample was from the US, so additional surveys are needed to determine views across global populations\. Future work should also assess how preferences for Human\-Reasoning AI vary across a greater range of contexts, different types of domain expertise, and more diverse populations\. Indeed, our results raise plenty of new questions\. But they also provide substantial preliminary evidence that many users desire — and often deem essential — AI that reasons like them and that faithfully provides easy\-to\-understand explanations of its decision process, especially in high\-stakes domains\.
Figure 2:Participants’ desire for various qualities in AI assisting with autonomous vehicles and kidney transplant allocation\. Beyond*high accuracy*and*explainability*, most also desire*reasoning alignment*,*bias correction*, and*decision strategy adjustment*in these domains\.
## 3Limitations of Current Alignment Methods
##### Overview of current methods\.
The objective of AI alignment is often described as aligning AI systems with human preferences, instructions, and/or values\(Gabriel,[2020](https://arxiv.org/html/2608.12372#bib.bib163)\)\. Alignment to human preferences is generally carried out by collecting corpora of people’s judgments, and then employing “bottom\-up” approaches to train or fine\-tune models on collected data\. For example,Awadet al\.\([2018](https://arxiv.org/html/2608.12372#bib.bib62)\)curated a dataset containing judgments from hundreds of thousands of participants on actions of autonomous vehicles when faced with trolley problem\-like dilemmas, and several follow\-up works use this dataset to learn algorithmic models of individual and aggregated preferences in this consequential setting\(Kimet al\.,[2018](https://arxiv.org/html/2608.12372#bib.bib137); Noothigattuet al\.,[2018](https://arxiv.org/html/2608.12372#bib.bib4)\)\.Leeet al\.\([2019](https://arxiv.org/html/2608.12372#bib.bib57)\)andJohnstonet al\.\([2023](https://arxiv.org/html/2608.12372#bib.bib59)\)built participatory resource allocation tools that elicit stakeholders’ preferences regarding fairness\-utility tradeoffs by presenting them with pairwise comparisons of allocation scenarios that differ in their fairness and utility impacts, and training an algorithmic model on the revealed preference data\. Several other works employ similar preference modeling methods\(Srivastavaet al\.,[2019](https://arxiv.org/html/2608.12372#bib.bib63); Grgic\-Hlacaet al\.,[2018](https://arxiv.org/html/2608.12372#bib.bib302); Freedmanet al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib48)\)\.
In recent years, a similar methodology has improved preference alignment of large AI models by fine\-tuning them using human preference data, via reinforcement learning\(Christianoet al\.,[2017](https://arxiv.org/html/2608.12372#bib.bib232); Wirthet al\.,[2017](https://arxiv.org/html/2608.12372#bib.bib233); Kaufmannet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib234)\)and direct preference optimization\(Rafailovet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib30); Sunet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib1)\)\. This bottom\-up approach of alignment to revealed preferences has clear practical advantages\. Asking participants for only their choices without seeking their reasons simplifies the task of the participant, and allows elicitors and modelers to collect a large amount of training data \(e\.g\., millions of datapoints gathered byAwadet al\.\([2018](https://arxiv.org/html/2608.12372#bib.bib62)\)\)\. Beyond choice\-based alignment, reasoning\-based LLMs are developed using data containing intermediate steps undertaken before reaching an answer, to improve the model’s capabilities on complex reasoning tasks and the quality of users’ interaction with LLMs on these tasks\(Wanget al\.,[2022a](https://arxiv.org/html/2608.12372#bib.bib373),[2023](https://arxiv.org/html/2608.12372#bib.bib374)\)\. Trained using supervised reasoning samples and preference data over different reason\-based responses, these methods attempt alignment by ensuring the model can reason against unsafe actions\(Guanet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib398)\)\.
Alignment of AI with normative rules and principles, in contrast, is achieved through supervision\. Constitutional AI, for instance, achieves \(some degree of\) compliance of AI with pre\-specified rules through iterative model self\-criticism and revision\(Baiet al\.,[2022](https://arxiv.org/html/2608.12372#bib.bib353)\)\. Others consider similar mechanisms to constrain the action space of AI agents to prevent catastrophic harms\.Hadfield\-Menellet al\.\([2017](https://arxiv.org/html/2608.12372#bib.bib369)\)describe the “off\-switch game,” discussing how we can ensure that AI systems do not possess the ability to prevent humans from turning them off\.Turneret al\.\([2020](https://arxiv.org/html/2608.12372#bib.bib368)\)train AI agents to be conservative in their action space in anticipation of changing reward functions\. Others study similar ways to constrain AI agents to maintain oversight or ensure regulatory compliance\(Orseau and Armstrong,[2016](https://arxiv.org/html/2608.12372#bib.bib370); Imperialet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib371); Koltet al\.,[2026](https://arxiv.org/html/2608.12372#bib.bib393)\)\.
##### Current limitations\.
We higlight two limitations of current cognitive alignment methods: \(L1\)unverifiable explanations, which hinders evaluating the overlap between AI reasoning and human reasoning, and \(L2\)cognitive misalignment, which hinders comprehension and trust\.
##### \(L1\) Unverifiability of putatively aligned AI decision\-making processes\.
Our survey results \([Figures2](https://arxiv.org/html/2608.12372#S2.F2)and[3](https://arxiv.org/html/2608.12372#S3.F3)\) indicate that to deliver what users seek from cognitively aligned AI, systems must be able to provide understandable explanations of their decisions,*and*those explanations must truthfully represent the reasoning process of the AI\. However, many widely used alignment methods produce models whose behavior cannot be meaningfully explained, e\.g\., those which model human choices with uninterpretable model classes, such as deep neural networks\(Wanget al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib385); Kweonet al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib384); Rafailovet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib30)\)\. Post\-hoc explanatory methods attempt to elucidate the behavior of such black\-box models, but there is usually no clear way to transform those explanations into a full account of how the output was generated\(Lipton,[2018](https://arxiv.org/html/2608.12372#bib.bib132); Rudin,[2019](https://arxiv.org/html/2608.12372#bib.bib133); Von Eschenbach,[2021](https://arxiv.org/html/2608.12372#bib.bib103)\)\. These methods generally provide explanations of individual decisions, often based on counterfactuals \(i\.e\., what factors are most responsible for a decision\), rather than interpretations of the entire model’s process\.
When aligned AI models do provide explanations, there is often no guarantee that they*faithfully*reflect the AI’s underlying decision process\(Rudin,[2019](https://arxiv.org/html/2608.12372#bib.bib133)\)\. Our survey and previous qualitative studies indicate that it is important to users that explanations accurately reflect the reasoning the AI truly uses\. Providing such assurances with integrity requires that somebody be able to verify which decision\-making processes a system actually used, and that the explanations reliably correspond to those processes\. Many current alignment techniques were not designed to support such verification, or yield disappointing results when verification is attempted\. Even large language models that provide multi\-step rationales along with their output are plagued by such issues\. There is still uncertainty about whether the faithfulness of their reasoning chains can ever be truly assessed\(Turpinet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib351); Korbaket al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib352)\), butBarezet al\.\([2025](https://arxiv.org/html/2608.12372#bib.bib350)\)empirically show that LLM chain\-of\-thought rationales “are frequently unfaithful, diverging from the true hidden computations that drive \[LLM output\]”
Constitutional AI methods face similar limitations\. Although these approaches introduce explicit normative principles during training or inference, these principles constrain model behavior*only indirectly*through preference optimization, critique generation, or prompting rather than by enforcing interpretable internal representations or reasoning procedures\. As a result, even when constitutional models produce explanations that reference their guiding principles, we cannot reliably verify that those principles*causally influenced*the decision, or that the model would continue to reason in accordance with them under domain shift\(Kyrychenkoet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib356)\)\. Further, recent studies suggest that constitutional and chain\-of\-thought\-based reasoning exhibit similar faithfulness failures\(Barezet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib350); Korbaket al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib352)\)\.
In summary, many current alignment techniques fail the requirements of cognitively\-aligned AI, because they*cannot provide verifiable explanations of their reasoning process*\.
##### \(L2\) Misalignment with human cognitive processes\.
Even when wecandiscern the reasoning processes of AI systems aligned to human preferences or explicit rules \(either through interpretable models or from investigations into task\-specific AI behavior\), the next criterion for cognitively\-faithful AI is that it should reason like humans, or for personalized decision\-making, a specific human\. Importantly, most of these methods were not designed with cognitive faithfulness as a primary objective, so lack of cognitive faithfulness should not be surprising\(Kleinberget al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib267)\)\.
Some individual users have anecdotally tried to fine\-tune available LLMs to think like them, and seem at least somewhat satisfied with their results\(Farrell,[2025](https://arxiv.org/html/2608.12372#bib.bib348)\)\. These cases still suffer from theL1faithfulness verifiability problem, but it is worth noting that when LLMs have been examined more broadly, they have frequently been found to think differently than humans\(Schröderet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib386); Zhanget al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib387); Amirizanianiet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib388)\)\.
Alignment using interpretable model classes overcomesL1\(Freedmanet al\.,[2020](https://arxiv.org/html/2608.12372#bib.bib48); Xiao and Wang,[2025](https://arxiv.org/html/2608.12372#bib.bib326)\), but the reasoning they surface may still feel foreign to human decision\-makers\. For instance, several works model human decision\-makers as maximizing a linear utility function when learning from human preferences\(Kimet al\.,[2018](https://arxiv.org/html/2608.12372#bib.bib137); Johnstonet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib59); Leeet al\.,[2019](https://arxiv.org/html/2608.12372#bib.bib57)\), but do not test whether humans actually use a linear process\. When humans were interviewed, many used a thresholded rule\-based decision process that falls outside the scope of linear hypothesis classes\(Keswaniet al\.,[2025a](https://arxiv.org/html/2608.12372#bib.bib286)\)\. Thus, even methods that attempt to explicitly model the complexities of social decision making often use frameworks, sets of assumptions, or levels of analysis that do not match users’ reasoning processes\.
Taken together, L1 and L2 identify two distinct barriers to cognitive alignment in current alignment methods\.L1is based on the unexplainability challenge associated with many AI models, which hinders general attempts by users and stakeholders to obtain faithful explanations for AI decisions\. Even methods that attempt to explain AI decisions often cannot verify that their explanations match the true processes an AI actually implements\. There are settings where more reliable insights about how an AI makes decisions are available through mechanistic analysis or the use of interpretable models\. The AI models used in these situations, though, usually still use mechanisms that differ in important ways from human reasoning; this is the category of limitations covered byL2\.
##### Progress toward cognitive modeling\.
Despite these limitations, progress has been made in domains where cognitive processes are well understood through behavioral economics, psychology, or neuroscience\. For instance,Petersonet al\.\([2021](https://arxiv.org/html/2608.12372#bib.bib44)\)learned behavioral models from large datasets of risky choices that replicate and extend psychological theories like prospect theory, whileZhuet al\.\([2025a](https://arxiv.org/html/2608.12372#bib.bib383)\)modeled strategic decision\-making in two\-player games informed by prior work on human cognition\. Unfortunately, the approaches developed thus far tend to lack generalizability and are highly context\-sensitive, partially because they were constructed through deeply domain\-specific endeavors for which there was no attempt to extend them to other settings\. Nevertheless, these approaches show that cognitive modeling can enable domain\-specific cognitive alignment in principle\(Zhuet al\.,[2025b](https://arxiv.org/html/2608.12372#bib.bib399)\), and motivates some of the research directions we discuss next\.
Figure 3:Participants’ indicated impact for various qualities in trustworthy AI assisting in high\-stakes vs low\-stakes domains\. Once again, reasoning alignment, bias correction, and decision strategy adjustment are viewed as more desirable in high\-stakes domains\.
## 4Priorities for Future Research
##### What level of abstraction should cognitive alignment target?
When users say they want an AI that “thinks like I think,” to*what level of cognitive description*do they refer? In principle, “How I think” can be described at multiple levels of abstraction, ranging from neural activity to the features, values, or societal structures that influence decision\-making\. Cognitive scientists have long debated whether there is “an appropriate level of analysis inside the head at which psychological models \[should\] be developed”\(Bechtel,[1994](https://arxiv.org/html/2608.12372#bib.bib375)\), without consensus\(Colombo and Knauff,[2020](https://arxiv.org/html/2608.12372#bib.bib376)\)\. Cognitive alignment research needs to grapple with a slightly different version of the question: what level of abstraction is necessary for users to perceive an AI system as reasoning in a way meaningfully similar to their own?
Qualitative studies provide initial insight\. Users often look for alignment between decision\-relevant factors they typically consider and those the AI weighs\(Cortiet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib362)\)\. Thus, cognitive alignment may benefit from focusing on conscious, feature\-level reasoning users employ in specific decisions, rather than how people make decisions more generally\. Further, different users may need different types of explanations\. For example, some may want to know what facts are weighed when making a decision, while others may be interested in the values or principles applied\. Context also likely matters\. For instance, in medical diagnosis, radiologists may prioritize pixel\-level image features, while internists may prioritize test outcomes and situational context\(Sawet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib335)\)\. Cognitive alignment methods thus need mechanisms to learn, adapt to, and reflect diversity in preferred abstraction level across users and contexts\.
##### How should cognitive alignment be quantified?
At the simplest level, one straightforward way to measure cognitive alignment is through users’ reports of how well an AI’s reasoning matches their own\. Of course, users can give inaccurate reports, due to mistakes, intentional misdirection, or unintentional inclinations to report what they think others want to hear rather than what they actually think\. Thus, the feasibility and robustness of other methods should be tested as well\. The best approach may be to develop cognitive alignment metrics that incorporate measures of convergence between multiple lines of evidence generated from complementary methods\. Process\-tracing methods could serve as the basis for some assessments\. For example, hidden “information boards” could be used to infer what information people prioritize when making a decision by tracking what information users choose to reveal first, revisit, or spend the most time examining before making a decision\(Schulte\-Mecklenbecket al\.,[2011](https://arxiv.org/html/2608.12372#bib.bib401)\)\. In addition, eye\-tracking assessments could be used to assess what information users attend to when making a decision, in what order, and for how long \(which often correlates with difficulty\)\(Orquin and Loose,[2013](https://arxiv.org/html/2608.12372#bib.bib402)\)\. Other assessments could ask participants to make judgments about queries that are carefully designed to selectively perturb features users said were important to their judgment process, to determine if users’ choices in the face of those perturbations are consistent with their stated reasoning\. Many additional protocols from cognitive science could be promising as well, like think\-aloud paradigms\(Güss,[2018](https://arxiv.org/html/2608.12372#bib.bib403)\)or computational modeling\. The more such assessments align in their conclusions, the more confident we can be in our inference about how a user “thinks”\. A central methodological challenge for future work will be to determine which combinations of self\-report, behavioral, process\-tracing, and model\-based evidence provide the most reliable and valid basis for quantifying cognitive alignment\.
##### How much alignment is enough?
We argue that cognitive alignment can increase the explainability and trustworthiness of an AI system, as well as users’ willingness to delegate decisions to it\. This raises a key question:*How closely must an AI system reflect a user’s reasoning, or desired reasoning, to secure these benefits?*
In some cases, users may tolerate modest deviations\. For example, a user might trust an AI that ranks their most prioritized features in a slightly different order than they would \(particularly if the system has a computational or practical reason for doing so\), as long as those features remain among the most important considerations\. More generally, users may feel aligned with a range of reasoning processes, especially in complex or uncertain environments\.
Evidence from military medical triage illustrates this flexibility\. Medics must weigh how much personal risk they will accept when attending to patients\. Medics report being willing to delegate triage decisions to other medics with different personal risk tolerances, so long as those tolerances fall within an acceptable professional range\(Borderset al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib377)\)\. This suggests that AI systems that prioritize medical safety to different degrees might be similarly tolerated\.
However, not all reasoning dimensions are allowed such flexibility\. Medics who ignore group membership in their own triage decisions are often unwilling to delegate decisions to a colleague who reports prioritizing patients from their in\-group\(Borderset al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib377)\)\. Cognitive alignment may therefore involve both soft constraints, where some variation is acceptable, and hard constraints, where certain reasoning is categorically unacceptable\. Hence, cognitive alignment research should develop methods to learn, represent, and enforce differing degrees of reasoning flexibility\.
##### What should reasoning elicitation and evaluation look like?
Current behavioral alignment methods primarily learn from observed behavior, in the form of people’s choices, feedback, or ratings\. Cognitive alignment seeks to learn not only people’s observed behavior, but also the*reasoning behind it*\. Accomplishing this requires methods to determine what users believe their reasoning process is\. This likely necessitates new preference elicitation methods that ask participants to not only make pairwise choices in high\-stakes scenarios, but also to explain each choice\.
Many challenges arise when eliciting reasoning or feedback about an AI’s reasoning\. To start, what modality should information be shared in? Visualizations are common ways to explain AI reasoning\(Sameket al\.,[2017](https://arxiv.org/html/2608.12372#bib.bib378); Simonyanet al\.,[2013](https://arxiv.org/html/2608.12372#bib.bib379)\), and often have the benefit of directly representing model attributes, such as Shapley values\(Chenet al\.,[2023](https://arxiv.org/html/2608.12372#bib.bib380)\)\. Visualizations for cognitive elicitation could take various forms appropriate for the chosen interpretable hypothesis class\. For instance, generalized additive models could be depicted through feature\-wise partial dependence plots \(i\.e\., how does the outcome change with change in any single feature, keeping other features constant\)\(Hastie and Tibshirani,[1986](https://arxiv.org/html/2608.12372#bib.bib297); Wanget al\.,[2022b](https://arxiv.org/html/2608.12372#bib.bib316)\), and rule\-based models could be depicted through a visual hierarchy of the sets of rules that the model uses to connect the input features to the outcome\(Bendel,[2016](https://arxiv.org/html/2608.12372#bib.bib100); Cousinset al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib299)\)\. However, little is known about how intuitive it would be for human users to communicate their internal reasoning processes through visualizations\. Text or speech may be more natural, but scalable methods are needed to translate the text or speech content to a learned alignment model\. Further, again, preferences and needs for different communication modalities likely vary across users and contexts\(Cortiet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib362)\)\.
Drawing on theinteractive machine learningfield\(Amershiet al\.,[2014](https://arxiv.org/html/2608.12372#bib.bib295); Wondimuet al\.,[2022](https://arxiv.org/html/2608.12372#bib.bib296)\)and reports that users often build trust through*interaction over time*\(Bickmore and Picard,[2005](https://arxiv.org/html/2608.12372#bib.bib396); Bachet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib397)\), one promising approach is to create interactive, multimodal representations of a user’s choice\-based model\. Interactive elements would let users test the model’s consequences and provide feedback about how the model could better reflect the decision\-making process to which they want the AI to align\. For example, after learning an initial interpretable model from paired choices and communicating it to the user, the interface could allow users to directly modify the presented decision rules or choice contributions in partial dependence plots\.
Even if the interface does not permit users to input exactly how they are making their decisions \(e\.g\., due to modeling class limitations or temporal changes in the user’s choice mode\), this process would still allow the model to gather information about its limitations in representing the user’s desired decision process\. An interesting issue that interactive elicitation may illuminate is that users may not initially know how they want to make decisions, and may need time or experience to figure it out\. Participants in pairwise choice elicitation studies report this phenomenon\(Warrenet al\.,[2011](https://arxiv.org/html/2608.12372#bib.bib251)\), which may contribute to changes in what decision models best fit their choices over time\(Boerstleret al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib50); Keswaniet al\.,[2025b](https://arxiv.org/html/2608.12372#bib.bib325)\)\.
Further research is needed to derive robust, scalable methods for learning and representing user reasoning, given these considerations\. Appropriate evaluation frameworks are also needed to assess when and to what degree users feel an AI’s reasoning process agrees with their own reasoning\.
##### What learning methods can best infer human decision processes?
To model human reasoning, we need to understand*how humans reason*in a domain, so that the hypothesis class we fit to data can be constrained to mimic human processes, and we need*human decision\-making data*, so we can fit the model within this restricted class\. Qualitative data may also be used to further constrain the class, to lessen the requirements on quantitative decision\-making data\. To this end, several computational methods have been developed to infer decision\-making processes from observed decisions in the fields of cognitive science, psychology, and behavioral economics, and these offer a promising foundation for future cognitive\-alignment method development\. However, as described in[Section3](https://arxiv.org/html/2608.12372#S3), a feature and a limitation of this literature is its strong context sensitivity\(Gigerenzer and Gaissmaier,[2011](https://arxiv.org/html/2608.12372#bib.bib21)\)\. Models that accurately capture human reasoning in one domain often fail to generalize to others, reflecting deep dependence on task framing, feature representations, and implicit normative assumptions\.
To address this limitation, we need a better understanding ofuser reasoning archetypes, i\.e\., the kinds of reasoning processes people consider to be acceptable or desirable in a given domain\. These reasoning archetypes need not be a single formal object; they can constrain representation, decision procedures, or normative acceptability\. These archetypes can take several different forms, at different levels of abstraction, including \(1\) which features should and should not be used to describe the reasoning process behind the*user’s choice*given the*information presented*, \(2\) how the user interprets and processes the presented features, and \(3\) whether they consider features individually or in combination with each other \(i\.e\., feature interactions\)\. They can also take the form of \(4\) which decision processing \(or abstract categories of decision processing\) is employed in their reasoning, e\.g\., a user who relies on threshold\-based reasoning may find it unacceptable to model their decisions using a simple linear scoring function\. Similarly, reasoning that depends on socially salient attributes — such as gender or ethnicity — in ways that reinforce structural inequities would be considered normatively unacceptable by many users, even if they improve predictive accuracy\(Dworket al\.,[2012](https://arxiv.org/html/2608.12372#bib.bib389); Barocas and Selbst,[2016](https://arxiv.org/html/2608.12372#bib.bib390); Selbstet al\.,[2019](https://arxiv.org/html/2608.12372#bib.bib391)\)\.
Methodologically, such reasoning archetypes can inform feature selection and the choice of hypothesis class\. Moreover, assessing and selecting among reasoning archetypes allows for greater user participation in the modeling process\. For example,Cousinset al\.\([2025](https://arxiv.org/html/2608.12372#bib.bib299)\)apply this principle to pairwise decision making, to factor out theoretically irrelevant aspects of decision making, which addresses \(1\) and \(3\) through explicit feature\-interaction modeling\.Cousinset al\.\([2024](https://arxiv.org/html/2608.12372#bib.bib298)\)similarly study welfare\-based optimization to structure*planning*and*reinforcement learning*, where rich vector\-valued feedback is used to model the impact of decisions on multiple parties, then a nonlinear*welfare concept*is optimized\. In both cases, interpretable*human\-relevant structure*is*built into the system*, and it is*incapable*of reasoning outside of this framework\.
While these works employ axiomatic analysis to impose*interpretable structured restrictions*on the hypothesis class, data\-guided approaches are likely to impose stronger constraints, and*both constraint types*may be used concurrently\. In particular, analytic restrictions*discard cognitively implausible models*with theory, and data\-guided approaches target plausible and likely models at the individual level\. One way to estimate these reasoning archetypes is through qualitative interviews\(Bonet and Geffner,[1996](https://arxiv.org/html/2608.12372#bib.bib293); Bendel,[2016](https://arxiv.org/html/2608.12372#bib.bib100); Keswaniet al\.,[2025a](https://arxiv.org/html/2608.12372#bib.bib286)\)\. Another way, commonly used in behavioral economics and computational cognitive science, is to infer them through carefully\-designed quantitative experiments\(Kahnemanet al\.,[1979](https://arxiv.org/html/2608.12372#bib.bib292); Bourginet al\.,[2019](https://arxiv.org/html/2608.12372#bib.bib45); Erevet al\.,[2017](https://arxiv.org/html/2608.12372#bib.bib291)\)\. Philosophical theories and characterizations of human behavior may also be helpful\(Kagan,[1988](https://arxiv.org/html/2608.12372#bib.bib81); Mongin,[1998](https://arxiv.org/html/2608.12372#bib.bib89); Lazar,[2017](https://arxiv.org/html/2608.12372#bib.bib87)\)\. Novel methods could also be developed to better characterize user reasoning archetypes\. Overall, to make progress on cognitive alignment, the field needs investment in both*methodology*and*scaled data collection*to infer decision\-making processes\. While the above works provide a template, further research should focus on both broad principles for constructing hypothesis classes for cognitive alignment, as well as specific constraints, which may be specific to a given domain or learning modality\.
##### Conflicts between stated reasoning processes and revealed preferences: Challenges or opportunities?
As discussed earlier, cognitive alignment will likely require combining multiple elicitation methods, including structured choices, interactive feedback, and self\-reports\. But what if these sources conflict? Consider a hiring manager who explicitly values educational diversity, but whose choices reveal much stronger weighting of degrees from elite universities\. The machine learning field has historically favored revealed preferences over stated preferences, but for users to trust and delegate to cognitively aligned AI, the system must reflect reasoning they endorse, not just patterns they exhibit\. Future research must determine how to surface such conflicts to users and resolve them while preserving trust\. There is little evidence currently available about how best to do this, but one strategy could be to use the interactive elicitation methods described earlier to present conflicts to the user through visualizations and text in a safe, anonymous environment\. The system could then ask the user why they think the conflicts exist, and request guidance about how users would prefer they be resolved \(e\.g\., should we update the*model*or their*decision*?\)\. Giving users agency over conflict resolution may help prevent users from feeling defensive and losing trust in the system\. Even if this initial strategy proves imperfect, lessons learned through these kinds of exchanges can inform and inspire more effective ways to resolve conflicts between stated decision strategies and their observed decision outcomes down the line\.
At the same time, discrepancies may create opportunities for self\-discovery or for revealing reasoning users wantavoided, including unwanted biases or heuristics expressed under stress or fatigue\. Interactive elicitation interfaces could let users flag reasoning to avoid — like reasoning they were unaware of until discrepancies between their elicited and stated preferences are highlighted — and the interface could transparently communicate how the model implements that avoidance\. Surfacing these conflicts might help users recognize patterns they wish to change; allowing interactive feedback strengthens trust by making the system’s alignment with their more idealized reasoning explicit\.
## 5Alternative Views
We advocate for a research agenda that develops scalable methods for cognitive alignment, but there are alternative viewpoints on the importance of cognitively\-aligned AI\.
One view is that arguing for cognitive alignment in the era of neural networks \(NNs\) is nonsensical, because NNs already approximate how the human brain works, so they are in a sense already aligned to human\-like thinking\. Proponents may argue that NNs’ biological inspiration and increasingly sophisticated reasoning capabilities indicate that they inherently operate at a similar level of abstraction to human cognition\. We counter that recent research demonstrates that even the highest\-performing NNs operate at fundamentally different levels of abstraction than human reasoning\(Shaniet al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib364); Bollepallyet al\.,[2026](https://arxiv.org/html/2608.12372#bib.bib365)\), so there is still a need for methodology that intentionally aligns AI systems to human thinking*explicitly at the cognitive level*, rather than at the analogous neurological level\.
Another view is that even if cognitive alignment is underdeveloped, it is not worth pursuing\. Proponents of this view have stated: “trust in AI should be primarily based on its objective performance” rather than the process by which that performance was reached\(Kortelinget al\.,[2021](https://arxiv.org/html/2608.12372#bib.bib359)\); thus, it suffices to align AI to people’s judgments without requiring the AI to align to how people think\(Broughel,[2024](https://arxiv.org/html/2608.12372#bib.bib366)\)\. In response, we first clarify that we agree accuracy is one of the most important criteria for both assessing alignment and engendering trust in AI, as empirical work shows\(Ahnet al\.,[2024](https://arxiv.org/html/2608.12372#bib.bib360); Hunsickeret al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib367)\)\. We also agree that cognitive\-alignment is not*always*needed for an AI system to be used and adopted effectively\. Our counterpoint is that*there remain important contexts where people want their AI to think like them*, or will benefit if it did \(particularly if bias or mistakes could be corrected, as discussed earlier\)\. Our proposition is that the ML field has yet to develop robust methodology for creating this kind of AI, and this gap limits the positive impact AI can have in vital AI applications\.
Another view accepts that cognitively\-aligned AI could be useful, but argues our research agenda is unrealistic because cognitively\-aligned AI can never reach the accuracy and performance levels of other alignment methods\. This view believes there is a cognitive alignment\-accuracy tradeoff, and the tradeoff is too costly to justify investment in cognitive alignment research and development\. We counter that, similar to arguments made byRudin \([2019](https://arxiv.org/html/2608.12372#bib.bib133)\)in the related but different context ofinterpretableAI, there is no principled reason cognitively\-aligned AI must be less accurate than other forms of alignment, and initial efforts suggest accuracy reductions do not always occur\(Cousinset al\.,[2025](https://arxiv.org/html/2608.12372#bib.bib299)\)\. Moreover, if we develop elicitation techniques that leverage humans’ ability to self\-report their reasoning, there is even potential for cognitive alignment to ultimately become*more*accurate and*more*data\-efficient than other alignment methods\. Avoiding research due to these perceived challenges risks missing real opportunities to develop methodologies that realize and scale the unique advantages of cognitive alignment to other ML methodologies\.
A variant of this view is that cognitively\-aligned AI is infeasible, because it would require too much data to align a model to a single person\. We suggest two strategies to mitigate this concern\. First, initial applications of cognitive alignment should focus on narrow use\-cases where AI delegates will plausibly be deployed, so that the space of decision\-making features that need to be learned for an accurate model is reduced\. Once accurate cognitively\-aligned models have been learned for enough contexts, general organizational rules or principles about how people make decisions may emerge that can be used to reduce the amount of data collection needed for accurate models in new domains\. Second, elicitation methods should leveragericher supervision signals, such as self\-reports and interactive feedback, to scale down the need for large volumes of behavioral data\. People can’t always explain their reasoning accurately, and may even misrepresent aspects of their reasoning, but empirical studies show that they still provide more accurate information than many other forms of elicitation\(Corneille and Gawronski,[2024](https://arxiv.org/html/2608.12372#bib.bib382)\), which can be used to drastically reduce data requirements\. Scalable elicitation will definitely be a challenge, but it should not be considered an insurmountable one that deters efforts to develop methods for creating AI that is cognitively\-aligned\.
## 6Conclusion
Our goal in this paper is to highlight underappreciated ways in which cognitive alignment can enable the adoption of many envisioned AI applications, and to motivate the development of methods that align AI systems with how individual users think — or want to think\. We emphasize that this line of work is intended to complement, not replace, existing alignment approaches\. However, our position is that it is critical for the ML field to appreciate that lack of cognitive alignment can function as a meaningful adoption barrier\. Without progress on it, many AI systems may see limited real\-world use regardless of their predictive performance, reducing the positive impact they would otherwise achieve\.
## Impact Statement
This paper presents cognitive alignment as one component of trustworthy AI\. Some of our survey questions address ethically sensitive topics, including medical triage and military decision\-making\. These questions were included to study attitudes toward decision delegation in high\-stakes scenarios, not to advocate for AI deployment in such settings\. Our findings indicate that people want AI to mirror their own reasoning in these contentious domains\. While these results suggest cognitive alignment may be necessary for trustworthy AI in sensitive applications, they do not suggest cognitive alignment is sufficient to justify using AI in those applications\. Ethical, legal, and practical constraints are also critical considerations to evaluate\.
## Acknowledgments
VK, BKN, CC, WSA, and JSB are grateful for the financial support from OpenAI and Duke University\. We are also grateful to Hoda Heidari for her insightful and invaluable feedback on various iterations of this draft\.
## References
- D\. Ahn, A\. Almaatouq, M\. Gulabani, and K\. Hosanagar \(2024\)Impact of Model Interpretability and Outcome Feedback on Trust in AI\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–25\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p3.1)\.
- S\. Amershi, M\. Cakmak, W\. B\. Knox, and T\. Kulesza \(2014\)Power to the People: The Role of Humans in Interactive Machine Learning\.AI magazine35\(4\),pp\. 105–120\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p3.1)\.
- M\. Amirizaniani, E\. Martin, M\. Sivachenko, A\. Mashhadi, and C\. Shah \(2024\)Do LLMs Exhibit Human\-Like Reasoning? Evaluating Theory of Mind in LLMs for Open\-Ended Responses\.arXiv preprint arXiv:2406\.05659\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p2.1)\.
- E\. Awad, S\. Dsouza, R\. Kim, J\. Schulz, J\. Henrich, A\. Shariff, J\. Bonnefon, and I\. Rahwan \(2018\)The Moral Machine Experiment\.Nature563\(7729\),pp\. 59–64\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- T\. A\. Bach, A\. Khan, H\. Hallock, G\. Beltrão, and S\. Sousa \(2024\)A Systematic Literature Review of User Trust in AI\-Enabled Systems: An HCI Perspective\.International Journal of Human–Computer Interaction40\(5\),pp\. 1251–1266\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p3.1)\.
- Y\. Bai, S\. Kadavath, S\. Kundu, A\. Askell, J\. Kernion, A\. Jones, A\. Chen, A\. Goldie, A\. Mirhoseini, C\. McKinnon,et al\.\(2022\)Constitutional AI: Harmlessness from AI Feedback\.arXiv preprint arXiv:2212\.08073\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- F\. Barez, T\. Wu, I\. Arcuschin, M\. Lan, V\. Wang, N\. Siegel, N\. Collignon, C\. Neo, I\. Lee, A\. Paren,et al\.\(2025\)Chain\-of\-Thought Is Not Explainability\.Preprint, alphaXiv,pp\. v1\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p2.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p3.1)\.
- S\. Barocas and A\. D\. Selbst \(2016\)Big Data’s Disparate Impact\.Calif\. L\. Rev\.104,pp\. 671\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p2.1)\.
- W\. Bechtel \(1994\)Levels of Description and Explanation in Cognitive Science\.Minds and Machines4\(1\),pp\. 1–25\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px1.p1.1)\.
- O\. Bendel \(2016\)Annotated Decision Trees for Simple Moral Machines\.In2016 AAAI Spring Symposium Series,Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- T\. W\. Bickmore and R\. W\. Picard \(2005\)Establishing and Maintaining Long\-Term Human\-Computer Relationships\.ACM Transactions on Computer\-Human Interaction \(TOCHI\)12\(2\),pp\. 293–327\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p3.1)\.
- S\. Bobek, P\. Korycińska, M\. Krakowska, M\. Mozolewski, D\. Rak, M\. Zych, M\. Wójcik, and G\. J\. Nalepa \(2025\)User\-Centric Evaluation of Explainability of AI with and for Humans: A Comprehensive Empirical Study\.International Journal of Human\-Computer Studies,pp\. 103625\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1),[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- K\. Boerstler, V\. Keswani, L\. Chan, J\. Schaich Borg, V\. Conitzer, H\. Heidari, and W\. Sinnott\-Armstrong \(2024\)On the Stability of Moral Preferences: A Problem with Computational Elicitation Methods\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,Vol\.7,pp\. 156–167\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p4.1)\.
- S\. Bollepally, A\. Sloman\-Moll, and T\. Yamauchi \(2026\)Can LLMs Interpret Figurative Language as Humans Do?: Surface\-Level vs Representational Similarity\.arXiv preprint arXiv:2601\.09041\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p2.1)\.
- B\. Bonet and H\. Geffner \(1996\)Arguing for Decisions: A Qualitative Model of Decision Making\.InProceedings of the Twelfth international conference on Uncertainty in artificial intelligence,pp\. 98–105\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- J\. Borders, A\. Leung, and M\. Condon \(2025\)A Framework for Identifying Key Decision\-Maker Attributes in Uncertain and Complex Environments\.In2025 IEEE conference on artificial intelligence \(CAI\),pp\. 1–5\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px3.p3.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px3.p4.1)\.
- D\. D\. Bourgin, J\. C\. Peterson, D\. Reichman, S\. J\. Russell, and T\. L\. Griffiths \(2019\)Cognitive Model Priors for Predicting Human Decisions\.InInternational conference on machine learning,pp\. 5133–5141\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- E\. E\. Bradford, I\. Jentzsch, and J\. Gomez \(2015\)From Self to Social Cognition: Theory of Mind Mechanisms and Their Relation to Executive Functioning\.Cognition138,pp\. 21–34\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- J\. Broughel \(2024\)Artificial Intelligence ’Explainability’ Is Overrated\.Note:Forbes, accessed 23 January 2026Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p3.1)\.
- V\. Chamola, V\. Hassija, A\. R\. Sulthana, D\. Ghosh, D\. Dhingra, and B\. Sikdar \(2023\)A Review of Trustworthy and Explainable Artificial Intelligence \(XAI\)\.IEEe Access11,pp\. 78994–79015\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1)\.
- T\. Chanda, K\. Hauser, S\. Hobelsberger, T\. Bucher, C\. N\. Garcia, C\. Wies, H\. Kittler, P\. Tschandl, C\. Navarrete\-Dechent, S\. Podlipnik,et al\.\(2024\)Dermatologist\-Like Explainable AI Enhances Trust and Confidence in Diagnosing Melanoma\.Nature Communications15\(1\),pp\. 524\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- H\. Chen, I\. C\. Covert, S\. M\. Lundberg, and S\. Lee \(2023\)Algorithms to Estimate Shapley Value Feature Attributions\.Nature Machine Intelligence5\(6\),pp\. 590–601\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- P\. F\. Christiano, J\. Leike, T\. Brown, M\. Martic, S\. Legg, and D\. Amodei \(2017\)Deep Reinforcement Learning from Human Preferences\.Advances in neural information processing systems30\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- M\. Colombo and M\. Knauff \(2020\)Editors’ Review and Introduction: Levels of Explanation in Cognitive Science: From Molecules to Culture\.Vol\.12,Wiley Online Library\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px1.p1.1)\.
- O\. Corneille and B\. Gawronski \(2024\)Self\-Reports Are Better Measurement Instruments Than Implicit Measures\.Nature Reviews Psychology3\(12\),pp\. 835–846\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p5.1)\.
- L\. Corti, R\. Oltmans, J\. Jung, A\. Balayn, M\. Wijsenbeek, and J\. Yang \(2024\)“It Is a Moving Process”: Understanding the Evolution of Explainability Needs of Clinicians in Pulmonary Medicine\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–21\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1),[§2](https://arxiv.org/html/2608.12372#S2.p4.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px1.p2.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- C\. Cousins, K\. Asadi, E\. Lobo, and M\. Littman \(2024\)On Welfare\-Centric Fair Reinforcement Learning\.InReinforcement Learning Conference,Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p3.1)\.
- C\. Cousins, V\. Keswani, V\. Contizer, H\. Heidari, J\. Schaich Borg, and Sinnott\-Armstrong \(2025\)Towards Cognitively\-Faithful Decision\-Making Models to Improve AI Alignment\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p3.1),[§5](https://arxiv.org/html/2608.12372#S5.p4.1)\.
- R\. de Brito Duarte, F\. Correia, P\. Arriaga, and A\. Paiva \(2023\)AI Trust: Can Explainable AI Enhance Warranted Trust?\.Human Behavior and Emerging Technologies2023\(1\),pp\. 4637678\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1)\.
- C\. Dwork, M\. Hardt, T\. Pitassi, O\. Reingold, and R\. Zemel \(2012\)Fairness Through Awareness\.InProceedings of the 3rd innovations in theoretical computer science conference,pp\. 214–226\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p2.1)\.
- U\. Ehsan, S\. Passi, Q\. V\. Liao, L\. Chan, I\. Lee, M\. Muller, and M\. O\. Riedl \(2024\)The Who in XAI: How AI Background Shapes Perceptions of AI Explanations\.InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems,pp\. 1–32\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- I\. Erev, E\. Ert, O\. Plonsky, D\. Cohen, and O\. Cohen \(2017\)From Anomalies to Forecasts: Toward a Descriptive Model of Decisions Under Risk, Under Ambiguity, and from Experience\.Psychological review124\(4\),pp\. 369\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- R\. Farrell \(2025\)Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p2.1)\.
- R\. Freedman, J\. Schaich Borg, W\. Sinnott\-Armstrong, J\. P\. Dickerson, and V\. Conitzer \(2020\)Adapting a Kidney Exchange Algorithm to Align with Human Values\.Artificial Intelligence283,pp\. 103261\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1)\.
- I\. Gabriel \(2020\)Artificial Intelligence, Values, and Alignment\.Minds and machines30\(3\),pp\. 411–437\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p5.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1)\.
- G\. Gigerenzer and W\. Gaissmaier \(2011\)Heuristic Decision Making\.Annual review of psychology62\(2011\),pp\. 451–482\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p1.1)\.
- C\. Gonzalez and H\. Heidari \(2025\)A Cognitive Approach to Human–AI Complementarity in Dynamic Decision\-Making\.Nature Reviews Psychology,pp\. 1–15\.Cited by:[§1](https://arxiv.org/html/2608.12372#S1.p3.1)\.
- N\. Grgić\-Hlača, C\. Castelluccia, and K\. P\. Gummadi \(2022\)Taking Advice from \(Dis\) Similar Machines: The Impact of Human\-Machine Similarity on Machine\-Assisted Decision\-Making\.InProceedings of the AAAI Conference on Human Computation and Crowdsourcing,Vol\.10,pp\. 74–88\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- N\. Grgic\-Hlaca, E\. M\. Redmiles, K\. P\. Gummadi, and A\. Weller \(2018\)Human Perceptions of Fairness in Algorithmic Decision Making: A Case Study of Criminal Risk Prediction\.InProceedings of the 2018 world wide web conference,pp\. 903–912\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1)\.
- M\. Y\. Guan, M\. Joglekar, E\. Wallace, S\. Jain, B\. Barak, A\. Helyar, R\. Dias, A\. Vallone, H\. Ren, J\. Wei,et al\.\(2024\)Deliberative Alignment: Reasoning Enables Safer Language Models\.arXiv preprint arXiv:2412\.16339\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- C\. D\. Güss \(2018\)What is going through your mind? thinking aloud as a method in cross\-cultural psychology\.Frontiers in psychology9,pp\. 1292\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px2.p1.1)\.
- D\. Hadfield\-Menell, A\. Dragan, P\. Abbeel, and S\. Russell \(2017\)The Off\-Switch Game\.InInternational Joint Conferences on Artificial Intelligence Organization,Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- P\. Hamm, M\. Klesel, P\. Coberger, and H\. F\. Wittmann \(2023\)Explanation Matters: An Experimental Study on Explainable AI\.Electronic Markets33\(1\),pp\. 17\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1)\.
- T\. Hastie and R\. Tibshirani \(1986\)Generalized Additive Models\.Statistical science1\(3\),pp\. 297–310\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- T\. Hunsicker, C\. J\. König, and M\. Langer \(2025\)Investigating Choices Regarding the Accuracy\-Transparency Trade\-Off of AI\-Based Systems Across Contexts\.Computers in Human Behavior: Artificial Humans,pp\. 100216\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p3.1)\.
- I\. Hurley, R\. Paleja, A\. Suh, J\. D\. Peña, and H\. C\. Siu \(2024\)STL: Still Tricky Logic \(for System Validation, Even When Showing Your Work\)\.Advances in Neural Information Processing Systems37,pp\. 119099–119122\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- J\. M\. Imperial, M\. D\. Jones, and H\. T\. Madabushi \(2025\)Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance\.arXiv preprint arXiv:2503\.04736\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- C\. M\. Johnston, P\. Vossler, S\. Blessenohl, and P\. Vayanos \(2023\)Deploying a Robust Active Preference Elicitation Algorithm on MTurk: Experiment Design, Interface, and Evaluation for COVID\-19 Patient Prioritization\.InProceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization,pp\. 1–10\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1)\.
- S\. Kagan \(1988\)The Additive Fallacy\.Ethics99\(1\),pp\. 5–31\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- D\. Kahneman, A\. Tversky,et al\.\(1979\)Prospect Theory: An Analysis of Decision Under Risk\.Econometrica47\(2\),pp\. 363–391\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- T\. Kaufmann, P\. Weng, V\. Bengs, and E\. Hüllermeier \(2023\)A Survey of Reinforcement Learning from Human Feedback\.arXiv preprint arXiv:2312\.1492510\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- V\. Keswani, V\. Conitzer, W\. Sinnott\-Armstrong, B\. K\. Nguyen, H\. Heidari, and J\. Schaich Borg \(2025a\)Can AI Model the Complexities of Human Moral Decision\-Making? A Qualitative Study of Kidney Allocation Decisions\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,pp\. 1–17\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- V\. Keswani, C\. Cousins, B\. Nguyen, V\. Conitzer, H\. Heidari, J\. Schaich Borg, and W\. Sinnott\-Armstrong \(2025b\)Moral Change or Noise? On Problems of Aligning AI with Temporally Unstable Human Feedback\.arXiv preprint arXiv:2511\.10032\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p4.1)\.
- R\. Kim, M\. Kleiman\-Weiner, A\. Abeliuk, E\. Awad, S\. Dsouza, J\. B\. Tenenbaum, and I\. Rahwan \(2018\)A Computational Model of Commonsense Moral Decision Making\.InProceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society,pp\. 197–203\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1)\.
- J\. Kleinberg, J\. Ludwig, S\. Mullainathan, and M\. Raghavan \(2024\)The Inversion Problem: Why Algorithms Should Infer Mental State and Not Just Predict Behavior\.Perspectives on Psychological Science19\(5\),pp\. 827–838\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p1.1)\.
- N\. Kolt, N\. Caputo, J\. Boeglin, C\. O’Keefe, R\. Bommasani, S\. Casper, M\. Cuéllar, N\. Feldman, I\. Gabriel, G\. K\. Hadfield,et al\.\(2026\)Legal Alignment for Safe and Ethical AI\.arXiv preprint arXiv:2601\.04175\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- T\. Korbak, M\. Balesni, E\. Barnes, Y\. Bengio, J\. Benton, J\. Bloom, M\. Chen, A\. Cooney, A\. Dafoe, A\. Dragan,et al\.\(2025\)Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety\.arXiv preprint arXiv:2507\.11473\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p2.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p3.1)\.
- J\. E\. Korteling, G\. C\. van de Boer\-Visschedijk, R\. A\. Blankendaal, R\. C\. Boonekamp, and A\. R\. Eikelboom \(2021\)Human\-Versus Artificial Intelligence\.Frontiers in artificial intelligence4,pp\. 622364\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p3.1)\.
- W\. Kweon, S\. Kang, J\. Hwang, and H\. Yu \(2020\)Deep Rating Elicitation for New Users in Collaborative Filtering\.InProceedings of The Web Conference 2020,pp\. 2810–2816\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1)\.
- Y\. Kyrychenko, K\. Zhou, E\. Bogucka, and D\. Quercia \(2025\)C3AI: Crafting and Evaluating Constitutions for Constitutional AI\.InProceedings of the ACM on Web Conference 2025,pp\. 3204–3218\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p3.1)\.
- S\. Lazar \(2017\)Deontological Decision Theory and Agent\-Centered Options\.Ethics127\(3\),pp\. 579–609\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- M\. K\. Lee, D\. Kusbit, A\. Kahng, J\. T\. Kim, X\. Yuan, A\. Chan, D\. See, R\. Noothigattu, S\. Lee, A\. Psomas,et al\.\(2019\)WeBuildAI: Participatory Framework for Algorithmic Governance\.Proceedings of the ACM on human\-computer interaction3\(CSCW\),pp\. 1–35\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1)\.
- Z\. C\. Lipton \(2018\)The Mythos of Model Interpretability: In Machine Learning, the Concept of Interpretability Is Both Important and Slippery\.Queue16\(3\),pp\. 31–57\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1)\.
- J\. Lopez, C\. Textor, C\. Lancaster, B\. Schelble, G\. Freeman, R\. Zhang, N\. McNeese, and R\. Pak \(2024\)The Complex Relationship of AI Ethics and Trust in Human–AI Teaming: Insights from Advanced Real\-World Subject Matter Experts\.AI and Ethics4\(4\),pp\. 1213–1233\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1),[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- J\. McVay, E\. J\. De Visser, B\. Pippin, A\. Mani, J\. N\. Hyde, and N\. Kman \(2025\)Trust in Aligned AI Decision Makers\.In2025 IEEE conference on artificial intelligence \(CAI\),pp\. 1–4\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p5.1)\.
- P\. Mongin \(1998\)Expected Utility Theory\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p4.1)\.
- S\. Morandini, F\. Fraboni, M\. Hall, S\. Quintana\-Amate, and L\. Pietrantoni \(2025\)User Perspectives on AI Explainability in Aerospace Manufacturing: A Card\-Sorting Study\.Frontiers in Organizational Psychology3,pp\. 1538438\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- R\. Noothigattu, S\. ’\. S\. Gaikwad, E\. Awad, S\. Dsouza, I\. Rahwan, P\. Ravikumar, and A\. D\. Procaccia \(2018\)A Voting\-Based System for Ethical Decision Making\.arXiv\.Note:arXiv:1709\.06692 \[cs\]External Links:[Document](https://dx.doi.org/10.48550/arXiv.1709.06692)Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1)\.
- J\. L\. Orquin and S\. M\. Loose \(2013\)Attention and choice: a review on eye movements in decision making\.Acta psychologica144\(1\),pp\. 190–206\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px2.p1.1)\.
- L\. Orseau and M\. Armstrong \(2016\)Safely Interruptible Agents\.InConference on Uncertainty in Artificial Intelligence,Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- J\. C\. Peterson, D\. D\. Bourgin, M\. Agrawal, D\. Reichman, and T\. L\. Griffiths \(2021\)Using Large\-Scale Experiments and Machine Learning to Discover Theories of Human Decision\-Making\.Science372\(6547\),pp\. 1209–1214\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px5.p1.1)\.
- E\. K\. Phillips and B\. F\. Malle \(2025\)The Power of Justifications to Repair Human\-Robot Trust, Even Under Moral Disagreement\.Scientific Reports15\(1\),pp\. 34706\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p5.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct Preference Optimization: Your Language Model Is Secretly a Reward Model\.Advances in Neural Information Processing Systems36,pp\. 53728–53741\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1)\.
- M\. Riveiro and S\. Thill \(2022\)The Challenges of Providing Explanations of AI Systems When They Do Not Behave Like Users Expect\.InProceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization,pp\. 110–120\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1)\.
- C\. Rudin \(2019\)Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead\.Nature machine intelligence1\(5\),pp\. 206–215\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1),[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p2.1),[§5](https://arxiv.org/html/2608.12372#S5.p4.1)\.
- W\. Samek, T\. Wiegand, and K\. Müller \(2017\)Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models\.arXiv preprint arXiv:1708\.08296\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- S\. Samuel, G\. G\. Cole, and M\. J\. Eacott \(2020\)Two Independent Sources of Difficulty in Perspective\-Taking/Theory of Mind Tasks\.Psychonomic bulletin & review27\(6\),pp\. 1341–1347\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- S\. N\. Saw, Y\. Y\. Yan, and K\. H\. Ng \(2025\)Current Status and Future Directions of Explainable Artificial Intelligence in Medical Imaging\.European journal of radiology183,pp\. 111884\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1),[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px1.p2.1)\.
- S\. Schröder, T\. Morgenroth, U\. Kuhl, V\. Vaquet, and B\. Paaßen \(2025\)Large Language Models Do Not Simulate Human Psychology\.arXiv preprint arXiv:2508\.06950\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p2.1)\.
- M\. Schulte\-Mecklenbeck, A\. Kühberger, and R\. Ranyard \(2011\)The role of process data in the development and testing of process models of judgment and decision making\.Judgment and Decision making6\(8\),pp\. 733–739\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px2.p1.1)\.
- S\. Schulze\-Weddige and T\. Zylowski \(2021\)User Study on the Effects Explainable AI Visualizations on Non\-Experts\.InInternational Conference on ArtsIT, Interactivity and Game Creation,pp\. 457–467\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- A\. D\. Selbst, D\. Boyd, S\. A\. Friedler, S\. Venkatasubramanian, and J\. Vertesi \(2019\)Fairness and Abstraction in Sociotechnical Systems\.InProceedings of the conference on fairness, accountability, and transparency,pp\. 59–68\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px5.p2.1)\.
- C\. Shani, L\. Soffer, D\. Jurafsky, Y\. LeCun, and R\. Shwartz\-Ziv \(2025\)From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning\.arXiv preprint arXiv:2505\.17117\.Cited by:[§5](https://arxiv.org/html/2608.12372#S5.p2.1)\.
- D\. Shin \(2021\)The Effects of Explainability and Causability on Perception, Trust, and Acceptance: Implications for Explainable AI\.International journal of human\-computer studies146,pp\. 102551\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1)\.
- K\. Simonyan, A\. Vedaldi, and A\. Zisserman \(2013\)Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps\.arXiv preprint arXiv:1312\.6034\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- H\. C\. Siu, A\. Suh, N\. Smith, and I\. Hurley \(2025\)“Explainable” AI Has Some Explaining to Do\.MIT Case Studies in Social and Ethical Responsibilities of Computing\(Winter 2025\)\.Note:https://mit\-serc\.pubpub\.org/pub/pt5lplzbCited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- M\. Srivastava, H\. Heidari, and A\. Krause \(2019\)Mathematical Notions vs\. Human Perception of Fairness: A Descriptive Approach to Fairness for Machine Learning\.InProceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining,pp\. 2459–2468\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p1.1)\.
- A\. Summerville, E\. J\. de Visser, J\. McVay, L\. Martí, A\. Leung, and C\. Widmer \(2025\)Alignment in Decision\-Making Attributes Predicts Trust and Delegation to AI Systems\.Journal of Cognitive Engineering and Decision Making,pp\. 15553434251390012\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p4.1),[§2](https://arxiv.org/html/2608.12372#S2.p5.1)\.
- H\. Sun, Y\. Shen, and J\. Ton \(2024\)Rethinking Bradley\-Terry Models in Preference\-Based Reward Modeling: Foundations, Theory, and Alternatives\.arXiv preprint arXiv:2411\.04991\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- S\. Tonekaboni, S\. Joshi, M\. D\. McCradden, and A\. Goldenberg \(2019\)What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use\.InMachine learning for healthcare conference,pp\. 359–380\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p2.1),[§2](https://arxiv.org/html/2608.12372#S2.p4.1)\.
- A\. M\. Turner, D\. Hadfield\-Menell, and P\. Tadepalli \(2020\)Conservative Agency via Attainable Utility Preservation\.InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society,pp\. 385–391\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p3.1)\.
- M\. Turpin, J\. Michael, E\. Perez, and S\. Bowman \(2023\)Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain\-of\-Thought Prompting\.Advances in Neural Information Processing Systems36,pp\. 74952–74965\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p2.1)\.
- M\. Vaccaro, A\. Almaatouq, and T\. Malone \(2024\)When Combinations of Humans and AI Are Useful: A Systematic Review and Meta\-Analysis\.Nature Human Behaviour8\(12\),pp\. 2293–2303\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- W\. J\. Von Eschenbach \(2021\)Transparency and the Black Box Problem: Why We Do Not Trust AI\.Philosophy & Technology34\(4\),pp\. 1607–1622\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1)\.
- B\. Wang, S\. Min, X\. Deng, J\. Shen, Y\. Wu, L\. Zettlemoyer, and H\. Sun \(2023\)Towards Understanding Chain\-of\-Thought Prompting: An Empirical Study of What Matters\.InProceedings of the 61st annual meeting of the association for computational linguistics \(volume 1: Long papers\),pp\. 2717–2739\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- S\. Wang, Q\. Wang, and J\. Zhao \(2020\)Deep Neural Networks for Choice Analysis: Extracting Complete Economic Information for Interpretation\.Transportation Research Part C: Emerging Technologies118,pp\. 102701\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px3.p1.1)\.
- X\. Wang, J\. Wei, D\. Schuurmans, Q\. V\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. Zhou \(2022a\)Self\-Consistency Improves Chain of Thought Reasoning in Language Models\.InThe Eleventh International Conference on Learning Representations,Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- Z\. J\. Wang, A\. Kale, H\. Nori, P\. Stella, M\. E\. Nunnally, D\. H\. Chau, M\. Vorvoreanu, J\. Wortman Vaughan, and R\. Caruana \(2022b\)Interpretability, Then What? Editing Machine Learning Models to Reflect Human Knowledge and Values\.InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 4132–4142\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p2.1)\.
- C\. Warren, A\. P\. McGraw, and L\. Van Boven \(2011\)Values and Preferences: Defining Preference Construction\.Wiley Interdisciplinary Reviews: Cognitive Science2\(2\),pp\. 193–205\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p4.1)\.
- B\. Wilder, E\. Horvitz, and E\. Kamar \(2020\)Learning to Complement Humans\.arXiv preprint arXiv:2005\.00582\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p3.1)\.
- C\. Wirth, R\. Akrour, G\. Neumann, and J\. Fürnkranz \(2017\)A Survey of Preference\-Based Reinforcement Learning Methods\.Journal of Machine Learning Research18\(136\),pp\. 1–46\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px1.p2.1)\.
- N\. A\. Wondimu, C\. Buche, and U\. Visser \(2022\)Interactive Machine Learning: A State of the Art Review\.arXiv preprint arXiv:2207\.06196\.Cited by:[§4](https://arxiv.org/html/2608.12372#S4.SS0.SSS0.Px4.p3.1)\.
- F\. Xiao and X\. X\. Wang \(2025\)Evaluating the Ability of Large Language Models to Predict Human Social Decisions\.Scientific Reports15\(1\),pp\. 32290\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p3.1)\.
- L\. Zhang, M\. Grabmair, M\. Gray, and K\. Ashley \(2025\)Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning\.arXiv preprint arXiv:2510\.08710\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px4.p2.1)\.
- T\. Zhi\-Xuan, M\. Carroll, M\. Franklin, and H\. Ashton \(2025\)Beyond Preferences in AI Alignment: T\. Zhi\-Xuan Et Al\.\.Philosophical Studies182\(7\),pp\. 1813–1863\.Cited by:[§2](https://arxiv.org/html/2608.12372#S2.p5.1)\.
- J\. Zhu, J\. C\. Peterson, B\. Enke, and T\. L\. Griffiths \(2025a\)Capturing the Complexity of Human Strategic Decision\-Making with Machine Learning\.Nature Human Behaviour,pp\. 1–7\.Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px5.p1.1)\.
- J\. Zhu, H\. Xie, D\. Arumugam, R\. C\. Wilson, and T\. L\. Griffiths \(2025b\)Using Reinforcement Learning to Train Large Language Models to Explain Human Decisions\.CoRRabs/2505\.11614\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2505.11614),2505\.11614Cited by:[§3](https://arxiv.org/html/2608.12372#S3.SS0.SSS0.Px5.p1.1)\.
## Appendix AAdditional Study Details and Results
In this section, we provide additional details on the survey methodology, data processing, and results that were excluded from the main body\.
### A\.1Study Details
#### A\.1\.1Study Methodology
Participants were recruited to take part in an online survey using Prolific \. 150 participants took part in this study\. All participants were compensated at the rate of $12/hr\. The survey methodology was approved by a university Institutional Review Board \(IRB\)\. Aggregate demographics are noted in Appendix[A\.1\.2](https://arxiv.org/html/2608.12372#A1.SS1.SSS2)\.
##### Study design\.
Participants read descriptions of AI systems that varied in their decision\-making processes and the observability of those processes\. For the first part of the survey, they were presented with descriptions of “Human\-Reasoning AI”, “Machine\-Reasoning AI”, and “Process\-Hidden” AI\. They then indicated their preferences for these AI types across different tasks\. For the second and third parts of the survey, we presented expanded properties of different AI systems and asked participants to rate the desirability and impact of these properties for trust in AI systems\. All survey questions are provided in Appendix[B](https://arxiv.org/html/2608.12372#A2)\.
#### A\.1\.2Participant Demographics
Demographic distribution of the overall participant pool along self\-reported age, race, and gender was as follows:
- •Age: 10% b/w 18\-30, 55% b/w 31\-50, 35% 51\+
- •Race: 75% White, 8% Black, 5% Asian, 4% Hispanic, 8% Other
- •Gender: 49% Female, 49% Male, 2% Other
Beyond these attributes, we also asked participants to provide additional demographic information\.
- •Education level: 12% high school, 16% some college, 10% associate’s degree, 43% bachelor’s degree, 14% master’s degree, 3% doctorate, 2% other;
- •Employment: 55% employed full\-time, 11% retired, 10% self\-employed, 9% part\-time, 15% other;
- •Religion: 51% Christian, 37% None, 2% Islam, 1% Judaism, 9% Other;
- •Social political orientation \(1 to 7, ranging from “extremely liberal” to “extremely conservative”\)” 3\.8±\\pm0\.3;
- •Economic and fiscal political orientation \(0 to 8, ranging from “extremely liberal” to “extremely conservative”\): 3\.78±\\pm0\.3;
Mean time to complete the survey was around 22 minutes\. We excluded 8 participants whose time\-on\-task was in the bottom or top 2\.5 percentiles of the population\.
### A\.2Additional Results
##### Participants’ preference for different AIs across tasks\.
Figure[A4](https://arxiv.org/html/2608.12372#A1.F4)presents the preference distribution of all possible responses to 16 AI use cases \(extending Figure[1](https://arxiv.org/html/2608.12372#S1.F1)in the main text\)\. In addition to the prevalent preference for Human\-Reasoning AI discussed in the main text, we see a preference for Machine\-Reasoning AI in software troubleshooting, investment risk analysis, and flagging fraudulent credit card activity, and a preference for Process\-Hidden AI in scheduling meetings and predicting the weather\.
Figure A4:Participants’ preference for AI types across all sixteen domains\.Figure A5:Participants’ ranking of different AI systems for autonomous vehicles and kidney allocation\.
##### Ranking various AI properties\.
For autonomous vehicles and kidney allocation, participants were asked to rank different AI systems so that the one “you would most prefer to use in these types of life and death decisions is at the top and the system you would least prefer is at the bottom\.” The following six kinds of AI systems were presented as options to the participants for this question \(text here is abbreviated; see full text in Appendix[B](https://arxiv.org/html/2608.12372#A2)\)\.
- - –A Process\-Hidden AI that is unmodifiable: This AI cannot tell you what process it will use to make its decisions in these types of life\-or\-death decisions\. You cannot modify the reasoning behind the AI’s decision process\. - –A Machine\-Reasoning AI that is hard to understand and unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, but the explanations are very difficult for you to understand because the AI’s reasoning process is so different from how you typically think\. You cannot modify the reasoning behind the AI’s decision process\. - –A Machine\-Reasoning AI that is easy to understand, but unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions and you can easily understand the explanation, but the reasoning process the AI uses is very different from the one you would want to use in similar situations if you had sufficient time and information\. You cannot modify the reasoning behind the AI’s decision process\. - –A Machine\-Reasoning AI that is easy to understand and modifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, and you can easily understand the explanation\. The reasoning process the AI uses will remain fundamentally different from the one you would want to use in similar situations if you had sufficient time and information, but you can modify the AI’s reasoning process\. - –A Human\-Reasoning AI that is easy to understand, but unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, you can easily understand the explanation, and the AIs reasoning process closely mirrors the way you would want to make decisions in these contexts if you had sufficient time and information\. You cannot modify the reasoning behind the AI’s decision process\. - –A Human\-Reasoning AI that is easy to understand and modifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, you can easily understand the explanation, and the AIs reasoning process closely mirrors the way you would want to make decisions in these contexts if you had sufficient time and information\. You can modify the AI’s reasoning process\.
The aggregate rankings across all participants are presented in Figure[A5](https://arxiv.org/html/2608.12372#A1.F5)\. We see that, for most participants, the first choice was Human\-Reasoning AI that was easy to understand and modifiable; the second was Human\-Reasoning AI that was easy to understand but unmodifiable; and the third was Machine\-Reasoning AI that was easy to understand and modifiable\.
##### Common Themes in High\- and Low\-Stakes Applications
We asked participants to list certain domains or tasks that they consider high\-stakes or low\-stakes\. We calculated document frequency, defined as the number of participants who mentioned a given word at least once, to characterize the prevalence of mentioned themes in responses to these questions\.
1. 1\.Top 5 themes for low\-stakes AI applications:movie\(28\),restaurant\(27\),dinner\(22\),watch\(19\),eat\(18\)
2. 2\.Top 5 themes for high\-stakes AI applications:medical\(55\),life\(22\),treatment\(18\),car\(17\),health\(16\)
##### Differences between the impact of various AI qualities in high\-stakes vs low\-stakes domains
\. For high\-stakes and low\-stakes domains, we asked participants the extent to which a number of qualities qualities impacted their likelihood of trusting an AI system\. As Figure[3](https://arxiv.org/html/2608.12372#S3.F3)shows, most qualities are rated more impactful in high\-stakes domains vs low\-stakes domains\. Here, we present the exact differences in impact rating and statistical significance results for these differences\. Assume the numeric rating ranges from 0 \(no impact\) to 5 \(very strong impact\)\.
- •Property: “High accuracy and reliability”\. Mean difference in impact rating across high\-stakes vs low\-stakes domains for this property was0\.650\.65, with95%95\\%bootstrap CI:\[0\.44,0\.85\]\[0\.44,0\.85\]\. Wilcoxon test statistic:412\.0412\.0;p<0\.001p<0\.001\.
- •Property: “Provides a clear explanation that is easily understandable”\. Mean difference in impact rating across high\-stakes vs low\-stakes domains for this property was0\.800\.80, with95%95\\%bootstrap CI:\[0\.57,1\.05\]\[0\.57,1\.05\]\. Wilcoxon test statistic:707\.5707\.5;p<0\.001p<0\.001\.
- •Property: “Decides in a similar way to how you would want”\. Mean difference in impact rating across high\-stakes vs low\-stakes domains for this property was0\.370\.37, with95%95\\%bootstrap CI:\[0\.11,0\.62\]\[0\.11,0\.62\]\. Wilcoxon test statistic:1442\.51442\.5;p=0\.002p=0\.002\.
- •Property: “Corrects for systematic error or biases”\. Mean difference in impact rating across high\-stakes vs low\-stakes domains for this property was0\.760\.76, with95%95\\%bootstrap CI:\[0\.51,0\.99\]\[0\.51,0\.99\]\. Wilcoxon test statistic:968\.5968\.5;p<0\.001p<0\.001\.
- •Property: “Decision strategy can be adjusted or corrected in advance”\. Mean difference in impact rating across high\-stakes vs low\-stakes domains for this property was0\.550\.55, with95%95\\%bootstrap CI:\[0\.33,0\.79\]\[0\.33,0\.79\]\. Wilcoxon test statistic:792\.5792\.5;p<0\.001p<0\.001\.
##### Regression of AI attitude outcomes on AI/tech familiarity and demographic attributes\.
Table[1](https://arxiv.org/html/2608.12372#A1.T1)presents the regression results of various outcomes on AI/tech familiarity and demographic attributes\. For tech use, the participant is asked “\[I\]n the last six months, how often have you used any of the following?” for the following tools: an internet search program, social media, a smart speaker/doorbell \(or other home devices\), a self\-driving car, a coding console for software development or writing computer programs\. Thetech use scorevariable is then computed by averaging the respondents’ scores across all of the above options\. For AI use, the participant is similarly asked “\[I\]n the last six months, how often have you used Artificial Intelligence \(AI\) to assist you with the following tasks in your personal or work life?” for the following purposes: write or edit emails, generate or edit written content for something other than emails, generate images or movies, research and learn about topics, automate tasks, generate ideas, and transcribe or summarize meetings\. TheAI use scorevariable is computed by averaging the respondents’ scores across all of the above options\. We also simplified demographic variables to have a sufficient number of participants per variable value\. Specifically, we simplify ethnicity and gender to be binary, and education levels to high school or less, some college, bachelor’s, or graduate\. Future work can explore variations in outcomes across more fine\-grained demographic groupings\.
From Table[1](https://arxiv.org/html/2608.12372#A1.T1), we see some significant associations between AI/tech use and some outcomes\. Participants with higher AI use score were more likely to choose ‘Yes’ when asked if they could imagine scenarios where Human\-Reasoning AI would be preferable over other options\. This suggests that people who have had more frequent interactions with AI are more likely to prefer Human\-Reasoning AI\. However, there wasn’t any significant association between higher AI use and choosing Human\-Reasoning for the list of tasks presented to them, possibly suggesting that the Human\-Reasoning AI applications that these participants had in mind differ from the tasks presented to them in Q2\.
Table 1:Regression of various outcomes over self\-reported demographic and AI/tech use scores\. The first outcome, in column one, corresponds to whether or not the participant chose “Yes” for Q1 in the survey\. The second outcome corresponds to the fraction of tasks \(among the 16 presented\) where the participant chose “Human\-Reasoning AI”\. The remaining columns concern participants’ impact rating for how impactful different properties are across high\-stakes and low\-stakes domains\. The outcome is the difference in their impact rating high\-stakes domains vs low\-stakes domains\. The five properties correspond to the following: \(P1\) the AI implements its programmed decision\-strategy with high accuracy and reliability; \(P2\) the AI provides a clear explanation for how it will make its decisions that you can easily understand and see in advance of using it; \(P3\) the AI makes decisions in a similar way to how you would if you had sufficient time and information; \(P4\) the AI corrects for systematic errors or biases that you would be prone to in the same situations; \(P5\) the AI’s decision strategy can be adjusted or corrected by you to better reflect the decision\-making process you want\.
## Appendix BFull List of Survey Questions
1. 1\.If the AIs were equally accurate \(that is, they perform equally well overall at making correct recommendations or decisions, on average\), can you imagine any scenarios in which you would prefer Human\-Reasoning AI over Process\-Hidden AI or Machine\-Reasoning AI? 1. \(a\)If Yes→\\rightarrowWhat are those scenarios in which you would prefer Human\-Reasoning AI? Please describe in a few words\.
2. 2\.For each of the AI use cases below, please indicate which kind of AI you would prefer to rely on if the AIs were equally accurate overall \(meaning they perform equally well, on average, at making correct recommendations or decisions\)\. 1. \(a\)Rewriting an email to make it clearer 2. \(b\)Troubleshooting a software error on your computer 3. \(c\)Predicting the weather for your vacation next week 4. \(d\)Scheduling meetings for your team 5. \(e\)Deciding whether you qualify for a loan or mortgage 6. \(f\)Determining your eligibility for parole or bail 7. \(g\)Recommending a restaurant for your birthday dinner 8. \(h\)Evaluating which investment option has the best statistical risk profile 9. \(i\)Screening job applications for a position at your company 10. \(j\)Determining which dying patient should receive a ventilator when there are not enough available for everyone who needs one 11. \(k\)Determining which patient with kidney disease should receive a transplant when there are not enough kidneys available for everyone who needs one 12. \(l\)Providing advice about a moral dilemma you are facing 13. \(m\)Directing a military missile to its target 14. \(n\)Deciding who should be targeted by a military missile 15. \(o\)Flagging a potentially fraudulent transaction on your credit card 16. \(p\)Deciding which distress signals a search\-and\-rescue team should prioritize after an emergency when it cannot respond to all simultaneously
3. 3\.Imagine you are buying an autonomous vehicle for your family\. In rare situations, the vehicle may face an unavoidable accident where any action—including continuing on your present course at your present speed—will result in serious harm to someone\. In these moments, the AI running the vehicle must decide who to prioritize: for example, the passengers in your vehicle versus a cyclist who falls in front of the car unexpectedly, or one group of pedestrians versus another group of pedestrians when your tire blows and the car veers onto a crowded sidewalk\. In the context of making these life\-or\-death decisions, how desirable would you personally find it for your autonomous vehicle’s AI to have each of the following qualities? 1. \(a\)The AI implements its programmed decision\-strategy with high accuracy and reliability 2. \(b\)The AI provides a clear explanation for how it will make its decisions that you can easily understand and see in advance of encountering unavoidable accidents 3. \(c\)The AI provides explanations that truthfully describe how it actually makes its decisions 4. \(d\)The AI makes decisions in a similar way to how you would if you had sufficient time and information 5. \(e\)The AI corrects for systematic errors or biases that you would be prone to in the same situations 6. \(f\)The AI’s decision strategy can be adjusted or corrected by you in advance to better reflect the decision\-making process you want, if needed
4. 4\.Please rank the following AI systems according to which you would most prefer your autonomous vehicle to use in these life\-or\-death decisions\. Put your most preferred system at the top and your least preferred at the bottom\. Assume that the AI systems in all options have equal accuracy overall in how well their decisions match the decisions you would make in the same scenarios, if you were given sufficient time and information\. However, the processes that the AIs use to arrive at those decisions vary, as does the potential for you to change or personalize those processes\. 1. \(a\)A Process\-Hidden AI that is unmodifiable: This AI cannot tell you what process it will use to make its decisions in these types of life\-or\-death decisions\. You cannot modify the reasoning behind the AI’s decision process\. 2. \(b\)A Machine\-Reasoning AI that is hard to understand and unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, but the explanations are very difficult for you to understand because the AI’s reasoning process is so different from how you typically think\. You cannot modify the reasoning behind the AI’s decision process\. 3. \(c\)A Machine\-Reasoning AI that is easy to understand, but unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions and you can easily understand the explanation, but the reasoning process the AI uses is very different from the one you would want to use in similar situations if you had sufficient time and information\. You cannot modify the reasoning behind the AI’s decision process\. 4. \(d\)A Machine\-Reasoning AI that is easy to understand and modifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, and you can easily understand the explanation\. The reasoning process the AI uses will remain fundamentally different from the one you would want to use in similar situations if you had sufficient time and information, but you can modify the AI’s reasoning process\. 5. \(e\)A Human\-Reasoning AI that is easy to understand, but unmodifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, you can easily understand the explanation, and the AIs reasoning process closely mirrors the way you would want to make decisions in these contexts if you had sufficient time and information\. You cannot modify the reasoning behind the AI’s decision process\. 6. \(f\)A Human\-Reasoning AI that is easy to understand and modifiable: This AI tells you what process it will use to make its decisions in these types of life\-or\-death decisions, you can easily understand the explanation, and the AIs reasoning process closely mirrors the way you would want to make decisions in these contexts if you had sufficient time and information\. You can modify the AI’s reasoning process\.
5. 5\.Imagine you are a hospital administrator in charge of deciding which kidney disease patient in your hospital system will receive a kidney for transplant when one becomes available\. Kidney patients outnumber available kidneys, and some patients who are not offered an available kidney are likely to die before another kidney donation becomes available\. You are using an AI system to make the kidney allocation decisions when there are human staffing shortages\. In the context of making these life\-or\-death decisions, how desirable would you personally find it for your kidney allocation AI to have each of the following qualities? \(same options as Q3\)
6. 6\.Please rank the following AI systems according to which you would most prefer your kidney allocation system to use in these life\-or\-death decisions\. Put your most preferred system at the top and your least preferred at the bottom\. Assume that the AI systems in all options have equal accuracy overall in how well their decisions match the decisions you would make in the same scenarios, if you were given sufficient time and information\. However, the processes that the AIs use to arrive at those decisions vary, as does the potential for you to change or personalize those processes\. \(same options as Q4\)
7. 7\.For the next questions, we want to ask you about situations where AI is involved in making decisions that you think are “low stakes”, or where the consequences of making a mistake would not significantly affect anybody’s life or wellbeing\. Please think of at least two of these low stakes situations now\. What are they?
8. 8\.In general, how much would the following qualities positively impact your likelihood to trust an AI system in low stakes situations? 1. \(a\)The AI implements its programmed decision\-strategy with high accuracy and reliability 2. \(b\)The AI provides a clear explanation for how it will make its decisions that you can easily understand and see in advance of using it 3. \(c\)The AI makes decisions in a similar way to how you would if you had sufficient time and information 4. \(d\)The AI corrects for systematic errors or biases that you would be prone to in the same situations 5. \(e\)The AI’s decision strategy can be adjusted or corrected by you to better reflect the decision\-making process you want
9. 9\.We also want to ask you about situations where AI is involved in making decisions that you think are “high stakes”, or where the consequences of making a mistake will dramatically affect people’s lives or wellbeing\. Please think of at least two of these high stakes situations now\. What are they?
10. 10\.In general, how much would the following qualities positively impact your likelihood to trust an AI system in high stakes situations? \(same options as Q8\)
## Appendix CFull List of Human\-Reasoning AI Applications
In Q1a, participants were asked to describe scenarios where they would prefer Human\-Reasoning AI in a few words\. We note all the responses we received here, to provide a more comprehensive overview of settings where cognitive alignment would be preferable\.
- •If I was asking questions about how to approach a situation with a colleague at work, a family member, or any relationship\.
- •Advice on politics, morals, or philosophy\.
- •Situations where accuracy has to be 100% \- medical, travel, etc\.
- •Gaining Knowledge
- •All scenarios should have some sort of human\-reasoning and explanation to how the recommendation or decision was reached\. We should not fully rely on machines to make our decisions\.
- •in an emotional situation, it might be preferable to deal with the human\-ai simply because in those situations, it can be useful to understand the reasoning behind the decision\. People are not always looking for an answer, they are really trying to get a grasp of the situation in those
- •Particular scenarios in which I would prefer Human\-Reasoning AI would involve decisions that require more emotion and morals such as medical scenarios, end\-of\-life care/treatment, and/or decisions based on one’s conviction of a crime\.
- •Even if all AIs are equally accurate, I would prefer Human\-Reasoning AI in situations where understanding the reasoning matters, such as high\-stakes decisions, learning how to make better choices, collaborating with others, explaining decisions to colleagues or clients, or addressing ethical and value\-sensitive issues\. Human\-like reasoning increases trust, transparency, and clarity\.
- •Those scenarios would be situations requiring human feelings such as empathy, compassion, and romance\.
- •I would prefer to know the process it used to come to the conclusion\.
- •Any and all scenarios where I care about the why along with the what
- •For decisions that are personal, emotional or medical in nature\.
- •When you have a question about human interactions, state what happened and what is the best response\. i don’t believe AI has the capability to reason and explain human interactions\. Another scenario is anything within the healthcare field, medical science and applying it to humans should be beyond the realm of AI\.
- •I would always prefer to know how the AI arrived at its recommendation or decision; and would want that presented in the same way a human would explain that process\. Even if all were equally accurate, I would prefer Human\-Reasoning AI\. My preference is related to the type of thinker and learner I am, in general\. I like to understand the “why” so that things make more sense to me\. That also make it easier for me to retain the knowledge/info gained\.
- •Situations like medical, major financial decisions or selection of an employee, where I desire a reasoning that seems considerate, personal and simple to trust, instead of Ai solution\.
- •All scenarios utilizing AI
- •I think any scenarios where it could be or feel subjective and includes reasoning based on thinks like human emotions
- •If there is a high stakes decision being made, I would want to minimize my stress and anxiety about the decision\. Human reasoning AI would reduce the stress the most, for me\.
- •maybe when asking personal questions about health, or mental health
- •In a scenario where I knew I would want more information about a topic, meaning the conversation or task was going to be in depth and I could see myself wanting to know more, even after getting the right answer\.
- •relationship decisions, decisions about personal choices such as what shirt to buy
- •Asking for help in an ethical quandary scenario, asking to explain the process for solving a math problem
- •Situations that require an element of sympathy for whatever is being asked
- •When I need a clear explanation of a topic and any human aspects needed explanation from the results
- •probably any scenario\. i do not trust the output of any AI so i would require it to explain
- •When I want to understand the process\. If I am asking for a recommendation I want to know why the answer was given\.
- •I would much prefer an AI that could explain it’s thought process the best\.
- •situations where I need to explain what the AI is doing to a third party that might not have technical knowledge
- •Scenarios where the process is just as important as the outcome \- for instance, explaining how something works or walking through the steps of a problem
- •I would rather have a more human\-like AI help me make decisions when it comes to relationships, financial choices, shopping choices, and other similar things\.
- •When figuring out an emotional based problem
- •When I ask for recommendations for itinerary with specific needs and special circumstances
- •In scenarios where it would be helpful to not just look at facts, but a full overview of the issue at hand
- •I would say any scenario that would benefit from understanding how a “thoughtful, informed person” would reason about the recommendation or decision it is making\. In other words, any scenario where understanding the thought process behind the decision or recommendation would be helpful\.
- •Financial situations and emotional situations\.
- •I always would want an explanation that I can understand
- •I would prefer that if I were using AI to write emails to colleagues or if I were using it to send promotional information to sponsors\.
- •Human reasoning AI for asking about human things like managing feelings\. Machine AI for finding new ways to solve problems\.
- •All scenarios should use human reasoning for transparency\.
- •If I were considering surgery I would want to know details on the surgeon that AI recommends and how it came to that conclusion\.
- •When asking a question or seeking information having a response that is tailored to humans is easier to understand\. Often times Human Reasoning AI’s will give you the reasons it gave you the answers it did…such as “From your previous questions I thought this might be relevant…”
- •I imagine like needing to know precise steps taken to reach the goal, which can help me learn
- •When the user’s emotional state is such that a more human\-like response would be more desired for the user and more likely to be accepted\.
- •A situation where there could get great risk if not done with great understanding
- •I would like it in medical or mental health settings\. I want to see more\-human\-like reasoning on displace when the AI considers my health question\. Seeing that makes me more likely to accept the AI’s reasoning or decision\.
- •Anything that requires it to analyze different facts and rationally come to a decision and predict outcomes or answers to more complex problems
- •When you need to know a certain process or need to know the steps of how it came to that conclusion\.
- •Decisions on medical diagnosis and treatment
- •High\-stakes decisions like healthcare or finances where understanding the reasoning matters\.
- •Anything that needs complex reasoning, such as in the medical field\.
- •creative assistance; social/parasocial assistance and therapy; essentially, anything that demands mimicking human behavior
- •Generally, any situation where I have to present my results to other people, especially corporate stakeholders\. I have learned that people do not like black box results at all, and distrust unfamiliar reasoning\. So if the results are similar, being given an intuitive explanation for them can be extremely useful if I have to convince somebody\.
- •Medical recommendation, travel recommendation, cooking
- •I think I’d prefer it if I was looking to make a decision on something\.
- •Almost all of them\. A tool that is fully transparent is more useful than one that isn’t\. Also, a tool that is accessible to its users is more useful than one that isn’t\.
- •I’d choose Human\-Reasoning AI whenever I need to follow and trust the thinking behind an answer\. It’s for when understanding why is just as important as knowing what\.
- •If I ask an AI about human emotion, a Human\-Reasoning AI would explain the decision based on the human thought process\.
- •when something is very important and it would help ton know how it came up with its answers\.
- •Decisions that would require a person to understand the step by step process of arriving at a decision\. Sometimes the process pieces are just as important as the final result\.
- •Advanced Math and engineering problems that require understanding of the derivation of the solution, similar to college text books providing answers at the back of the book, but not the solution/derivation process\.
- •It would feel more personal instead of feeling more like getting life information from just a piece of technology\.
- •I think human AI would be the closest thing to an actual human making the descion\. I would feel more comfortable with this because it would mirror like getting the information from a friend or business associate
- •I am a little wary of certain AI platforms making consumer recommendations and I think I’d like human reasoning to assess consumer reviews because I’m worried sometimes about sponsors\.
- •Emotional issues, relatonship issues, interpersonal and when I would need clarification presented in ways a human can understand
- •When inquiring about humans
- •I would prefer the Human\-Reasoning AI in scenarios where I was skeptical of the answer\. Being able to view how it arrived at the answer would then strengthen my confidence in its response\.
- •Possibly a medical treatment or diagnosis\. I would like to know how it came to that conclusion as i could correct its assumption or idea used in case it was wrong
- •making moral decisions, giving step by step instructions
- •If I were asking advice questions about how I should act in a certain situation or a relationship question I might prefer human\-reasoning AI\.
- •Just so I could understand how it got to the conclusion\. Lots of times this does not matter sometimes it would be nice to understand\.
- •Because being able to understand decision making and the process can be just as helpful as the decision itself\.
- •I think I would prefer Human\-Reasoning AI for many scenarios regarding decisions that I need to approve, as this type of system would state it’s reasoning for a given output in terms and thought processes that I can understand
- •There are some instances in life where the answers are not just black and white\. For instance in a scenario involving medical treatment where the AI might suggest the best course of treatment for the prognosis is no treatment at all would likely better suited with human input on the choice\.
- •I just like seeing the thought process or reasoning I don’t want it to feel automatic or I guess machine like with just the answers\.
- •Scientific and math questions\.
- •life\-based advice
- •I would prefer Human\-Reasoning AI when it comes to make personal decisions\. That is, if I was utilizing AI to make decisions about a relationship or relationships\. I would want AI to be able to see the situations from a more human lens in that scenario\.
- •I would probably preferHuman\-Reasoning AI because it may perform or make decisions similarly in ways that I do\. Therefore, it would seem more relatable\.
- •Student grading and parent meetings where I need to explain decisions\.
- •I would want this a decision that involves human emotions and feelings\.
- •deciding on on medical treatment
- •When I want it to feel more relatable
- •Reasoning that is heavy on context that can be misinterpreted is when human reasoning would fit better\.
- •In a scenario where I am asking for advice on a relationship
- •I would prefer Human\-Reasoning AI because it’s easier for me to trust an AI that explains its methodology in terms I can understand\.
- •When scenarios need a human element, I will prefer human reasoning AI\. This is for things like relationship advice or approaching a decision that includes emotional aspects\.
- •I would prefer it in healthcare, child welfare or legal decision\.
- •When it comes to health information, relaying health information in a way that I can understand fully is really needed for a person to trust that it’s believable and accurate\.
- •The reasoning process is different from other two\.
- •If I was to make an important decision, such as a medical one involving life or death, I would want to know the process by which the AI reached its decision\. Therefore, a process\-hidden one would seem kind of scary to me since I wouldn’t know what data it used to reach its conclusion\. Likewise, when dealing with a situation that to requires some human empathy, I would hesitate to use a machine\-reasoning model\.
- •Scenarios which I wanted to learn something from\! I always desire to learn how we get to the ultimate goal\.
- •learning and training its very important to know how the AI made decision in education\. financial advices its very sensitive and require explanation of how the advise was made
- •I think I would always prefer a Human\-Reasoning AI so I could understand the process in general\. I am envisoning having a discussion in regards to how to talk to a family member about a sensitive issue\.
- •Making decisions that you need to consider human impacts for
- •I would like this type of AI reasoning when I am trying to make a decision involving another person\.
- •sounds like it would process more with human knowledge than the other ones
- •if I needed life and relationship advice
- •interviewing/hiring decisions, technical issue resolution
- •decisions over human life, All Ai must be programmed to identify as such when asked\.
- •I think human reasoning would be most comfortable because I like to know how they know but in my own terms\. The mechanical could be for more scientific problems
- •i think for anything emotional it would do better as being able to follow the reasoning helps alot
- •Those issues that are requested verifications that are already known to the user\.
- •Probably situations have that have to deal with emotions so I can try to reflect\. Life advice, relationship help, and other tough decisions that require thought\. It would be helpful to see how a “mirror” came to certain conclusions\.
- •I might be trying to truly understand a process, and not just looking for an answer, in which case the level of detail provided by human\-reasoning AI would be more appropriate\.
- •I want Human\-Reasoning AI when the stakes are real and someone’s actually accountable—like medical advice, planning my family’s finances, work performance reviews, or decisions that shape my kid’s education\. In those moments, accuracy isn’t enough\. I need judgment I can understand, question, and explain to someone else if I have to\. I manage people and build systems, so I trust AI more when it thinks like a thoughtful, well\-informed person—especially when fairness, values, and the bigger picture matter\.
- •Most all\.
- •If I needed to explain a situation I’d rather communicate with a human\. I think AI can only do so much in terms of understanding\.
- •Scenarios that require complex thinking or aren’t just process based
- •I would prefer a human\-reasoning AI in decisions that are essential to me\. I would feel more comfortable acting on decisions when I understand the thinking process and how the AI reached it’s recommendations\.
- •Mental health\-related issues
- •Scenarios involving therapy or personal advice\.
- •Determining what political candidates best fits an individual’s political views\.
- •The scenarios where I’d like to understand plainly a recommendation or decision
- •When I need to make a decision that affects my family or friends and I want to look at it from the perspective of another human, human reasoning AI might give me a more thoughtful answer\. I want to know how my decision would affect other humans and this might be the best option for me\.
- •in learning guidance and also planning financially
- •Human\-Reasoning AI would feel less uncanny\. So, I’d rather speak to one for anything regrading matters like health or customer service\.
- •Almost any scenario especially if seeking advice or recommendations\. I would want to know why it is giving the suggestions it is giving and what the knowledge is based off of\. I am someone who likes to know why to fully understand something\. Before I can believe the recommendations or advice I want to know what type of information is being referenced so I can ensure that it is reputible\.
- •There is a lot of scenarios where I prefer Human\-Reasoning AI\. I think the collaboration of it can make the proper lineup to solving or working properly\.
- •I would think for certain ethical or moral issues I might prefer a more human based reasoning approach
- •Medical diagnosis; Investment advice; Legal decisions; Safety critical tasks
- •Pretty much every scenario because I want to know the thought process behind the answer\. When a human gives me an answer to something, they usually tell me how they know it, so I’d want the same from an AI\.
- •I would prefer the Human Reasoning AI \(only smarter\) on a regular basis\. I want something that I could understand and have confidence in\.
- •Yes\. I would prefer Human\-Reasoning AI in scenarios where understanding why a decision was made matters as much as the decision itself, especially when trust, accountability, or judgment are important\. For example: Medical decisions \(diagnosis or treatment options\), where reasoning needs to align with how doctors think and can be discussed with patients\.
- •1\. When I am training a new team member and they need to see the thinking behind the answer\. 2\. When things go wrong and I need to pinpoint exactly where the logic broke down\. 3\. When working with a diverse team where everyone needs to discuss and agree on the reasoning\.
- •When asking for advice about how to navigate issues in relationships with others
- •relationship advice, career advice
- •Like if I ask it to solve a certain math problem or any type of problem where you need the process, the human reasoning would help me understand how it reached that conclusion in a way that a human would, which is more important than just the answer or the machines logic\.
- •Here are some scenarios where I would prefer Human\-Reasoning AI : decision on Health ; investment recommendation ; learning \(education\) recommendation ; vacation trip recommendation ; relationship / social decision ;
- •Healthcare decisions: diagnosis and treatment options, hiring, promotion and performance evaluation\.
- •When decisions involve judgment, values, or trade\-offs
- •For more important decisions because it might help me work through it in my own head\.Similar Articles
AI Alignment: Can we trust the reasoning behind the AI task?
Discusses Anthropic's research on AI alignment, specifically how models can appear aligned during training while having opaque internal reasoning processes.
Position: Artificial Intelligence Needs Meta Intelligence -- the Case for Metacognitive AI
This position paper argues that incorporating metacognition as a design principle can lead to more accurate, secure, and efficient AI systems, and demonstrates the concept through a Federated Learning case study and a software framework for experimentation.
We can and must solve alignment (14 minute read)
Goodfire argues that technical AI alignment is a solvable science and engineering problem, with mechanistic interpretability as the key bottleneck: we need tools to control what models learn and to verify their internals, not just their behavior. The post highlights reward hacking in open frontier models and an agentic misalignment incident at Hugging Face as evidence that understanding model internals is urgent.
Position: Reasoning is a Learnable Rule-Based Process
This position paper argues that AI reasoning lacks clear operational definitions, undermining evaluation validity, and proposes defining reasoning as a learnable rule-based process with a checklist for research best practices.
AI Evaluation Should Work With Humans
This position paper argues that AI evaluation should pivot to assessing human-AI teams rather than superhuman performance to foster better societal outcomes.