Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
Summary
This paper introduces a Human-Centric Reflective Architecture (HCRA) for human-AI collaborative decision-making, formulating the task as a stochastic game and using reinforcement learning with linguistic feedback. The proposed framework significantly improves decision-making effectiveness and recommendation quality.
View Cached Full Text
Cached at: 07/07/26, 04:34 AM
# Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
Source: [https://arxiv.org/html/2607.03025](https://arxiv.org/html/2607.03025)
\(5 June 2009\)
###### Abstract\.
The use of Large Language Models \(LLMs\) across diverse areas of human activity—ranging from everyday tasks to safety\-critical applications—aims to enhance decision\-making effectiveness with minimal human feedback\. Concurrently, it seeks to align decisions with human expectations, preferences, and needs while mitigating risks associated with AI non\-determinism\. However, humans frequently over\- or under\-rely on AI recommendations, and current AI systems remain poorly calibrated to human expectations\. To address these challenges, we introduce a human\-AI collaborative decision\-making framework designed to augment human capabilities and align AI agents with human preferences and expectations\. Specifically, this paper \(a\) formulates the collaborative decision\-making task as a stochastic game between an AI agent and a human player, and \(b\) proposes the Human\-Centric Reflective Architecture \(HCRA\), which integrates human\-calibrated models with reinforcement learning agents that leverage linguistic feedback in an iterative, reflective process\. Evaluation results demonstrate that HCRA significantly enhances decision\-making effectiveness and delivers high\-quality recommendations\.
Agentic AI, Human\-centric AI, Collaborative Decision Making, Large Language Models, Language Agents, Reinforcement Learning, Alignment
††copyright:acmlicensed††journalyear:2018††doi:XXXXXXX\.XXXXXXX††conference:14th EETN Conference on Artificial Intelligence; September 09–11, 2026; Chania, Greece††isbn:978\-1\-4503\-XXXX\-X/2018/06††ccs:Computing methodologies Stochastic games††ccs:Computing methodologies Learning from critiques††ccs:Computing methodologies Reinforcement learning††ccs:Information systems Decision support systems## 1\.Introduction
In many collaborative decision\-making environments, humans and AI agents engage in an interactive feedback loop\. This iterative process aims to: \(a\) guide agents toward human\-acceptable recommendations, \(b\) enable users to express and refine their preferences and expectations, and \(c\) allow agents to learn from historical interactions\. Such tight interaction is vital in high\-stakes domains \(e\.g\.,\(Beedeet al\.,[2020](https://arxiv.org/html/2607.03025#bib.bib1)\),\(Fischer\-Abaigaret al\.,[2024](https://arxiv.org/html/2607.03025#bib.bib8)\)\) where humans must retain ultimate decision\-making authority\.
However, effective human\-AI collaboration remains a challenge\. When partnering with AI, human performance often falls short of expectations\(Liuet al\.,[2021](https://arxiv.org/html/2607.03025#bib.bib6)\)\. While AI explanations can influence user decisions\(Li and Yin,[2025](https://arxiv.org/html/2607.03025#bib.bib13)\), they frequently fail to improve system understanding or properly calibrate trust\(Liuet al\.,[2021](https://arxiv.org/html/2607.03025#bib.bib6)\)\(Buçincaet al\.,[2020](https://arxiv.org/html/2607.03025#bib.bib2)\)\(Bansalet al\.,[2021b](https://arxiv.org/html/2607.03025#bib.bib3)\)\(Zhanget al\.,[2020](https://arxiv.org/html/2607.03025#bib.bib4)\)\(Wang and Yin,[2021](https://arxiv.org/html/2607.03025#bib.bib5)\)\(Vasconceloset al\.,[2022](https://arxiv.org/html/2607.03025#bib.bib7)\)\. This shortfall is especially critical when users must dedicate limited cognitive bandwidth to time\-sensitive, high\-consequence problems\(Papadopouloset al\.,[2024](https://arxiv.org/html/2607.03025#bib.bib29)\)\. Furthermore, high recommendation accuracy alone is insufficient to build trust; effective collaboration requires AI systems to optimize for human utility alongside core task objectives\(Bansalet al\.,[2021a](https://arxiv.org/html/2607.03025#bib.bib14)\),\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\)\.\.
Focusing on LLM\-assisted agents, our work is motivated by two core limitations\. First, miscalibrated human expectations regarding AI capabilities often trigger over\- or under\-reliance\(Bansalet al\.,[2019](https://arxiv.org/html/2607.03025#bib.bib15)\)\(Zhang,[2023](https://arxiv.org/html/2607.03025#bib.bib16)\)\(Boet al\.,[2025](https://arxiv.org/html/2607.03025#bib.bib17)\)\. Second, rather than relying solely on large\-scale training followed by aggregated alignment \(e\.g\., RLHF\(Ziegleret al\.,[2020](https://arxiv.org/html/2607.03025#bib.bib11)\)\(Ouyanget al\.,[2022](https://arxiv.org/html/2607.03025#bib.bib12)\)\), we leverage test\-time scaling\(Snellet al\.,[2024](https://arxiv.org/html/2607.03025#bib.bib18)\)\(Jaech and et al\.,[2024](https://arxiv.org/html/2607.03025#bib.bib19)\)and contextual test\-time tuning\. This approach establishes a human\-centric workflow where high\-quality decisions are made in\-context based on user utility, enabling agents to dynamically refine recommendation quality and mitigate issues stemming from LLM non\-determinism\.
To address these issues, we propose a human\-centric reflective architecture \(HCRA\) for collaborative decision\-making that integrates reinforcement learning \(RL\) language agents with human\-calibrated models\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\), in an iterative reflective process\(Xuet al\.,[2025](https://arxiv.org/html/2607.03025#bib.bib20)\)111Datasets and code are available in[https://github\.com/AILabDsUnipi/HCRA](https://github.com/AILabDsUnipi/HCRA)\.
This article makes the following contributions:
\- Game\-Theoretic Formulation: We model the collaborative decision\-making process as a stochastic game between an AI agent and a human\. Within this framework, we define the human\-utility optimized for user expectations and preferences\. This utility maximization criterion provides formal convergence and termination guarantees for the iterative decision\-making process\.
\- Architecture Design: We propose HCRA, a framework that optimizes human\-AI collaboration through an iterative reflection process\. HCRA tunes AI recommendations at test time by leveraging historical interactions, an AI recommendation calibration model, and a human acceptance model\.
\- Empirical Evaluation: We validate HCRA within the domain of tourism recommendations—a setting where inaccuracies can have real\-world consequences\.222See, for example,[https://www\.webpronews\.com/ai\-hallucinations\-in\-travel\-apps\-lead\-to\-fake\-landmarks\-and\-dangers/](https://www.webpronews.com/ai-hallucinations-in-travel-apps-lead-to-fake-landmarks-and-dangers/)\.Our results demonstrate the framework’s effectiveness and recommendation quality, highlighting the vital role of human behavior modeling and human\-centric objective functions\.
## 2\.Related work
This work contributes to the growing and important efforts to assist AI\-assisted decision\-making\(Laiet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib9)\)\. The most basic paradigm here involves the AI agent to perform an assistive role by providing a recommendation that a human may accept or reject\.
A critical challenge is whether we can achieve human\-AI collaboration that can outperform the human and AI alone, contributing to high\-quality recommendations via an effective decision\-making process\. Addressing this challenge, several recent studies propose different designs to help humans allocate appropriate trust to AI based on confidence scores\. Empirical results suggest that people have poorly\-calibrated self confidence and cannot reliably reflect well\-calibrated confidence scores\. This motivates research questions in\(Bansalet al\.,[2021a](https://arxiv.org/html/2607.03025#bib.bib14)\),\(Maet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib21)\)and\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\), emphasizing on the importance of AI advice confidence for human decisions\. This research is in contrast to shaping humans’ self confidence for AI assisted decision\-making \(e\.g\. as in\(Takayanagiet al\.,[2025](https://arxiv.org/html/2607.03025#bib.bib23)\)\)\. The authors in\(Maet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib21)\)try to answer this question by proposing a framework that considers well\-calibrated confidence scores from both humans and AI\. However, the framework proposed in\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\)highlights that modifying the confidence of an AI advice by exploiting a human behavior model can enhance human\-AI collaboration: Authors propose optimizing the transformation of the AI advice confidence so that the transformed confidence score matches the human perception of AI confidence scores\. In so doing they optimize an AI system with respect to the human user, towardshuman\-calibrated AI\.
In addition to these efforts, authors in\(Benz and Rodriguez,[2024](https://arxiv.org/html/2607.03025#bib.bib22)\), following a more fundamental approach, show that if the confidence values of \(human\) decision makers satisfy a natural alignment property with respect to the confidence they have on their own predictions \(human\-alignment\) then the decision makers can make optimal decisions\.
Building on these results, in this paper we optimize an AI language agent with respect to human behavior models\. In contrast to\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\)we consider that the human does not have any information on what the true advise could be \(and, there is not necessarily an initial guess\), but he/she may have certain requirements that should be satisfied by the final decision\. Furthermore, in contrast to a single\-step decision\-making task, we formulate and implement an iterative reflective collaborative process which transparently, without human intervention, aims at maximizing the utility of humans\.
The reflective process is inspired by Reflexion\(Shinnet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib26)\): It represents a novel paradigm in RL for language agents, with no requirements regarding gradient\-based optimization methods\. Reflexion enables agents to learn from trial\-and\-error through linguistic feedback rather than parameter updates, without requiring expensive model fine\-tuning\.
Reflexion is formalized as an iterative optimization process\. As described in\(Shinnet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib26)\), in any iteration of the reflective processtt, an actor produces a trajectoryτt\\tau\_\{t\}to solve a sequential task, by interacting with the environment\. An evaluator model produces a scalar reward signalrtr\_\{t\}that improves as task\-specific performance increases\. To amplifyrtr\_\{t\}to a feedback form that can be used for improvement by an LLM, the self\-reflection model analyzes the set of \{ \(τi,ri\\tau\_\{i\},r\_\{i\}\), i=1,…,t\} to produce a summarysrtsr\_\{t\}, which is stored in a memory buffer\. The actor, evaluator, and self\-reflection models work together through iterations until the evaluator deems the final trajectoryτT\\tau\_\{T\}at iterationTT, to be correct\. The memory components of Reflexion, i\.e\., the short term memory for trajectory history and the long term memory implementing a persistent knowledge base that informs future decision\-making, are crucial to its effectiveness\.
Reflexion’s key innovation lies in transforming sparse rewards on trajectories into linguistic feedback signals\. Reflexion’s linguistic reflections provide specific insights into failure modes and suggest targeted corrective actions in an iterative process\. This iterative process of trial, error, self\-reflection, and persistent memory, as shown in\(Shinnet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib26)\), enables the agent to improve its decision\-making ability and facilitates learning by addressing recommendation deficiencies\.
Aiming to a human\-centric decision\-making process, we integrate well\-trained predictive human models into the agent’s reflective loop\. These models steer the selection of recommendations toward human\-centric objectives, aiming to maximize expected human utility, dictating the reflective iterative process termination conditions\. Trained on empirical data gathered from real users, these models accurately simulate human acceptance behavior\. By substituting simulated interactions for real\-world responses, this approach allows the system to evaluate decisions without requiring continuous, manual human feedback at every step\. Furthermore, we consider simple, one\-shot decision\-making tasks rather than sequential ones\.
## 3\.Problem formulation
The iterative reflective decision\-making process is formulated as a stochastic game between the AI and the human: At any iteration \(time step\)tt, the process is at statests\_\{t\}and players choose actions as required by the decision\-making process\. They get their payoffs, and the process proceeds to the next iteration\. Formally, the iterative reflective decision\-making stochastic process has the following components:
\-SS: A set of games \(states\)\.st∈Ss\_\{t\}\\in Sis specified to be of the form<Qrt,Cs,\(Ret,cft\),gt\><Qr\_\{t\},Cs,\(Re\_\{t\},cf\_\{t\}\),g\_\{t\}\>, where
1. \(1\)QrtQr\_\{t\}is the request attt, formed by the initial requestQr0Qr\_\{0\}specified by the human, and enhanced at subsequent iterations\.
2. \(2\)CsCsis a set of human preferences or constraints that should be satisfied by the agent recommendation\. We may distinguish between preferences \(a\.k\.a soft constraints\) and hard constraints, although subsequently we refer to any of these as “constraints”\. Constraints are expressed in natural language, but they can be formulated as \(attrattropopvaluevalue\), whereattrattris a real\-world domain\-specific attribute \(e\.g\. time of task completion, duration, amount of resources, style/skill\),opopcan be any of the operators in\{=,≤,≥,\>,<\}\\\{=,\\leq,\\geq,\>,<\\\}, or any domain\-specific specification of a relation, andvaluevalueis any of the numerical or categorical values forattrattr\(e\.g\.criminalitycriminalityininareaareaisisveryverylowlow,menu\_price<100menu\\\_price<100Euros\)\.
3. \(3\)\(Ret,cft\)\(Re\_\{t\},cf\_\{t\}\)is the AI recommendation at time steptt, comprising the contentRetRe\_\{t\}of the recommendation and the confidencecftcf\_\{t\}assessed by the agent\.
4. \(4\)gtg\_\{t\}is the human\-calibrated confidence of the recommendation, to be shown to the human\. Actually,gtg\_\{t\}is a transformation ofcftcf\_\{t\}to meet human\-calibrated expectations from AI recommendations \(explained subsequently\)\.
ConstraintsCsCsmust be satisfied by the AI recommendationRetRe\_\{t\}\. In case there is not any violation, we refer to that situation as an “agreement”\.
\-NN: This is the set of players, including the human \(denoted byhh\), here represented by a human model, and the AI agent \(denoted byagag\)\.
Althoughagaghas full access to the components of a states∈Ss\\in S, the human is assumed to observe the recommendationRetRe\_\{t\}and the human\-calibrated confidence of the recommendationgtg\_\{t\}, at any time steptt\.
\- Set of actions:A=Ah×AagA=A^\{h\}\\times A^\{ag\}, whereAiA^\{i\}is a set of actions available to playeri∈\{h,ag\}i\\in\\\{h,ag\\\}\. Specifically, actionsatha^\{h\}\_\{t\}inAhA^\{h\}are in\[0,1\]\[0,1\]and specify the probability of the human to accept the recommendation at that state\. Given that humans are represented by a human behavior model,atha^\{h\}\_\{t\}is thepredictedpredictedprobability of the human to accept the recommendation at time stepttand is provided by the functionhacc:S→\[0,1\]h\_\{acc\}:S\\rightarrow\[0,1\]\. Actionsataga^\{ag\}\_\{t\}are inAag=\{Reflective\_text\(st\)\|st∈S,t=0,1,2…\}A^\{ag\}=\\\{Reflective\\\_text\(s\_\{t\}\)\|s\_\{t\}\\in S,t=0,1,2\.\.\.\\\}, which is the set of reflective texts that can be provided as linguistic feedback at any time steptt\.
\-P:S×A×S→\[0,1\]P:S\\times A\\times S\\rightarrow\[0,1\]is the state transition function from statests\_\{t\}to statest\+1s\_\{t\+1\}after the execution of the joint players’ actionat=\(ath,atag\)∈Aa\_\{t\}=\(a\_\{t\}^\{h\},a\_\{t\}^\{ag\}\)\\in A\.
\- Reward functionsrhr^\{h\}andragr^\{ag\}:rh:S→ℝr^\{h\}:S\\rightarrow\\mathbb\{R\}is a real\-valued payoff function for the human player, andrag:S→<ℝ×ℝ\>r^\{ag\}:S\\rightarrow<\\mathbb\{R\}\\times\\mathbb\{R\}\>is the payoff function for the agent comprising persts\_\{t\}the correctness and the agreementassessmentsassessmentswith regards toRetRe\_\{t\}, denoted byCorrt^\\widehat\{Corr\_\{t\}\}andAggrt^\\widehat\{Aggr\_\{t\}\}, respectively\. We consider these as assessments, given that no player has access to the ground truth\.
Although alternative formulations of the stochastic process are possible \(e\.g\. specifying agent actions to be the set of possible recommendations\), the above specification emphasizes on the importance of the linguistic reflective feedback provided by RL language agents during the collaborative process\.
### 3\.1\.Rewards
The rewardrag\(st\)r^\{ag\}\(s\_\{t\}\)that the AI agent gets after the execution of the joint actionat=\(at−1ag,at−1h\)a\_\{t\}=\(a^\{ag\}\_\{t\-1\},a^\{h\}\_\{t\-1\}\)in a statest−1s\_\{t\-1\}results from the evaluation of the statests\_\{t\}in terms of correctness and agreement to constraints\. This is done by an evaluation component, whose functionality is specified subsequently\.
The reward of the human in iterationttis defined as follows:
rh\(st\)=c\(st\)×log\(hacc\(st\)\)\+\(1−c\(st\)\)×log\(1−hacc\(st\)\)r^\{h\}\(s\_\{t\}\)=c\(s\_\{t\}\)\\times log\(h\_\{acc\}\(s\_\{t\}\)\)\+\(1\-c\(s\_\{t\}\)\)\\times log\(1\-h\_\{acc\}\(s\_\{t\}\)\)where,c\(st\)∈\{0,1\}c\(s\_\{t\}\)\\in\\\{0,1\\\}is theassessedassessedclass of the AI recommendation at statests\_\{t\}, depending on the assessed correctness and the assessed agreement of the AI recommendationRetRe\_\{t\}\. We considerc\(st\)=1c\(s\_\{t\}\)=1when the recommendationRetRe\_\{t\}is assessed to be correct and with agreement to constraints, andc\(st\)=0c\(s\_\{t\}\)=0in all other cases\. In a more general case, we can considerc\(st\)c\(s\_\{t\}\)to be the probability that at statests\_\{t\},RetRe\_\{t\}is correct and with agreement, but this is not the case in this work\. This reward function assigns a high penalty when the human \(model\) either \(a\) accepts a recommendation that is assessed to be incorrect or with no agreement to constraints, or \(b\) does not accept a recommendation that is assessed to be correct and with agreement to constraints\. Humans get the highest reward when they accept a proposal that is assessed to be correct and in agreement to constraints\.
It must be noted that state transitions and thus rewards players get depend on the joint action of players\. Also,rhr^\{h\}depends on the components of the agent rewardragr^\{ag\}\.
### 3\.2\.Objective
The objective of the overall task is to maximize the human utility, or else, to minimize the expected human lossℒT\\mathcal\{L\}\_\{T\}, afterTTiterations\. To formulate this, we can consider that at each time steptt, the human player plays the lottery
\[Lt:hacc\(st\),ℒrefl\(st\):\(1−hacc\(st\)\)\]\[L\_\{t\}:h\_\{acc\}\(s\_\{t\}\),\\mathcal\{L\}\_\{refl\}\(s\_\{t\}\):\(1\-h\_\{acc\}\(s\_\{t\}\)\)\]where, if the human accepts the decision atttwith probabilityhacc\(st\)h\_\{acc\}\(s\_\{t\}\)the loss experienced by the human isL\(st\)=−rh\(st\)L\(s\_\{t\}\)=\-r^\{h\}\(s\_\{t\}\)\. Else, with probability\(1−hacc\(st\)\)\(1\-h\_\{acc\}\(s\_\{t\}\)\), the human experiences the expected lossℒrefl\(st\)=\(hacc\(st\)L\(st\+1\)\+\(1−hacc\(st\+1\)ℒrefl\(st\+1\)\)\)\\mathcal\{L\}\_\{refl\}\(s\_\{t\}\)=\\big\(h\_\{acc\}\(s\_\{t\}\)L\(s\_\{t\+1\}\)\+\(1\-h\_\{acc\}\(s\_\{t\+1\}\)\\mathcal\{L\}\_\{refl\}\(s\_\{t\+1\}\)\)\\big\)by proceeding to subsequent iterations of the process\. In so doing, at the final iterationTTthe expected human loss is
ℒT=\\displaystyle\\mathcal\{L\}\_\{T\}=\[hacc\(sT\)∏j=1T−1\(1−hacc\(sj\)\)L\(sT\)\]\+\\displaystyle\[h\_\{acc\}\(s\_\{T\}\)\\prod^\{T\-1\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)L\(s\_\{T\}\)\]\+\[∑i=1T−1hacc\(si\)L\(si\)∏j−1i−1\(1−hacc\(sj\)\]\\displaystyle\[\\sum^\{T\-1\}\_\{i=1\}h\_\{acc\}\(s\_\{i\}\)L\(s\_\{i\}\)\\prod\_\{j\-1\}^\{i\-1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\]The first term in brackets specifies the expected loss in case the human accepts the recommendationReTRe\_\{T\}generated at the final time stepTT\(previous recommendations have not been accepted\)\. The second bracket specifies the expected loss of accepting any of the previous recommendations generated at time steps1,…,\(T−1\)1,\.\.\.,\(T\-1\)\.
Ideally, the final iterationTToccurs when the recommendation atTTis accepted with probability equal to 1\. However, this can not be easily done, since no human model has access to the ground correctness of any recommendation\. Otherwise, at time stepTTany of the following should occur: \(a\) There is indifference between accepting the recommendation generated atTTand any of the recommendations generated at steps 1…T−1T\-1, or \(b\) The expected loss at stepTTis smaller than the loss of accepting any of the previous recommendations generated at time steps1…\(T−1\)1\.\.\.\(T\-1\)\.
Thus, for the process to terminate at stepTT, it must hold that
\[hacc\(sT\)∏j=1T−1\(1−hacc\(sj\)\)L\(sT\)\]≤\\displaystyle\[h\_\{acc\}\(s\_\{T\}\)\\prod^\{T\-1\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)L\(s\_\{T\}\)\]\\leq\[∑i=1T−1hacc\(si\)L\(si\)∏j−1i−1\(1−hacc\(sj\)\]\\displaystyle\[\\sum^\{T\-1\}\_\{i=1\}h\_\{acc\}\(s\_\{i\}\)L\(s\_\{i\}\)\\prod\_\{j\-1\}^\{i\-1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\]Thus, to bound the loss near to zero, for any arbitrarily small real numberϵ\>0\\epsilon\>0, it should hold that:
\[hacc\(sT\)∏j=1T−1\(1−hacc\(sj\)\)L\(sT\)\]≤ϵ≤\\displaystyle\[h\_\{acc\}\(s\_\{T\}\)\\prod^\{T\-1\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)L\(s\_\{T\}\)\]\\leq\\epsilon\\leq\[∑i=1T−1hacc\(si\)L\(si\)∏j−1i−1\(1−hacc\(sj\)\]\\displaystyle\[\\sum^\{T\-1\}\_\{i=1\}h\_\{acc\}\(s\_\{i\}\)L\(s\_\{i\}\)\\prod\_\{j\-1\}^\{i\-1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\]Taking the first inequality, the objective is to reach a statesTs\_\{T\}where the followingtermination conditionholds:
\(1\)L\(sT\)≤ϵhacc\(sT\)∏j=1T−1\(1−hacc\(sj\)\)\\displaystyle L\(s\_\{T\}\)\\leq\\frac\{\\epsilon\}\{h\_\{acc\}\(s\_\{T\}\)\\prod^\{T\-1\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}The denominator of the right inequality term involves the multiplication of numbers less than 1\. Thus, this ratio at a time stepTTshall be larger thanL\(sT\)L\(s\_\{T\}\)\.
The following theorem shows that with the objective to minimizeℒT\\mathcal\{L\}\_\{T\}, the termination condition specified by the inequality \(1\) guarantees the termination of the iterative process\.
###### Theorem 3\.1\.
The iterative decision\-making process with the stated termination condition terminates for anyϵ≥0\\epsilon\\geq 0\.
The proof of this theorem \(provided in Appendix B\) shows that the upper bound of the difference in loss in two subsequent loss function updates fromt−1t\-1tottand fromtttot\+1t\+1increases by a factor that is greater or equal to 1, and is inversely proportional to the probability of the human to reject the recommendation made in statests\_\{t\}\.
The additional importance of this theorem is that, any choice ofϵ\\epsilonwill terminate the process, independently of whether the probability of user acceptance to recommendations increases\. Choosing the value ofϵ\\epsilonwe can tune the effectiveness of the process in terms of the maximum rounds of iteration\. However, it must be noted that termination does not implysuccessfulsuccessfultermination: The process may terminate either by accepting an erroneous recommendation that satisfies human constraints \(recall that the human has no information on the correctness of any AI recommendation\) or by not accepting a recommendation that is both correct and satisfies the constraints\.
## 4\.Human\-centric reflective architecture \(HCRA\)
Figure 1\.The overall human\-centric reflective architecture\.As it is shown in Figure[1](https://arxiv.org/html/2607.03025#S4.F1), the proposed human\-centric reflective architecture consists of five functional components: \(a\) those of thereflective agent\(shown in yellow\), i\.e\., the actor and the self reflection LLM, \(b\) those that constitute thehuman behavior model\(shown in red\), i\.e\. the human calibration model and the human acceptance model, and \(c\) theevaluator\(indicated in blue\)\. The long\-term and short\-term memories are experience buffers for the agent, with different time horizons\.
Subsequently, we describe each component and explain their dependencies and their joint operation\. Numbers in Figure[1](https://arxiv.org/html/2607.03025#S4.F1)specify the order of components’ execution in each iteration of the decision\-making process, based on their input requirements and dependencies, as specified by arrows\.
Long term memory: The long term memory stores a long history of past interactions\. Specifically, the long term memory is updated with the content of the short term memory, when the iterative process per request terminates\. This creates a comprehensive knowledge base that informs future decision\-making\.
Short term memory: The short term memory stores the interactions regarding the last request\. Specifically, the short term memory stores per iteration the original question, actor recommendation, evaluator assessments, the actor confidence, and generated reflective text\.
When generating reflective text or recommendations, the self\-reflection component and the actor, respectively, retrieve stored interactions to identify specific patterns enabling targeted feedback generation and recommendations\. As already pointed out in Section 2, memory components facilitate addressing specific deficiencies in making recommendations\.
Actor: This is an LLM that serves as the primary recommendation generator in our reflective decision\-making process\. At each iterationtt, the actor receives as input theQrtQr\_\{t\}, which is the initialQr0Qr\_\{0\}enhanced with the three most recent interactions fetched from the short term memory\. It provides the AI recommendationRetRe\_\{t\}, along with the AI confidence,cftcf\_\{t\}\. The template for the actor prompt is specified in Figure 2\(top\)\.
Figure 2\.Prompt templatesEvaluator: The evaluator is an LLM that assesses the correctness and the agreement of the recommendationRetRe\_\{t\}\. The template for the evaluator prompt is provided in Figure 2 \(bottom\)\. In terms of correctness, the agent evaluates the factual accuracy of theRetRe\_\{t\}in relation toQrtQr\_\{t\}, assigning a score of 1 toCorrt^\\widehat\{\{Corr\_\{t\}\}\}if it assesses that the recommendation is correct; otherwise, it assigns \-1\. To assess agreement, the evaluator focuses exclusively on specified constraintsCsCs\. For each constraint, the evaluator assigns a score of 1 if theRetRe\_\{t\}is assessed to align with it, otherwise it assigns \-1 \. The overall agreementAggrt^\\widehat\{\{Aggr\_\{t\}\}\}is determined to be 1 if assessments in relation to individual constraints are 1; otherwise, it is \-1\.
Self\-reflection: This LLM analyzes the actor’s past recommendations and generates reflective text \(linguistic feedback\) to guide the generation of the next AI recommendation\. It exploits \(a\) the long term memory, fetching per historical request: the question, the final actor recommendation, and the final reflective text, \(b\) the interactions stored in short term memory \(if any\), and \(c\) the lastRetRe\_\{t\}provided by the actor,Corrt^\\widehat\{\{Corr\_\{t\}\}\}provided by the evaluator, andhacc\(st\)h\_\{acc\}\(s\_\{t\}\)provided by the human acceptance model\. The self\-reflection model prompt is specified in Figure 2 \(middle\)\. The feedback targets the improvement of three key aspects of the actor’s behavior in iterationt\+1t\+1: The assessed correctness of the recommendationCorrt\+1^\\widehat\{\{Corr\_\{t\+1\}\}\}, the actor confidencecft\+1cf\_\{t\+1\}, and the probabilityhacc\(st\+1\)h\_\{acc\}\(s\_\{t\+1\}\)of the human to accept the recommendation\. An example of reflective text along with the corresponding recommendation is provided in Figure[3](https://arxiv.org/html/2607.03025#S4.F3)\.
Figure 3\.Example of reflective text for an actor\-generated recommendation\.Human acceptance model: This is part of the human behavior model, and realizes the functionhacch\_\{acc\}\. This model simulates humans evaluatingRetRe\_\{t\}and outputs the probability that humans accept the AI recommendation at time steptt\. It takes as input the human\-calibrated confidence ofRetRe\_\{t\},gtg\_\{t\}, the assessmentAggrt^\\widehat\{\{Aggr\_\{t\}\}\}, together with human demographic features\.
Human calibration model: This model modifies the actor’s confidencecftcf\_\{t\}, and provides the human\-calibrated confidencegtg\_\{t\}forRetRe\_\{t\}\. The goal of training this model is to identify the optimal confidence range that increaseshacc\(st\)h\_\{acc\}\(s\_\{t\}\), with respect to the correctness assessedCorrt^\\widehat\{\{Corr\_\{t\}\}\}ofRetRe\_\{t\}, reflecting the human expectation from the generated recommendation\.
The proposed approach is human\-centric as evidenced by the following: First, the agent action \(the reflective text\) depends on the assessed probability that the AI recommendation is acceptable by humans, and historical human\-AI interactions\. Second, the formulated reward of the agent depends on human\-calibrated expectations \(gtg\_\{t\}\), needs \(QrtQr\_\{t\}\) and specified constraints \(CC\), and third, the objective aims at minimizing the final human lossℒT\\mathcal\{L\}\_\{T\}over a horizon ofTTiterations\.
### 4\.1\.The overall human\-calibrated decision\-making process
As specified in Figure[1](https://arxiv.org/html/2607.03025#S4.F1)by the ordering of components execution, initially, the human provides the questionQr0Qr\_\{0\}, with constraintsCsCsthat the final decision must satisfy\. WhileCsCsremain constant throughout the iterative process, in subsequent iterationst\>0t\>0,QrtQr\_\{t\}is enhanced with the three most recent interactions stored in short term memory, and the reflective text\.
GivenQrtQr\_\{t\}andCsCs, an iteration of the reflective process involves the execution of the following components in order:
\(1\) The actor LLM generates the recommendation\(Ret,cft\)\(Re\_\{t\},cf\_\{t\}\), givenQrtQr\_\{t\}\.
\(2\) The evaluator takes this recommendation,QrtQr\_\{t\}andCsCs, and providesCorrt^\\widehat\{\{Corr\_\{t\}\}\}andAggrt^\\widehat\{\{Aggr\_\{t\}\}\}\. WhileCorrt^\\widehat\{\{Corr\_\{t\}\}\}is exploited by the human\-calibration model and the self\-reflection model, theAggrt^\\widehat\{\{Aggr\_\{t\}\}\}is exploited by the human acceptance model\.
\(3\) The human calibration model provides the human\-calibrated confidencegtg\_\{t\}, which is exploited by the human acceptance model\.
\(4\) The human acceptance model predicts the recommendation acceptance probabilityhacch\_\{acc\}, which is provided to the self reflection model\.
\(5\) The self reflection model, givenRet,Corrt^,hacc\(st\)Re\_\{t\},\\widehat\{\{Corr\_\{t\}\}\},h\_\{acc\}\(s\_\{t\}\), the interactions stored in the short term memory and the last interactions per request stored in the long term memory, generates the reflective text\.
\(6\) The short term memory is then updated to include the new interaction, comprisingQrt,\(Ret,cft\),Corrt^,Aggrt^Qr\_\{t\},\(Re\_\{t\},cf\_\{t\}\),\\widehat\{\{Corr\_\{t\}\}\},\\widehat\{\{Aggr\_\{t\}\}\}, and the generated reflective text\.
In this work, humans are represented by behavioral models\. While humans can theoretically actively participate in the loop by providing feedback on the actor’s recommendations—thereby tuninggtg\_\{t\}and the acceptance probabilityhacc\(st\)h\_\{acc\}\(s\_\{t\}\)—this study utilizes a generalized human behavior model trained on real\-world human data, as detailed in the next section\. Actively involving human participants and leveraging live interventions during the iterative process are beyond the scope of this work and are reserved for future research\.
## 5\.Training the models
The training process involves the training of the human behavior models, and the training of the reflective agent\. The training of the human behavior models follows a two\-stages training process: We first train the human acceptance model, and then we exploit this well trained model to train the human calibration model\.


Figure 4\.The effect of \(left\) agreement onhacch\_\{acc\}and \(right\) of correctness on human\-calibrated confidenceg\(st\)g\(s\_\{t\}\)\.Human acceptance model: This model is a three\-layer fully connected neural network \(10×24, 24×12, and 12×1\) with ReLU activation\. It is trained using binary cross\-entropy loss, along with the Adam optimizer\(Kingma and Ba,[2017](https://arxiv.org/html/2607.03025#bib.bib28)\)and early stopping to prevent overfitting\. To train the model we construct a dataset using data from the dataset described in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)\. The constructed dataset is inclusive with interactions from four tasks \(art, cities, sarcasm, census\) and many humans from multiple countries and continents\. Important details about the dataset, the features used, and its use for the training of models are included in Appendix A\.
Towards a general and inclusive human behavior model for the purposes of this study, we use the entire constructed dataset\. The human acceptance model achieves a ROC\-AUC score of 0\.78 on the test set of unseen interactions\. Although differing training configurations prevent a direct comparison, this result closely aligns with the 0\.81 ROC\-AUC baseline reported for the activation model in\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\)\.
Figure[4](https://arxiv.org/html/2607.03025#S5.F4)\(left\) presents a partial dependence plot for the human acceptance model\. The plot shows the average predicted acceptance probability on the test set of unseen interactions when agreement is fixed across samples\. Positive \(negative\) values in x\-axis indicate agreement \(resp\. disagreement\) with constraints\. The absolute values in the x\-axis indicate the “raw” confidence \(cftcf\_\{t\}\) of the AI recommendation in time steptt\. The plot reveals that humans accept \(reject\) recommendations that respect \(resp\. disagree with\) their constraints\. The behavior of the model is intuitive, considering that humans evaluate recommendations mainly based on agreement with constraints, possessing no information regarding correctness\. Details about the dataset are in Appendix A\.1\.
Human calibration model:Similarly to how AI confidence is treated in\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\), we consider the inverse sigmoid ofht=\(Corrt^⋅cft\)h\_\{t\}=\(\\widehat\{Corr\_\{t\}\}\\cdot cf\_\{t\}\)given the recommendation\(Ret,cft\)\(Re\_\{t\},cf\_\{t\}\)in iterationttand the assessed correctness by the evaluator\. Thus, the human\-calibration confidence model optimizes the functiong:\[0,1\]×\{−1,1\}→\[0,1\]g:\[0,1\]\\times\\\{\-1,1\\\}\\rightarrow\[0,1\]that maps \(cftcf\_\{t\},Corr^t\\widehat\{Corr\}\_\{t\}\) to \[0,1\]\. The functionggis of the following form:
gt=g\(cft,Corr^t\)g\_\{t\}=g\(cf\_\{t\},\\widehat\{Corr\}\_\{t\}\)=11\+e−sign\(Corr^t\)\(α⋅sign\(Corr^t\)⋅ht\+β\)=\\frac\{1\}\{1\+e^\{\-sign\(\\widehat\{Corr\}\_\{t\}\)\(\\alpha\\cdot sign\(\\widehat\{Corr\}\_\{t\}\)\\cdot h\_\{t\}\+\\beta\)\}\}=11\+e−Corr^t\(α⋅cft\+β\)=\\frac\{1\}\{1\+e^\{\-\\widehat\{Corr\}\_\{t\}\(\\alpha\\cdot cf\_\{t\}\+\\beta\)\}\}
whereα,β∈ℝ≥0\\alpha,\\beta\\in\\mathbb\{R\}\_\{\\geq 0\}\. In the case where\(α,β\)=\(1,0\)\(\\alpha,\\beta\)=\(1,0\)thenggis the sigmoid function onhth\_\{t\}\. Similarly to\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\), whileα\\alphamodulates the rate at whichggincreases as the confidence of the recommendation increases,β\\betaadjusts the minimumgg\. Providing the human calibration model with the correctness assessmentCorr^\\widehat\{Corr\}is intuitive, since humans also consider correctness when assessing the confidence and determining their subsequent reaction to the AI proposal\.
Given the well\-trained human acceptance model and the dataset used for training that model, we optimize the parameters\(α,β\)\(\\alpha,\\beta\)ofggto minimize the expected loss:
\(α∗,β∗\)=argminα,β𝔼t\[−log\(hacc\(st\)\)\]\(\\alpha^\{\*\},\\beta^\{\*\}\)=\\arg\\min\_\{\\alpha,\\beta\}\\;\\mathbb\{E\}\_\{t\}\[\-\\log\\big\(h\_\{\\text\{acc\}\}\(s\_\{t\}\)\\big\)\]
We finally use\|α\|\|\\alpha\|,\|β\|\|\\beta\|\. Using this loss function we aim to penalize the model more whenhacc\(s\)h\_\{acc\}\(s\)is low, rather whenhacc\(s\)h\_\{acc\}\(s\)is high\. In this way, we encourage the model to adjust its confidence in order to increase the likelihood of the user to accept the recommendation given the human\-calibrated confidence\. Given the optimizedgg, the percentage of accepted recommendations to the total number of interactions in the test dataset is 51%, compared to 49% when the confidence is not human\-calibrated\.
Figure[4](https://arxiv.org/html/2607.03025#S5.F4)\(right\) illustrates the behavior of the trained human calibration model, mapping the relationship between the assessed correctness of recommendations and AI confidence values\. Specifically, the model scales up the confidence scores for correct recommendations \(represented by positive x\-axis values\) within moderate AI confidence intervals\. This indicates that the model successfully identifies a specific confidence range where users are most receptive to accurate suggestions\. Conversely, incorrect recommendations \(negative x\-axis values\) consistently yield low scores\(<0\.4\)\(<0\.4\), demonstrating that the calibration model effectively penalizes highly confident but incorrect AI outputs\.
Self\-Reflection model:HCRA operates according to the refle\-xion\-based paradigm that enables language agents to provide linguistic feedback supporting the generation of correct advices\. Although the actor serves as the primary recommendation generator, the self\-reflection model guides the process providing reflective text, towards enhancing the efficiency of the actor and minimizing the loss of the human\. The self\-reflection LLM model generates linguistic feedback, crucially exploiting the short and long term memories which accumulate experience of past interactions\. This experience enables the AI agent to be increasingly effective to recommendation and feedback provision considering reward signals and patterns across similar scenarios\.
## 6\.Experiments
We evaluate our framework within the domain of touristic decision\-making\. This environment can be highly critical due to strict human requirements and constraints\. These include budget limits, accessibility needs, dietary restrictions, and rigid transportation schedules tied to specific events\. Furthermore, safety constraints are paramount, as recommendations must actively prevent users from navigating unsafe areas\.
Experiments aim to show that: \(a\) HCRA\-driven decision\-making is more effective than a reflective agentic architecture without models of human behavior; \(b\) the human\-centric aspects of HCRA play a significant role in effectiveness of the collaborative approach; \(c\) the proposed architecture facilitates high\-quality recommendations when the actor is able to take advantage of an exploratory response\-generation process, without being overly non\-deterministic; \(d\) the role of long\-term memory is crucial for the agent to be effective in answering new requests\.
We measure the quality of recommendations in terms of successful terminations, depending on factual correctness \(which neither the AI nor the human models can access\) and agreement to constraints\. A termination is considered successful if at the final time step the acceptance probability is greater than0\.50\.5, the recommendation is correct and in agreement to the constraints\.
Effectiveness is measured by means of \(a\) the number of iterations needed for the iterative process to terminate, \(b\) the percentage of successful terminations to the total number of terminations\.
Furthermore, we provide human loss at the final iteration, demonstrating that behavioral models successfully minimize human loss while simultaneously enhancing recommendation effectiveness\.
We have formed 32 questions at different levels of difficulty, depending on the scarcity of information required to form a decision\. The iterative process for answering an individual question is a “trial”\. Each trial requires a number of iterations until reaching the terminating condition, and can terminate either successfully or not\. We execute 10 independent runs, each one involving one trial per question \(i\.e\., 32 trials per run\)\. At the beginning of a run, we randomly sample human demographic features \(age, gender, socioeconomic status, education, programming experience, and AI preference scores\) from the human\-behaviour models training dataset\. These features apply to all trials performed in the run\. In addition, human constraints are randomly selected per trial by multiple criteria for food, transportation, accommodation, activities, shopping, nightlife, and budget attributes\. In addition to these 32 questions we have formed 10 questions that are either similar to some of the 32, or more complex, in the sense that their recommendation involves the combination of multiple decisions\. Using these 10 questions, we evaluate the effectiveness of HCRA to answer questions exploiting historical interactions stored in the long\-term memory\. In our case these historical interactions are gathered after responding to the 32 questions\. All questions are specified in Appendix C\.
The implemented HCRA uses the DeepSeek\-V3\-0324 model playing the roles of the actor, evaluator and self\-reflection model\. We vary the temperatures of the actor and of the evaluator, as specified in different experimental settings, to evaluate the exploratory nature of the actor and the effect of noisy evaluator assessments\. Theϵ\\epsilonparameter of the iterative process termination condition has been set to 0\.01\.
### 6\.1\.Experimental results
First, to evaluate the impact of the actor model temperature parameter and investigate the role of the actor temperature to steer the quality of decision of HCRA, we performed experiments using different actor temperature values\. We report on the effectiveness of HCRA in terms of success rate and average number of iterations \(Avg Iters\), and we provide the average loss in the last iteration \(Avg Loss\) of the iterative process\.
Table 1\.Performance with different actor temperaturesTable 1 summarizes the performance metrics across different configurations, with comprehensive data compiled in Appendix C\.3\. The empirical results demonstrate that HCRA exhibits notable robustness to changes in actor temperature, maintaining a high success rate across most settings\. The architecture struggles only at an extreme temperature of 1\.9, where it fails to guide the generation process toward successful outcomes\. Notably, setting the temperature below 1\.0 increases both the final average loss and the mean number of iterations compared to a baseline temperature of 1\.0\. Conversely, raising the temperature to 1\.5 further minimizes the final iteration loss, though this occurs at the expense of an increased iteration count and a slightly lower overall success rate\. With a high temperature \(equal to 1\.9\) HCRA is unable to steer the process towards successful responses, while both, the average loss and the average number of iterations increase\. These results show that temperatures higher or lower than 1 fail to steer decision making effectively\. For temperatures lower than 1, the actor becomes more rigid and struggles to incorporate the received feedback, resulting in an increased number of iterations and high average loss\. Similarly, for temperatures higher than 1, the actor struggles to respond successfully due to non\-determinism, even after receiving the reflective text, resulting in an increased number of iterations\. Therefore temperature equal to 1 allows the agent to steer the decision making process effectively, taking advantage of an exploratory process for the generation of responses\. These conjectures are further supported by results reported in Figures[13](https://arxiv.org/html/2607.03025#A3.F13),[16](https://arxiv.org/html/2607.03025#A3.F16),[19](https://arxiv.org/html/2607.03025#A3.F19),[22](https://arxiv.org/html/2607.03025#A3.F22),[25](https://arxiv.org/html/2607.03025#A3.F25),[28](https://arxiv.org/html/2607.03025#A3.F28),[31](https://arxiv.org/html/2607.03025#A3.F31)in Appendix C\.3\.
To delve into the results for actor temperature equal to 1, we provide results regarding the effectiveness of the process in terms of the successful trials and the number of iterations that they require: Figure[5](https://arxiv.org/html/2607.03025#S6.F5)provides the proportion of successful trials \(red bars\) to the total number of trials \(gray bars\) in relation to the required number of iterations \(x\-axis\)\. Results show the effectiveness of HCRA: It achieves an overall success rate of 58\.4%, with the majority of trials \(57\.5%\) reaching termination at the third or fourth iteration\. It must be noted that the x\-axis starts from 2, given that no trials ended earlier\. Notably, 85\.7% of 3\-iteration trials and 86\.0% of 4\-iteration trials resulted in successful terminations\. The distribution shows a preference for early termination when recommendations meet acceptance criteria, with more iterations \(5–10\) representing difficult cases requiring extensive refinement of recommendations\. Based on this performance, the actor temperature is set equal to 1 for all subsequent experiments\.
Figure 5\.Distribution of total vs successful trials \(actor temperature=1\.0\)\.Focusing on the required number of iterations per question, Figure[6](https://arxiv.org/html/2607.03025#S6.F6)shows the number of iterations required per question \(sorted by difficulty left to right\) in 10 independent trials and in the form of boxplots\. This reveals substantial variation in iteration requirements across different questions, with most of the trials requiring at most 8 iterations and 4 outliers requiring at most 10\. The average number of iterations for all trials is 4\.78\. The results show consistent behaviour with a median number of 3\.5–6\.5 iterations for most questions\.
Figure 6\.Iterations required per question \(actor temperature =1\.0\)\.Figure[7](https://arxiv.org/html/2607.03025#S6.F7)shows the number of iterations required for the successful trials per question\. As shown, these trials require an average number of 3\.84 iterations\. This is an improvement over the average number of iterations for all trials, suggesting that successful trials tend to terminate faster than unsuccessful ones, which is quite reasonable considering that unsuccessful trials should imply a kind of difficulty\. The median number of iterations across most questions ranges from 3\.0 to 4\.0, indicating robust HCRA behavior regardless of question difficulty, with only a subset of 4 questions requiring a higher median number of iterations \(up to 6\.0\), and one question reaching the maximum number of 10 iterations\.
Figure 7\.Iterations required for successful termination of questions \(actor temperature=1\.0\)\.Table 2\.Performance of the trained reflective model\.Table[2](https://arxiv.org/html/2607.03025#S6.T2)shows the effectiveness of the reflective process in answering the additional 10 questions, taking advantage of the interactions gathered when it iterated to answer the 32 questions\. The results show that HCRA manages to answer these new and complex questions effectively, significantly reducing the average number of iterations \(4\.32\), while significantly increasing the percentage of successful terminations \(75%\), compared to the scores achieved when responding without exploiting past experience\.
### 6\.2\.Ablation Study
The contribution of the human calibration model:In the absence of the human calibration model, the reflective process relies directly on the AI recommendation confidencecftcf\_\{t\}\. As shown in Table[3](https://arxiv.org/html/2607.03025#S6.T3), this yields degraded performance, compared to the scores reported in Table[1](https://arxiv.org/html/2607.03025#S6.T1)for actor temperature set to 1\.0\. As shown, the success rate drops to 53\.8% with an increased average number of iterations \(5\.02\)\. The average loss value remains relatively stable in this setting compared to those in Table[1](https://arxiv.org/html/2607.03025#S6.T1)\. These results suggest that the human calibration model contributes to the overall HCRA performance, steering the reflective process towards more effective behavior\. Detailed results for this setting are shown in Appendix C\.3, Figures[28](https://arxiv.org/html/2607.03025#A3.F28),[29](https://arxiv.org/html/2607.03025#A3.F29),[30](https://arxiv.org/html/2607.03025#A3.F30)\.
Table 3\.Performance without human calibrationThe impact of an unreliable evaluator:In relation to the role of the evaluator, Table[4](https://arxiv.org/html/2607.03025#S6.T4)reports on the performance of HCRA with an evaluator that makes unreliable assessments: At each iteration, with probability 0\.5, a random noise term drawn uniformly from\[−1,1\]\[\-1,1\]is added to the evaluator’s correctness probability assessment, clipping the result to \[0,1\]\. The result is compared to 0\.7 to make the final boolean correctness assessment\. The value of the agreement assessment is also changed from 1 to−1\-1or vice versa\. The reported results take into account only the original assessments of the evaluator since these reflect the actor’s true performance\. The significant drop in success rate highlights the crucial role of the evaluator in HCRA: An unreliable evaluator adds noise in the reflective process, failing to steer HCRA to correct recommendations\. Detailed results for this setting are shown in Appendix C\.3, Figures[31](https://arxiv.org/html/2607.03025#A3.F31),[32](https://arxiv.org/html/2607.03025#A3.F32),[33](https://arxiv.org/html/2607.03025#A3.F33)\.
Table 4\.Performance with an unreliable evaluatorThe HCRA without the human acceptance model is the baseline architecture whose evaluation is reported in the next paragraph\.
Baseline comparison:To demonstrate the effectiveness of HCRA, we compare against a baseline system that uses factual correctness as the termination criterion, without exploiting human models and the human loss\. Figure[8](https://arxiv.org/html/2607.03025#S6.F8)reveals the dramatic difference in termination behavior: The baseline system exhibits premature termination with 50% of trials \(161/320\) terminating after just one iteration and 69% completing within two iterations\. This contrasts sharply with the results reported for HCRA where the majority \(57,5%\) of trials terminate at the third or fourth iteration\. The baseline rapid termination is due \(a\) to the lack of any kind of assessment of recommendations’ agreement to human constraints, and \(b\) to the fact that the probability that the human accepts the recommendation is not exploited in any way towards the final decision\. While this approach achieves faster convergence \(3\.50 iterations on average\), it sacrifices human collaboration and recommendation quality with respect to human constraints, while, due to the lack of human models, recommendations are not guaranteed to be acceptable by humans\.
Figure 8\.Iterations required for termination when using reflection and factual correctness as termination criterion\.
## 7\.Concluding remarks
Aiming to improve human\-AI collaboration, this work formulates the human\-centric AI\-assisted decision\-making process as a stochastic game played between the AI agent and a human\. Based on this formulation we propose HCRA that leverages models of human behaviour in a reflective process towards refining the recommendations provided by the AI agent according to human objectives and constraints\. Experimental results show that the integration of human\-calibrated AI models contributes to providing successful recommendations to humans effectively, i\.e\., in few iterations, favoring explorative behavior and flexibility in responding\. To ensure utility of the HCRA architecture in specific real\-life settings that necessitate the modeling of idiosyncratic human behavior, it needs the following: \(a\) A highly tuned model of humans’ behavior, trained in the domain and context of system use, with data from the humans that collaborate with the system; \(b\) a language agent that is able to provide domain\-specific recommendations and reflect on them effectively\. In settings where fast adaptation to humans is necessary, human behavior models should be trained incrementally to improve their abilities in their context of use, leveraging humans’ reactions\.
Regarding future work, although we do not consider long\-horizon tasks, we believe that our work can be extended to include such tasks\. A recent survey on RL for long horizon interactive LLM agents can be found in\(Chenet al\.,[2025](https://arxiv.org/html/2607.03025#bib.bib24)\)\. Finally, involving humans in the process, in real\-life contexts, is critical to further enhance and assess the utility of the proposed architecture\.
## References
- G\. Bansal, B\. Nushi, E\. Kamar, E\. Horvitz, and D\. S\. Weld \(2021a\)Is the most accurate ai the best teammate? optimizing ai for teamwork\.Proceedings of the AAAI Conference on Artificial Intelligence35\(13\),pp\. 11405–11414\.External Links:[Link](https://ojs.aaai.org/index.php/AAAI/article/view/17359),[Document](https://dx.doi.org/10.1609/aaai.v35i13.17359)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1),[§2](https://arxiv.org/html/2607.03025#S2.p2.1)\.
- G\. Bansal, B\. Nushi, E\. Kamar, W\. S\. Lasecki, D\. S\. Weld, and E\. Horvitz \(2019\)Beyond accuracy: the role of mental models in human\-ai team performance\.Proceedings of the AAAI Conference on Human Computation and Crowdsourcing7\(1\),pp\. 2–11\.External Links:[Link](https://ojs.aaai.org/index.php/HCOMP/article/view/5285),[Document](https://dx.doi.org/10.1609/hcomp.v7i1.5285)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- G\. Bansal, T\. Wu, J\. Zhou, R\. Fok, B\. Nushi, E\. Kamar, M\. T\. Ribeiro, and D\. Weld \(2021b\)Does the whole exceed its parts? the effect of ai explanations on complementary team performance\.InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems,CHI ’21,New York, NY, USA\.External Links:ISBN 9781450380966,[Link](https://doi.org/10.1145/3411764.3445717),[Document](https://dx.doi.org/10.1145/3411764.3445717)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- E\. Beede, E\. Baylor, F\. Hersch, A\. Iurchenko, L\. Wilcox, P\. Ruamviboonsuk, and L\. M\. Vardoulakis \(2020\)A human\-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy\.InProceedings of the 2020 CHI Conference on Human Factors in Computing Systems,CHI ’20,New York, NY, USA,pp\. 1–12\.External Links:ISBN 9781450367080,[Link](https://doi.org/10.1145/3313831.3376718),[Document](https://dx.doi.org/10.1145/3313831.3376718)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p1.1)\.
- N\. L\. C\. Benz and M\. G\. Rodriguez \(2024\)Human\-aligned calibration for ai\-assisted decision making\.External Links:2306\.00074,[Link](https://arxiv.org/abs/2306.00074)Cited by:[§2](https://arxiv.org/html/2607.03025#S2.p3.1)\.
- J\. Y\. Bo, S\. Wan, and A\. Anderson \(2025\)To rely or not to rely? evaluating interventions for appropriate reliance on large language models\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,CHI ’25,New York, NY, USA\.External Links:ISBN 9798400713941,[Link](https://doi.org/10.1145/3706598.3714097),[Document](https://dx.doi.org/10.1145/3706598.3714097)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- Z\. Buçinca, P\. Lin, K\. Z\. Gajos, and E\. L\. Glassman \(2020\)Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems\.InProceedings of the 25th International Conference on Intelligent User Interfaces,IUI ’20,New York, NY, USA,pp\. 454–464\.External Links:ISBN 9781450371186,[Link](https://doi.org/10.1145/3377325.3377498),[Document](https://dx.doi.org/10.1145/3377325.3377498)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- K\. Chen, M\. Cusumano\-Towner, B\. Huval, A\. Petrenko, J\. Hamburger, V\. Koltun, and P\. Krähenbühl \(2025\)Reinforcement learning for long\-horizon interactive llm agents\.External Links:2502\.01600,[Link](https://arxiv.org/abs/2502.01600)Cited by:[§7](https://arxiv.org/html/2607.03025#S7.p2.1)\.
- U\. Fischer\-Abaigar, C\. Kern, N\. Barda, and F\. Kreuter \(2024\)Bridging the gap: towards an expanded toolkit for ai\-driven decision\-making in the public sector\.Government Information Quarterly41\(4\),pp\. 101976\.External Links:ISSN 0740\-624X,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.giq.2024.101976),[Link](https://www.sciencedirect.com/science/article/pii/S0740624X24000686)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p1.1)\.
- A\. Jaech and et al\. \(2024\)OpenAI o1 system card\.External Links:2412\.16720,[Link](https://arxiv.org/abs/2412.16720)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- D\. P\. Kingma and J\. Ba \(2017\)Adam: a method for stochastic optimization\.External Links:1412\.6980,[Link](https://arxiv.org/abs/1412.6980)Cited by:[§5](https://arxiv.org/html/2607.03025#S5.p2.1)\.
- V\. Lai, C\. Chen, A\. Smith\-Renner, Q\. V\. Liao, and C\. Tan \(2023\)Towards a science of human\-ai decision making: an overview of design space in empirical human\-subject studies\.InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency,FAccT ’23,New York, NY, USA,pp\. 1369–1385\.External Links:ISBN 9798400701924,[Link](https://doi.org/10.1145/3593013.3594087),[Document](https://dx.doi.org/10.1145/3593013.3594087)Cited by:[§2](https://arxiv.org/html/2607.03025#S2.p1.1)\.
- Z\. Li and M\. Yin \(2025\)Utilizing human behavior modeling to manipulate explanations in ai\-assisted decision making: the good, the bad, and the scary\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- H\. Liu, V\. Lai, and C\. Tan \(2021\)Understanding the effect of out\-of\-distribution examples and interactive explanations on human\-ai decision making\.Proc\. ACM Hum\.\-Comput\. Interact\.5\(CSCW2\)\.External Links:[Link](https://doi.org/10.1145/3479552),[Document](https://dx.doi.org/10.1145/3479552)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- S\. Ma, Y\. Lei, X\. Wang, C\. Zheng, C\. Shi, M\. Yin, and X\. Ma \(2023\)Who should i trust: ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai\-assisted decision\-making\.InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems,CHI ’23,New York, NY, USA\.External Links:ISBN 9781450394215,[Link](https://doi.org/10.1145/3544548.3581058),[Document](https://dx.doi.org/10.1145/3544548.3581058)Cited by:[§2](https://arxiv.org/html/2607.03025#S2.p2.1)\.
- L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. Lowe \(2022\)Training language models to follow instructions with human feedback\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- G\. Papadopoulos, A\. Bastas, G\. A\. Vouros, I\. Crook, N\. Andrienko, G\. Andrienko, and J\. M\. Cordero \(2024\)Deep reinforcement learning in service of air traffic controllers to resolve tactical conflicts\.Expert Systems with Applications236,pp\. 121234\.External Links:ISSN 0957\-4174,[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.eswa.2023.121234),[Link](https://www.sciencedirect.com/science/article/pii/S0957417423017360)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. Yao \(2023\)Reflexion: language agents with verbal reinforcement learning\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§C\.3\.2](https://arxiv.org/html/2607.03025#A3.SS3.SSS2.p5.1),[§2](https://arxiv.org/html/2607.03025#S2.p5.1),[§2](https://arxiv.org/html/2607.03025#S2.p6.8),[§2](https://arxiv.org/html/2607.03025#S2.p7.1)\.
- C\. Snell, J\. Lee, K\. Xu, and A\. Kumar \(2024\)Scaling llm test\-time compute optimally can be more effective than scaling model parameters\.External Links:2408\.03314,[Link](https://arxiv.org/abs/2408.03314)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- T\. Takayanagi, R\. Hashimoto, C\. Chen, and K\. Izumi \(2025\)The impact and feasibility of self\-confidence shaping for ai\-assisted decision\-making\.External Links:2502\.14311,[Link](https://arxiv.org/abs/2502.14311)Cited by:[§2](https://arxiv.org/html/2607.03025#S2.p2.1)\.
- H\. Vasconcelos, M\. Jörke, M\. Grunde\-McLaughlin, T\. Gerstenberg, M\. S\. Bernstein, and R\. Krishna \(2022\)Explanations can reduce overreliance on AI systems during decision\-making\.CoRRabs/2212\.06823\.External Links:[Link](https://doi.org/10.48550/arXiv.2212.06823),[Document](https://dx.doi.org/10.48550/ARXIV.2212.06823),2212\.06823Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- K\. Vodrahalli, R\. Daneshjou, T\. Gerstenberg, and J\. Zou \(2022a\)Do humans trust advice more if it comes from ai? an analysis of human\-ai interactions\.InProceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society,AIES ’22,New York, NY, USA,pp\. 763–777\.External Links:ISBN 9781450392471,[Link](https://doi.org/10.1145/3514094.3534150),[Document](https://dx.doi.org/10.1145/3514094.3534150)Cited by:[Figure 9](https://arxiv.org/html/2607.03025#A1.F9),[§A\.1](https://arxiv.org/html/2607.03025#A1.SS1.p2.1),[§A\.1](https://arxiv.org/html/2607.03025#A1.SS1.p4.3),[§A\.1](https://arxiv.org/html/2607.03025#A1.SS1.p5.2),[§A\.2](https://arxiv.org/html/2607.03025#A1.SS2.p2.1),[Table 5](https://arxiv.org/html/2607.03025#A1.T5),[§5](https://arxiv.org/html/2607.03025#S5.p2.1)\.
- K\. Vodrahalli, T\. Gerstenberg, and J\. Zou \(2022b\)Uncalibrated models can improve human\-ai collaboration\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§A\.1](https://arxiv.org/html/2607.03025#A1.SS1.p4.3),[§1](https://arxiv.org/html/2607.03025#S1.p2.1),[§1](https://arxiv.org/html/2607.03025#S1.p4.1),[§2](https://arxiv.org/html/2607.03025#S2.p2.1),[§2](https://arxiv.org/html/2607.03025#S2.p4.1),[§5](https://arxiv.org/html/2607.03025#S5.p3.1),[§5](https://arxiv.org/html/2607.03025#S5.p5.7),[§5](https://arxiv.org/html/2607.03025#S5.p6.9)\.
- X\. Wang and M\. Yin \(2021\)Are explanations helpful? a comparative study of the effects of explanations in ai\-assisted decision\-making\.InProceedings of the 26th International Conference on Intelligent User Interfaces,IUI ’21,New York, NY, USA,pp\. 318–328\.External Links:ISBN 9781450380171,[Link](https://doi.org/10.1145/3397481.3450650),[Document](https://dx.doi.org/10.1145/3397481.3450650)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- F\. Xu, Q\. Hao, Z\. Zong, J\. Wang, Y\. Zhang, J\. Wang, X\. Lan, J\. Gong, T\. Ouyang, F\. Meng, C\. Shao, Y\. Yan, Q\. Yang, Y\. Song, S\. Ren, X\. Hu, Y\. Li, J\. Feng, C\. Gao, and Y\. Li \(2025\)Towards large reasoning models: a survey of reinforced reasoning with large language models\.External Links:2501\.09686,[Link](https://arxiv.org/abs/2501.09686)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p4.1)\.
- P\. Zhang \(2023\)Taking advice from chatgpt\.arXiv preprint arXiv:2305\.11888\.Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
- Y\. Zhang, Q\. V\. Liao, and R\. K\. E\. Bellamy \(2020\)Effect of confidence and explanation on accuracy and trust calibration in ai\-assisted decision making\.InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency,FAT\* ’20,New York, NY, USA,pp\. 295–305\.External Links:ISBN 9781450369367,[Link](https://doi.org/10.1145/3351095.3372852),[Document](https://dx.doi.org/10.1145/3351095.3372852)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p2.1)\.
- D\. M\. Ziegler, N\. Stiennon, J\. Wu, T\. B\. Brown, A\. Radford, D\. Amodei, P\. Christiano, and G\. Irving \(2020\)Fine\-tuning language models from human preferences\.External Links:1909\.08593,[Link](https://arxiv.org/abs/1909.08593)Cited by:[§1](https://arxiv.org/html/2607.03025#S1.p3.1)\.
APPENDIX
## Appendix ADataset and Training of Models
### A\.1\.Dataset
The human behavior models are crucial components of HCRA\. In this section we present details for the dataset used for training these models\.
Our starting point is the dataset described in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\), which includes samples from real human responses\.
For the purposes of our work we selected a subset of samples from that dataset, where the term “sample” denotes a single interaction comprising a question, an AI recommendation, and a person response\. We consider samples of “agreement” and “disagreement”, considering ”agreement” as specified in the main part of the article with respect to person constraints\. The terms “acceptance” and “rejection” indicate the person’s reaction to the provided recommendation\. Below we specify how these samples have been selected\.
To determine cases of acceptance using this dataset, we computed the absolute difference between the featuresresponse1response\_\{1\}andresponse2response\_\{2\}\(i\.e\. the difference in confidence between the first and the final person response, as provided in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)\)\. When this difference is greater than0\.0350\.035then samples were labeled with “accept”; otherwise, they were labeled with “reject”\. This is justified in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)and further evidence is provided in\(Vodrahalliet al\.,[2022b](https://arxiv.org/html/2607.03025#bib.bib10)\)by the fact that this difference indicates whether persons integrate the AI recommendation in their own response\. Actually, this happens when \(a\) the AI recommendation is provided with high confidence and is opposite to the person’s response, and \(b\) in case the recommendation and the person’s initial response share the same label, the person’s initial response has low confidence and the AI has high confidence\.
To determine cases of \(dis\)agreement we are using the featuresadviceadviceandresponse1response\_\{1\}as defined in the dataset\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)\. The sign of each of these features shows whether the corresponding choice is \(in\)correct in a binary classification task\. When both features have the same sign, this indicates that they both, AI and the person, selected the \(in\)correct answer, which we classify as “agreement”\. Different signs indicate that one choice was correct, while the other was incorrect, which we classify as disagreement\.
Table 5\.Distribution of samples in the original dataset provided in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)in the different cases of \(non\)acceptance and \(dis\)agreement\.Table[5](https://arxiv.org/html/2607.03025#A1.T5)presents the number of valid samples in the original dataset in different cases of \(non\)acceptance and \(dis\)agreement\. A sample is considered valid if it includes non\-null values in the features we use as inputs to the human\-behavior model \(these features are specified subsequently in A\.2\)\.
Table 6\.Distribution of samples in our balanced dataset\.As it can be observed, the dataset is unbalanced as far as agreement/ disagreement is concerned, and the proportion of samples in the disagreement/acceptance class is larger to those in the disagreement/rejection\. Through a series of experiments, we observed that this strongly influences the predicted human acceptance probability to the provided AI recommendations under the assumptions made in this work\. Therefore, we created a new version of the dataset in which, the cases for agreement and disagreement are balanced and there is also a balance in the distribution of samples with labels “accept” and “reject” in each case\. This has been done in a meticulous manner as it is described subsequently, so as to support in our setting the tendency of persons to evaluate AI responses primarily based on agreement with the constraints, rather than on correctness, given that, as assumed, humans do not have access to any hint or information concerning the ground correctness of AI responses\.
To construct such a balanced dataset with the largest possible number of samples, we started from the under\-represented case, namely the case disagreement/reject that contains 3152 samples\. Using this as a reference, we selected the same number of samples for the agreement/accept case\. To allocate samples to the cases we experimented with different ratios of samples to the disagreement/reject and agreement/accept cases compared to their complimentary ones, and we adopted a ratio 75/25, where 75% corresponds to the majority classes \(disagreement/reject and agreement/accept\) with 3152 samples each, and 25% to the complementary cases \(agreement/reject and disagreement/accept\) with 1050 samples each\.
Figure 9\.Distribution of samples in tasks, in the original dataset\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\)Figure 10\.Distribution of samples in tasks, in our datasetAiming to aggregate training towards a general and inclusive human\-behavior model, we selected samples from all available tasks \(art, sarcasm, cities, census\)\. With respect to the 75/25 ratio we ensure that the proportion of samples per task is maintained in all four classes \(agreement/accept, agreement/reject, disagreement/accept, and disagreement/reject\) as illustrated in Figures[9](https://arxiv.org/html/2607.03025#A1.F9),[10](https://arxiv.org/html/2607.03025#A1.F10), and is similar to their proportion in the original dataset\. Moreover, we ensured that the final dataset includes samples from persons across multiple countries and continents, thereby preserving diversity in demographic features\. The relevant statistics are shown in Figure[11](https://arxiv.org/html/2607.03025#A1.F11)\.
Figure 11\.Distribution of samples in geographical regions, in our datasetTable 7\.Country names and their abbreviationsFigure[12](https://arxiv.org/html/2607.03025#A1.F12)shows the distribution of the AI recommendation confidence for the samples labeled with “accept” \(left\) and “reject” \(right\)\. The distributions, although with differences, indicate that all confidence levels are represented in a similar manner in the two classes\.
Figure 12\.Distribution of recommendation confidence in our dataset for the cases where it is accepted \(left\) and when it is rejected \(right\)
### A\.2\.Input features for the human behavior model
In this section, we present the features exploited by the human behavior models, i\.e\., the human acceptance and the human calibration models\.
Table 8\.Input features for human acceptance modelTable[8](https://arxiv.org/html/2607.03025#A1.T8)specifies the features used as input to the human acceptance model\. Features 1\-7 represent person\-specific features that remain constant for a person across all interactions, while features 8\-10 are state\-dependent features that change during the iterative reflective process\. Features 3, 4, 6 and 7 are as specified in\(Vodrahalliet al\.,[2022a](https://arxiv.org/html/2607.03025#bib.bib27)\), preserving their domains and scales\. More specifically:
- •Feature 3: Socioeconomic status is a self\-assessed feature measured on a 1–10 scale, where 10 indicates individuals who are most advantaged in terms of income, education, and job opportunities, and 1 represents those who are least advantaged\.
- •Feature 4: Education level is a self\-assessed feature measured on a scale from 1 to 8\. A value of 1 indicates “Don’t know / not applicable,” while higher values correspond to higher education levels, ranging from \(2\) no formal qualifications, \(3\) secondary education, \(4\) high school diploma, \(5\) technical or community college, \(6\) undergraduate degree, \(7\) graduate degree and \(8\) doctorate degree\.
- •Feature 6: This feature quantifies the person’s confidence in AI performance before the person engages in the task\. It takes values in the range\[−1,1\]\[\-1,1\], where 1 indicates that AI is expected to perform better and−1\-1that a human would perform better\.
- •Feature 7: This question measures person familiarity with AI\. It takes values in the range\[0,1\]\[0,1\], where 0 indicates never and 1 indicates very frequent use of AI\.
The input features for the human calibration model are shown in Table[9](https://arxiv.org/html/2607.03025#A1.T9)\.
Table 9\.Input features for human calibration model
### A\.3\.Training the human behavior models
#### A\.3\.1\.Human acceptance model
Exploiting the features specified in our dataset, we applied Min–Max normalization to the “socioeconomic status” and “education” features and z\-score normalization for the “age’ feature\. The ground truth for the human acceptance model is as described in A\.1\.1\.
It must be noted that during the training of the human acceptance model, neither the human calibration confidencegtg\_\{t\}nor the agreement assessmentAggr^t\\widehat\{Aggr\}\_\{t\}are available\. Instead, we use \(i\) the absolute value of theadviceadvicefeature from the original dataset in place ofgtg\_\{t\}, and \(ii\) the derived agreement feature, as defined in A\.1, in place ofAggr^t\\widehat\{Aggr\}\_\{t\}\. Although theAggr^t\\widehat\{Aggr\}\_\{t\}and the agreement feature are not identical, both quantify how closely the AI recommendation is with what the human knows, believes and/or prefers\.
#### A\.3\.2\.Human calibration model
The human behavior model is trained independently and integrated into the overall process when it is well\-trained\. As a result, the correctness assessmentCorr^t\\widehat\{Corr\}\_\{t\}required for the training of the human calibration model is not available during training\. Therefore, we train the human calibration model using the correctness of the recommendation derived directly from the data set\. Specifically, if theadviceadvicefeature in the original dataset has a positive sign, it is considered correct \(1\), whereas a negative sign indicates incorrect recommendation \(\-1\)\.
## Appendix BTermination of the reflective process
#### B\.0\.1\.Theorem\.
The iterative decision\-making process with the stated termination condition terminates\.
#### B\.0\.2\.Proof\.
Letk=min\{hacc\(st\+1\)hacc\(st\),t=0,1,2,…\}k=min\\\{\\frac\{h\_\{acc\}\(s\_\{t\}\+1\)\}\{h\_\{acc\}\(s\_\{t\}\)\},t=0,1,2,\\dots\\\}\.
‖Lt\+1−Lt‖≤ϵhacc\(st\+1\)∏j=1t\(1−hacc\(sj\)\)−ϵhacc\(st\)∏j=1t−1\(1−hacc\(sj\)\)\\displaystyle\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\\leq\\frac\{\\epsilon\}\{h\_\{acc\}\(s\_\{t\+1\}\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}\-\\frac\{\\epsilon\}\{h\_\{acc\}\(s\_\{t\}\)\\prod^\{t\-1\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}
Given thathacc\(st\+1\)hacc\(st\)≥k\\frac\{h\_\{acc\}\(s\_\{t\+1\}\)\}\{h\_\{acc\}\(s\_\{t\}\)\}\\geq k, then
\(2\)‖Lt\+1−Lt‖≤ϵ−kϵ\(1−hacc\(st\)\)khacc\(st\)∏j=1t\(1−hacc\(sj\)\)=ϵ\(1−k\)khacc\(st\)\)∏tj=1\(1−hacc\(sj\)\)\+ϵ∏j=1t\(1−hacc\(sj\)\)\\displaystyle\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\\leq\\frac\{\\epsilon\-k\\epsilon\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}\{kh\_\{acc\}\(s\_\{t\}\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}=\\frac\{\\epsilon\(1\-k\)\}\{kh\_\{acc\}\(s\_\{t\}\)\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}\+\\frac\{\\epsilon\}\{\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}
\(a\) Whenk≥1k\\geq 1then\(1−k\)≤0\(1\-k\)\\leq 0, so \(2\) implies that
‖Lt\+1−Lt‖≤ϵ∏j=1t\(1−hacc\(sj\)\)\\displaystyle\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\\leq\\frac\{\\epsilon\}\{\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}Therefore,‖Lt\+1−Lt‖‖Lt−Lt−1‖≤1\(1−hacc\(st\)\)\\frac\{\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\}\{\\\|L\_\{t\}\-L\_\{t\-1\}\\\|\}\\leq\\frac\{1\}\{\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}, i\.e\. the difference in loss in two subsequent iterations increases by a factor greater or equal to 1\. Therefore, the process will reach a pointTTwhere the terminating condition is satisfied\.
\(b\) Whenk<1k<1then\(1−k\)<1\(1\-k\)<1, so we distinguish two cases:
\(b\.1\)ϵ\(1−k\)k\>1\\frac\{\\epsilon\(1\-k\)\}\{k\}\>1, ifϵ\>k\(1−k\)\\epsilon\>\\frac\{k\}\{\(1\-k\)\}\.
This implies that the first term of the right part of the inequality \(1\), which is greater than the second term, increases fast asttincreases, getting much bigger than 1\. In this case
‖Lt\+1−Lt‖‖Lt−Lt−1‖≤hacc\(st−1\)hacc\(st\)\(1−hacc\(st\)\)≤1k\(1−hacc\(st\)\)\\displaystyle\\frac\{\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\}\{\\\|L\_\{t\}\-L\_\{t\-1\}\\\|\}\\leq\\frac\{h\_\{acc\}\(s\_\{t\-1\}\)\}\{h\_\{acc\}\(s\_\{t\}\)\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}\\leq\\frac\{1\}\{k\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}
\(b\.2\)ϵ\(1−k\)k≤1\\frac\{\\epsilon\(1\-k\)\}\{k\}\\leq 1, ifϵ≤k\(1−k\)\\epsilon\\leq\\frac\{k\}\{\(1\-k\)\}\.
In this case \(1\) implies
‖Lt\+1−Lt‖≤1hacc\(st\)\)∏tj=1\(1−hacc\(sj\)\)\+k\(1−k\)∏j=1t\(1−hacc\(sj\)\)=\\displaystyle\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\\leq\\frac\{1\}\{h\_\{acc\}\(s\_\{t\}\)\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}\+\\frac\{k\}\{\(1\-k\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}=1−k\+khacc\(st\)\(1−k\)hacc\(st\)\)∏tj=1\(1−hacc\(sj\)\)≤1\(1−k\)hacc\(st\)\)∏tj=1\(1−hacc\(sj\)\)\\displaystyle\\frac\{1\-k\+kh\_\{acc\}\(s\_\{t\}\)\}\{\(1\-k\)h\_\{acc\}\(s\_\{t\}\)\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}\\leq\\frac\{1\}\{\(1\-k\)h\_\{acc\}\(s\_\{t\}\)\)\\prod^\{t\}\_\{j=1\}\(1\-h\_\{acc\}\(s\_\{j\}\)\)\}
Thus, in a similar way to \(b\.1\) it holds that
‖Lt\+1−Lt‖‖Lt−Lt−1‖≤hacc\(st−1\)hacc\(st\)\(1−hacc\(st\)\)≤1k\(1−hacc\(st\)\)\\displaystyle\\frac\{\\\|L\_\{t\+1\}\-L\_\{t\}\\\|\}\{\\\|L\_\{t\}\-L\_\{t\-1\}\\\|\}\\leq\\frac\{h\_\{acc\}\(s\_\{t\-1\}\)\}\{h\_\{acc\}\(s\_\{t\}\)\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}\\leq\\frac\{1\}\{k\(1\-h\_\{acc\}\(s\_\{t\}\)\)\}
Therefore, in both sub\-cases of \(b\) the upper bound of the difference in loss in two subsequent loss function updates increases by a factor greater or equal to 1\. Therefore, the process will reach a pointTTwhere the terminating condition is satisfied\.
## Appendix CExperimental Setting
### C\.1\.Questions
The following list provides the complete set of questions clustered and ordered by increased difficulty\. Clusters of questions of the same difficulty are within subsequent table rules\. For each question we specify the available options for the specification of constraints\. The last cluster of 10 questions are those used for evaluating the learning abilities of HCRA, exploiting gathered experience\.
### C\.2\.Configuration of components
This section contains details about the hyperparameters and configurations of all components in HCRA\.
#### C\.2\.1\.Termination hyperparameter\.
Results reported are from HCRA with the parameterϵ\\epsilonset to0\.010\.01, affecting the termination of the iterative reflective process\.
Higher values ofϵ\\epsilonare appropriate for decision\-making tasks where different issues have to be reconciled and potential complexities in constraints and recommendations should be resolved, thus for tasks requiring a large number of iterations\. Small values ofϵ\\epsilonare appropriate for decision\-making in question\-answering settings requiring few iterations\. Based on our experimental results,ϵ=0\.01\\epsilon=0\.01allows for sufficient response refinement, achieving high rates of successful decisions, i\.e\., of responses meeting all three criteria: acceptance probability exceeding 0\.5, assessed correctness, and assessed agreement with respect to person’s needs \(request\) and constraints\.
We also limit the number of iterations up to a maximum number providing a safety net for edge cases requiring an excessive number of iterations\. This is not a hyperparameter of the method, but we specify this here for reasons of comprehensiveness\. We have set a maximum of 10 iterations which has been reached only in very few cases, and exclusively during ablation study conditions, indicating that the proposed HCRA rarely requires this high number of iterations: Experimental results show that most successful terminations occur within 3\-7 iterations\. This also depends on the proper set of theϵ\\epsilonvalue\.
#### C\.2\.2\.Language model configurations\.
All HCRA components requiring a language model are instantiated by DeepSeek\-V3\-0324\. This LLM has different configurations, as shown in Table[10](https://arxiv.org/html/2607.03025#A3.T10), depending on the role it plays\.
Table 10\.LLMs HyperparametersHigher actor temperatures \(1\.5, 1\.9\) encourage more exploratory behavior, enabling the actor to generate diverse responses across iterations rather than repeatedly producing similar outputs\. However, they slightly increase the average number of required iterations, while decreasing the success rate, particularly for temperature 1\.9 \(29\.1%\)\.
Setting the evaluator temperature to 0\.0 ensures a consistent evaluation of correctness and agreement in identical or closely similar inputs\. This behaviour is crucial for the HCRA effectiveness, as higher evaluator temperatures would undermine the reflective process, thus the learning process, and termination due to non\-determinism\. This has been observed in the controlled experiment with an unreliable evaluator, as shown in the main paper\.
The self\-reflection temperature of 1\.0 corresponds to the default temperature setting of the underlying language model, allowing the reflective agent to generate varied reflective texts while maintaining coherence and relevance\.
Maximumm Tokens \(MaxTkns\) limits are configured to balance response quality with computational efficiency\. Setting the actor MaxTkns=150 encourages concise recommendations while allowing sufficient detail for comprehensive responses\. The MaxTkns=150 evaluator enables support for both correctness analysis and agreement across multiple constraint dimensions\. Finally, the self\-reflection model MaxTkns=300 enables comprehensive feedback that addresses multiple aspects of the actor performance and human reaction including correctness, confidence, and the acceptance prediction\.
Top logprobs \(TopLPs\) set to 1 for the actor enables confidence scoring for the single generated token\. The evaluator and self\-reflection models do not require confidence scores, therefore logprobs are omitted for these models\.
#### C\.2\.3\.Human Behavior Model Parameters\.
The human acceptance model employs a 3\-layer fully connected neural network architecture \(10×24, 24×12, 12×1\) with ReLU activation functions\. The human calibration model optimizes parametersα,β∈ℝ≥0\\alpha,\\beta\\in\\mathbb\{R\}\_\{\\geq 0\}for the confidence transformation functiongg, as described in the main part of the article\.
### C\.3\.Detailed experimental results
#### C\.3\.1\.Experimental setup\.
Experiments are done in 10 independent runs, each one answering 32 questions, totaling 320 trials\. Each trial concerns the decision making process for one question in one run\. For each run, the order of the questions is randomized to prevent any systematic ordering effect from influencing the results\. The sample size of 320 trials provides sufficient evidence for detecting performance differences in different experimental settings\. The 10\-run design accounts for randomness not only in question ordering, but also in sampling demographic information and in setting constraints\.
Demographic sampling occurs once per run, applying consistent person characteristics across all 32 trials within that run\. The demographic features, detailed in Table[8](https://arxiv.org/html/2607.03025#A1.T8), include age, gender, socioeconomic status, education level, programming experience, and AI preference scores\. This simulates realistic person interactions, where an individual person maintains consistent characteristics while having different constraints in various requests\.
Constraints’ are set independently among trials across seven preference criteria \(food, transportation, accommodation, activities, shopping, nightlife, budget\)\. These are set by a predefined pool of valid constraints per question in our dataset\. Rather than randomly sampling from all possible constraints values, we curate realistic constraints combinations for each question to ensure meaningful evaluation scenarios\. For example, for a question asking ”Give me a high\-end Michelin restaurant in the X area,” having a person constraint of ”low budget” would create a probably inherently contradictory scenario where potentially no response could satisfy both the question’s implicit high\-end requirement and the person’s budget constraint\. This would prevent the model from ever reaching overall agreement, regardless of response quality\. This design ensures realistic scenarios where constraints align with requirements of questions, enabling meaningful evaluation of the HCRA ability to balance correctness with agreement in practical decision\-making contexts\.
Although the main experimental results are provided in the main part of paper, here we provide detailed results as graphs for different actor temperatures: for temperature=0\.1 in Figures[13](https://arxiv.org/html/2607.03025#A3.F13),[14](https://arxiv.org/html/2607.03025#A3.F14),[15](https://arxiv.org/html/2607.03025#A3.F15), for temperature=0\.5 in Figures[16](https://arxiv.org/html/2607.03025#A3.F16),[17](https://arxiv.org/html/2607.03025#A3.F17),[18](https://arxiv.org/html/2607.03025#A3.F18), for temperature=1\.0 in Figures[19](https://arxiv.org/html/2607.03025#A3.F19),[20](https://arxiv.org/html/2607.03025#A3.F20),[21](https://arxiv.org/html/2607.03025#A3.F21), for temperature=1\.5 in Figures[22](https://arxiv.org/html/2607.03025#A3.F22),[23](https://arxiv.org/html/2607.03025#A3.F23),[24](https://arxiv.org/html/2607.03025#A3.F24), and for temperature=1\.9 in Figures[25](https://arxiv.org/html/2607.03025#A3.F25),[26](https://arxiv.org/html/2607.03025#A3.F26),[27](https://arxiv.org/html/2607.03025#A3.F27)
#### C\.3\.2\.Ablation study\.
To evaluate the contribution of individual components in HCRA, we devised three configurations alongside the full proposed architecture\. Experimental settings in the ablation study maintain the experimental setup and components’ configurations used in the experiments with HCRA, with actor temperature =1\.0\.
The HCRA configurations for the ablation study are the following ones:
No human calibration: This configuration \(denoted “No\-g”\) bypasses the human calibration model by directly using the raw actor confidence,cftcf\_\{t\}at each iterationtt\. Table[3](https://arxiv.org/html/2607.03025#S6.T3)in the main paper provide scores for this setting, while detailed results are shown in Figures[28](https://arxiv.org/html/2607.03025#A3.F28),[29](https://arxiv.org/html/2607.03025#A3.F29),[30](https://arxiv.org/html/2607.03025#A3.F30)\.
No Acceptance Model: This configuration replaces the acceptance model’s output with a fixed probability of 0\.5\. However, in this configuration the human calibration model does not play any functional role\. This results into the following baseline configuration\.
Baseline: This configuration removes the human behavior model entirely, using only factual correctness as the termination criterion, as done by Reflexion\(Shinnet al\.,[2023](https://arxiv.org/html/2607.03025#bib.bib26)\)\. The system stops the iterative process upon achieving evaluator correctness, regardless of person constraints or acceptance probability\. Although this configuration does not account neither for human acceptance of the recommendation, neither of evaluation/satisfaction of human constraints, we present it here for comprehensiveness of the ablation study, to investigate the number of iterations required until forming a correct response\. The success of the question\-answering task cannot be evaluated given the lack of human acceptance probability of recommendations\. Figure[8](https://arxiv.org/html/2607.03025#S6.F8)in the main paper provides the results for this setting,
Unreliable evaluator: The evaluator makes unreliable assessments for the correctness and agreement of actor recommendations, as described in Section[6\.2](https://arxiv.org/html/2607.03025#S6.SS2)\. Table[4](https://arxiv.org/html/2607.03025#S6.T4)provides the aggregated scores and detailed results are shown in Figures[31](https://arxiv.org/html/2607.03025#A3.F31),[32](https://arxiv.org/html/2607.03025#A3.F32),[33](https://arxiv.org/html/2607.03025#A3.F33)\.
Figure 13\.Distribution of total vs successful tasks \(actor temperature=0\.1\)Figure 14\.Iterations required per question \(actor temperature= 0\.1\)Figure 15\.Iterations required for successful termination of questions \(actor temperature=0\.1\)Figure 16\.Distribution of total vs successful tasks \(actor temperature=0\.5\)Figure 17\.Iterations required per question \(actor temperature= 0\.5\)Figure 18\.Iterations required for successful termination of questions \(actor temperature=0\.5\)Figure 19\.Distribution of total vs successful tasks \(actor temperature=1\)Figure 20\.Iterations required per question \(actor temperature= 1\)Figure 21\.Iterations required for successful termination of questions \(actor temperature=1\)Figure 22\.Distribution of total vs successful tasks \(actor temperature=1\.5\)Figure 23\.Iterations required per question \(actor temperature=1\.5\)Figure 24\.Iterations required for successful termination of questions \(actor temperature= 1\.5\)Figure 25\.Distribution of total vs successful tasks \(actor temperature=1\.9\)Figure 26\.Iterations required per question \(actor temperature= 1\.9\)Figure 27\.Iterations required for successful termination of questions \(actor temperature=1\.9\)Figure 28\.No g: Distribution of total vs successful tasks \(actor temperature=1\.0\)Figure 29\.No g: Iterations required per question \(actor temperature=1\.0\)Figure 30\.No g: Iterations required for successful termination of questions \(actor temperature= 1\.0\)Figure 31\.Unreliable evaluator: Distribution of total vs successful tasks \(actor temperature=1\.0\)Figure 32\.Unreliable evaluator: Iterations required per question \(actor temperature=1\.0\)Figure 33\.Unreliable evaluator: Iterations required for successful termination of questions \(actor temperature= 1\.0\)
## Appendix DPrompt templates
Figure 34\.Prompt templates for the actor, evaluation and self\-reflection models\.
## Appendix EIndicative examples of the reflective process
Figure 35\.The reflective process for Question 24\.Figure 36\.The reflective process for Question 16\.Similar Articles
Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era
This paper proposes TEAM-Design, a budgeted rule for allocating replay tasks to evaluate human-AI workflow effectiveness compared to human-only or agent-only alternatives, with applications in clinical and coding settings.
From Consumption to Reflection: Designing Human-AI Relations for Stable Reasoning
This paper introduces Relational Reflective Intelligence (RRI), an inference-time governance layer that uses auditable reasoning loops to stabilize human-AI reasoning, addressing cognitive vulnerabilities shared by humans and LLMs.
BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs
Introduces BoardroomAI, a framework for human-steerable multi-agent deliberation using dependency-aware evolving decision graphs, enabling selective repair of affected artifacts after human interventions. Evaluated on synthetic interventions, showing efficiency while preserving unaffected nodes.
Learning to Decide with AI Assistance under Human-Alignment
This paper studies the problem of learning to make optimal decisions with AI assistance under human-alignment, showing that alignment can reduce the complexity of learning, and provides regret bounds.
Human-AI Agent Interaction as a Neuroplastic Training Environment
This paper proposes that the iterative loop of human-AI agent interaction (request, response, appraisal, revision) is a high-frequency neuroplastic training environment that can reinforce negative psychological patterns through repetition, and suggests it can be leveraged for beneficial cognitive training.