TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback
Summary
The paper introduces TPvG, a moral decision framework for LLMs that incorporates consequence feedback, revealing that sequential decisions with feedback alter LLM moral choices, diverging from human patterns.
View Cached Full Text
Cached at: 09/01/26, 12:32 PM
# TPvG: A Moral Decision Framework for Large Language Models from One-Shot to Sequential Feedback
Source: [https://arxiv.org/html/2608.28610](https://arxiv.org/html/2608.28610)
Fangyuan Zhang1Dong Yu1Pengyuan Liu1,2 1Beijing Language and Culture University 2Peking University 202421198123@stu\.blcu\.edu\.cn yudong@blcu\.edu\.cnliupengyuan@pku\.edu\.cn
###### Abstract
Existing LLM moral evaluations typically present models with isolated moral vignettes and elicit a single\-shot decision, neglecting a factor known to profoundly influence human moral behavior: consequence feedback\. We introduce TPvG \(Text\-based Pain\-versus\-Gain\), adapted from a human moral paradigm, which embeds consequence feedback into an everyday moral dilemma of not harming others versus maximising self\-gain\. TPvG comprises five moral decision tasks, progressing from minimal\-context one\-shot choices to sequential decisions with explicit consequence feedback\. Our results show that LLM moral decisions were strongly affected by decision format \(one\-shot versus sequential\), and explicit receiver feedback produced heterogeneous effects across models\. Furthermore, LLM responses to explicit receiver feedback diverged from the human reference pattern, suggesting potentially different decision processes\. These findings highlight the need to evaluate whether LLM moral behavior remains stable in high\-stakes interactive settings\.
Figure 1:Overview of TPvG construction\. \(a\) Original human PvG paradigm\. \(b\) Text\-based dataset construction\. \(c\) Shared system prompt\. The bottom panel shows the five TPvG task formats from one\-shot choices to sequential decisions with increasing contextual and feedback information\.## 1Introduction and Background
Large language models \(LLMs\) are now widely used across many different contexts\(Bommasaniet al\.,[2021](https://arxiv.org/html/2608.28610#bib.bib72); Fraiwan and Khasawneh,[2023](https://arxiv.org/html/2608.28610#bib.bib1)\)\. People increasingly rely on them to make or advise on moral decisions, making it important to evaluate the quality of their moral decisions\(Krügelet al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib2)\)\. As LLMs move into more interactive settings, moral decisions may unfold over repeated actions with accumulating feedback\. This raises a basic question: do LLMs make the same moral decisions in one\-shot vignettes and in sequential tasks with consequence feedback?
Most evaluations of LLM moral judgment rely on moral vignettes administered in a one\-shot format\. One line of work focuses on norm and ethics evaluation, asking models to judge whether actions are ethical, acceptable, or consistent with social norms\(Hendryckset al\.,[2021](https://arxiv.org/html/2608.28610#bib.bib52); Jianget al\.,[2021](https://arxiv.org/html/2608.28610#bib.bib29); Forbeset al\.,[2020](https://arxiv.org/html/2608.28610#bib.bib48); Emelinet al\.,[2021](https://arxiv.org/html/2608.28610#bib.bib49); Ziemset al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib50)\); another examines moral conflict and trade\-offs, probing encoded moral beliefs, exception\-making capacity, autonomous\-vehicle dilemmas, utilitarian reasoning, and multi\-dimensional ethical judgment\(Scherreret al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib30); Jinet al\.,[2022](https://arxiv.org/html/2608.28610#bib.bib31); Takemoto,[2024](https://arxiv.org/html/2608.28610#bib.bib7); Jiaoet al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib6)\)\. Despite differences in content and framing, these evaluations share a common structural feature: the model encounters each scenario once and produces a single response, with no mechanism for the outcome of that response to inform subsequent decisions\. This one\-shot structure supports scalable cross\-model comparison, but leaves unanswered whether model choices remain stable when feedback from prior moral decisions is returned to the model and incorporated into subsequent choices\.
Human behavioral evidence suggests that consequence structure matters: hypothetical moral choices diverge from real\-consequence choices once harm outcomes become concrete, and moral behavior shifts across sequential decisions as a function of prior choices and harm feedback\(Bostynet al\.,[2018](https://arxiv.org/html/2608.28610#bib.bib5); Bostyn and Roets,[2022](https://arxiv.org/html/2608.28610#bib.bib4); Frechen and others,[2022](https://arxiv.org/html/2608.28610#bib.bib81)\)\. For example, in the Pain\-versus\-Gain \(PvG\) paradigm, participants kept significantly less money when feedback about another person’s pain was real rather than hypothetical\(FeldmanHallet al\.,[2012](https://arxiv.org/html/2608.28610#bib.bib15)\)\.
Recent work has placed LLMs in sequential environments, including text games, card\-game tasks, multi\-agent social dilemmas, and repeated games, showing that interaction history and feedback can shape model behavior\(Qinet al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib53); Piattiet al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib54); Akataet al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib55); Tennantet al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib56)\)\. Other studies introduce multi\-step moral dilemmas or prompt models to reason about downstream consequences\(Wuet al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib58); Selet al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib57)\)\. However, these settings primarily target strategic cooperation, reward pursuit, sequential competence, or imagined consequences rather than moral consistency under harm\-relevant feedback\. A more fundamental obstacle further compounds this gap: moral psychology has long relied on hypothetical vignettes and thought experiments in part because ethical concerns render naturalistic study of real\-consequence moral behavior largely impractical\(O’Connoret al\.,[2022](https://arxiv.org/html/2608.28610#bib.bib3)\), leaving human baseline data from consequential moral paradigms scarce\. Without such baselines, it remains difficult to assess whether LLM moral outputs reflect processes analogous to those underlying human moral decision\-making\. The Pain\-versus\-Gain \(PvG\) paradigm\(FeldmanHallet al\.,[2012](https://arxiv.org/html/2608.28610#bib.bib15)\)is among the rare exceptions—conducted under ethically sanctioned laboratory conditions, it provides human moral choice data under real\-consequence feedback alongside a grounded harm\-relevant dilemma structure, making it a natural candidate for adaptation into an LLM evaluation framework\.
We introduceTPvG, a text\-based PvG benchmark for testing LLM moral decisions beyond one\-shot vignettes\. We ask whether decision format and consequence feedback alter model choices, and whether these patterns align with human PvG data\. These results suggest that LLM moral evaluation should not be limited to one\-shot scenarios, but should include sequential, consequence\-sensitive settings that test whether moral behavior remains stable as decisions unfold over time\.
Table 1:Money Keptacross the five TPvG task formats\.Δ\(T\-S\)\\Delta\(T\\text\{\-\}S\)andΔ\(T\-E\)\\Delta\(T\\text\{\-\}E\)denote Trial\-by\-Trial minus Scenario and Trial\-by\-Trial minus Enriched, respectively\. Shaded cells indicate significant model\-level contrasts after Benjamini–Hochberg correction\.
## 2Methodology
We constructed TPvG by adapting stimuli from the original PvG study\(FeldmanHallet al\.,[2012](https://arxiv.org/html/2608.28610#bib.bib15)\)into five text\-based task scenarios \(Figure[1](https://arxiv.org/html/2608.28610#S0.F1)\)\. The benchmark includes two response\-question formats and an 11\-level receiver\-feedback scale for Real TPvG\.
STEP 1: Annotation\.Two psychology experts annotated the original PvG study to distinguish directly reusable information from construction\-required information\. Directly reusable information included the experimental background, participant roles, consent procedure, spatial arrangement, task procedure, initial endowment, and the qualitative inverse relation between spending and shock intensity\.
STEP 2: Textual Operationalization\.The two psychology experts then operationalized construction\-required information under predefined rewriting guidelines\. Objective descriptions were limited to observable room, apparatus, and setting details; subjective descriptions were grounded in the original low\-intensity sample stimulation; and interactional details were restricted to neutral exchanges unrelated to the task\. For sequential TPvG levels, we also annotated the spending–shock mapping to help models interpret each decision consequence\.
STEP 3: Feedback Construction and Validation\.For Real TPvG, receiver feedback was constructed as ordinal textual descriptions of visible hand movement for shock levels 0–10, based on the original PvG procedure and apparatus information and Digitimer DS7A specifications\(Digitimer Ltd\.,[2019](https://arxiv.org/html/2608.28610#bib.bib86)\)\. GPT\-4o was used only for auxiliary ordinal validation, yielding a mean Spearman correlation ofρ=\.994\\rho=\.994across 30 shuffled rating orders\.
STEP 4: Question Instantiation\.We instantiated response questions for the five TPvG levels\. Scenario and Enriched Scenario TPvG used “What would you choose to do?”, whereas Trial\-by\-Trial, Near\-Real, and Real TPvG used “How much do you choose to spend this trial?”
Together, these steps generated five controlled TPvG tasks varying contextual richness, decision format, participant interaction, baseline shock experience, physical\-money framing, and explicit receiver feedback\.
## 3Experimental Setup
### Models\.
We evaluate 11 contemporary LLMs spanning open\-weight and proprietary model families, including Llama\(Grattafioriet al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib59)\), Qwen/Qwen2\.5\(Baiet al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib62); Yanget al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib63)\), Mistral\(Jianget al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib61)\), Gemma\(Gemma Team,[2024](https://arxiv.org/html/2608.28610#bib.bib66)\), GPT\-4o\(OpenAI,[2024](https://arxiv.org/html/2608.28610#bib.bib67)\), DeepSeek\-V3\(DeepSeek\-AI,[2024](https://arxiv.org/html/2608.28610#bib.bib64)\), and Centaur\(Binzet al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib60)\)\.
### Metrics\.
We useMoney Keptas the primary outcome and mean adjacent absolute change \(MAC\) to measure decision variability in sequential tasks\. LowerMoney Keptindicates greater harm prevention, whereas higher values indicate greater self\-gain\.
## 4Results and Analysis
### Do Response Formats Affect LLM Moral Decisions?
Answer: Yes\.Table[1](https://arxiv.org/html/2608.28610#S1.T1)summarizesMoney Keptacross the five TPvG conditions\. Models retained little money in the two one\-shot conditions but substantially more in the three sequential conditions\. Aggregating within response format confirmed this shift: sequential conditions produced higherMoney Keptthan one\-shot conditions \(ΔMoney Kept=£6\.03\\Delta\\textit\{Money Kept\}=\\pounds 6\.03,pBH<\.01p\_\{\\mathrm\{BH\}\}<\.01\)\.
Trial\-by\-Trial TPvG also increasedMoney Keptrelative to both one\-shot baselines\. Compared with Scenario TPvG, 9 of 11 models retained more money; compared with Enriched Scenario TPvG, 8 of 11 did so\. Both contrasts were significant after correction \(Δ\(T\-S\)=£6\.92\\Delta\(T\\text\{\-\}S\)=\\pounds 6\.92,Δ\(T\-E\)=£6\.73\\Delta\(T\\text\{\-\}E\)=\\pounds 6\.73, bothpBH<\.01p\_\{\\mathrm\{BH\}\}<\.01\)\. This pattern indicates that LLM moral decisions are influenced by response format, motivating evaluation under sequential decision\-making conditions\.
### Do LLM Moral Decisions Change Within the Same Response Format?
Within response format, descriptive enrichment alone did not changeMoney Kept\(Enriched–Scenario:Δ=£0\.18\\Delta=\\pounds 0\.18,pBH=1\.00p\_\{\\mathrm\{BH\}\}=1\.00\), and Near\-Real TPvG did not differ from Trial\-by\-Trial TPvG \(Δ=£0\.04\\Delta=\\pounds 0\.04,pBH=1\.00p\_\{\\mathrm\{BH\}\}=1\.00\)\. By contrast, Real TPvG reducedMoney Keptrelative to Near\-Real TPvG \(Δ=£−2\.46\\Delta=\\pounds\{\-2\.46\},pBH=\.033p\_\{\\mathrm\{BH\}\}=\.033\) and Trial\-by\-Trial TPvG \(Δ=£−2\.42\\Delta=\\pounds\{\-2\.42\},pperm=\.016p\_\{\\mathrm\{perm\}\}=\.016\)\. At the individual\-model level, the Real–Near\-Real decrease was significant for Qwen2\.5\-14B, GPT\-4o, Gemma\-2\-9B, and DeepSeek\-V3\.
### Are Feedback Effects Consistent Across Models?
Feedback effects were heterogeneous across models\. Among the four models with significant Real–Near\-Real decreases inMoney Kept, Qwen2\.5\-14B and GPT\-4o showed increased decision variability in Real TPvG \(ΔMAC=\+\.026\\Delta\\mathrm\{MAC\}=\+\.026and\+\.052\+\.052\), Gemma\-2\-9B showed a descriptive decrease \(ΔMAC=−\.053\\Delta\\mathrm\{MAC\}=\-\.053\), and DeepSeek\-V3 changed little \(ΔMAC=\+\.014\\Delta\\mathrm\{MAC\}=\+\.014\)\. Among models without significantMoney Keptreductions, Llama3\-8B nevertheless showed increased MAC, whereas most others showed no reliable change or remained at floor\. Thus, feedback reshaped both outcomes and trajectories unevenly across models\. Thus, consequence feedback affected not only final decision outcomes, as indexed byMoney Kept, but also the stability of trial\-by\-trial decision trajectories, with heterogeneous patterns across models\.
Figure 2:Mean per\-trial Money Kept trajectories under Near\-Real and Real TPvG for the four models with significant Real–Near\-Real reductions in Money Kept\.
### How Do LLM TPvG Results Compare With Human PvG Patterns?
Figure 3:Descriptive comparison between the human PvG reference pattern and the LLM TPvG results\. Human values are taken from the original PvG study and are used only as a descriptive reference\.LLM results matched the human reference pattern in the broad one\-shot versus sequential contrast: both showed higher money retention under sequential settings\. However, the sequential\-condition profiles diverged\. As shown in Figure[3](https://arxiv.org/html/2608.28610#S4.F3), humans retained the least money in Trial\-by\-Trial, with retention increasing across more concrete and consequential conditions\. LLMs showed the opposite pattern, retaining the least in Real TPvG, where receiver feedback was explicitly represented in text\.
This divergence suggests that humans and LLMs may rely on different mechanisms during repeated moral decision\-making\. Human behavior may reflect increasing salience of the duty not to harm, whereas LLM outputs appear sensitive to harm\-related surface cues in the prompt—consistent with prior work on LLM moral and safety behavior\(Scherreret al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib30); Krügelet al\.,[2023](https://arxiv.org/html/2608.28610#bib.bib2); Cheunget al\.,[2025](https://arxiv.org/html/2608.28610#bib.bib18); Röttgeret al\.,[2024](https://arxiv.org/html/2608.28610#bib.bib83)\)\. These findings highlight the importance of examining LLM decisions across sequential settings rather than relying solely on one\-shot evaluations\.
## 5Conclusion
We introduced TPvG, a text\-based PvG framework for evaluating LLM moral decision\-making in harm\-versus\-self\-gain dilemmas\. Across 11 LLMs,Money Keptvaried sharply by decision format, and responses to explicit receiver feedback were heterogeneous\. These findings suggest that one\-shot moral evaluations should be complemented by sequential, consequence\-relevant settings that test whether LLM moral behavior remains stable as decisions unfold\.
## Limitations
Our study focuses on a controlled, text\-based adaptation of one PvG paradigm\. This design allows us to isolate decision format and feedback effects, but it does not capture real monetary stakes, embodied pain, or social responsibility\. The human comparison is also descriptive, as the reference values come from the original PvG study rather than a matched text\-only experiment\. Finally, although we evaluate 11 LLMs and include robustness checks, the results may vary with future models, prompts, and deployment settings\. These limitations point to a natural extension of the framework to broader moral domains and matched human–model studies\.
## Ethics Statement
This work studies morally sensitive scenarios involving self\-benefit, harm reduction, and painful outcomes\. All experiments were text\-based simulations; no participant or model faced real shocks, monetary loss, or actual harm\. We do not interpret model outputs as evidence of moral agency, subjective concern, or emotional experience\. The purpose of the evaluation is to diagnose how LLM outputs change under controlled textual task formats, not to certify models as morally competent decision\-makers\.
## References
- E\. Akata, L\. Schulz, J\. Coda\-Forno, S\. J\. Oh, M\. Bethge, and E\. Schulz \(2025\)Playing repeated games with large language models\.Nature Human Behaviour9,pp\. 1380–1390\.External Links:[Document](https://dx.doi.org/10.1038/s41562-025-02172-y),[Link](https://doi.org/10.1038/s41562-025-02172-y)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- J\. Bai, S\. Bai, Y\. Chu, Z\. Cui, K\. Dang, X\. Deng, Y\. Fan, W\. Ge, Y\. Han, F\. Huang,et al\.\(2023\)Qwen technical report\.External Links:2309\.16609,[Document](https://dx.doi.org/10.48550/arXiv.2309.16609),[Link](https://arxiv.org/abs/2309.16609)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- M\. Binz, E\. Akata, M\. Bethge, F\. Brändle, F\. Callaway, J\. Coda\-Forno, P\. Dayan, C\. Demircan, M\. K\. Eckstein,et al\.\(2025\)A foundation model to predict and capture human cognition\.Nature644,pp\. 1002–1009\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09215-4),[Link](https://doi.org/10.1038/s41586-025-09215-4)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- R\. Bommasani, D\. A\. Hudson, E\. Adeli, R\. Altman, S\. Arora, S\. von Arx, M\. S\. Bernstein, J\. Bohg, A\. Bosselut, E\. Brunskill,et al\.\(2021\)On the opportunities and risks of foundation models\.External Links:2108\.07258,[Document](https://dx.doi.org/10.48550/arXiv.2108.07258),[Link](https://arxiv.org/abs/2108.07258)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p1.1)\.
- D\. H\. Bostyn and A\. Roets \(2022\)Sequential decision\-making impacts moral judgment: how iterative dilemmas can expand our perspective on sacrificial harm\.Journal of Experimental Social Psychology98,pp\. 104244\.External Links:[Document](https://dx.doi.org/10.1016/j.jesp.2021.104244),[Link](https://doi.org/10.1016/j.jesp.2021.104244)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p3.1)\.
- D\. H\. Bostyn, S\. Sevenhant, and A\. Roets \(2018\)Of mice, men, and trolleys: hypothetical judgment versus real\-life behavior in trolley\-style moral dilemmas\.Psychological Science29\(7\),pp\. 1084–1093\.External Links:[Document](https://dx.doi.org/10.1177/0956797617752640),[Link](https://doi.org/10.1177/0956797617752640)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p3.1)\.
- V\. Cheung, M\. Maier, and F\. Lieder \(2025\)Large language models show amplified cognitive biases in moral decision\-making\.Proceedings of the National Academy of Sciences122\(25\),pp\. e2412015122\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2412015122),[Link](https://doi.org/10.1073/pnas.2412015122)Cited by:[§4](https://arxiv.org/html/2608.28610#S4.SS0.SSS0.Px4.p2.1)\.
- DeepSeek\-AI \(2024\)DeepSeek\-V3 technical report\.External Links:2412\.19437,[Document](https://dx.doi.org/10.48550/arXiv.2412.19437),[Link](https://arxiv.org/abs/2412.19437)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- Digitimer Ltd\. \(2019\)DS7A and DS7AH high voltage constant current stimulators\.Note:Product documentationAccessed 2026\-07\-14External Links:[Link](https://www.digitimer.com/product/life-science-research/stimulators/ds7a-ds7ah-hv-current-stimulator)Cited by:[§2](https://arxiv.org/html/2608.28610#S2.p4.1)\.
- D\. Emelin, R\. Le Bras, J\. D\. Hwang, M\. Forbes, and Y\. Choi \(2021\)Moral stories: situated reasoning about norms, intents, actions, and their consequences\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 698–718\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.54),[Link](https://aclanthology.org/2021.emnlp-main.54)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- O\. FeldmanHall, D\. Mobbs, D\. Evans, L\. Hiscox, L\. Navrady, and T\. Dalgleish \(2012\)What we say and what we do: the relationship between real and hypothetical moral choices\.Cognition123\(3\),pp\. 434–441\.External Links:[Document](https://dx.doi.org/10.1016/j.cognition.2012.02.001),[Link](https://doi.org/10.1016/j.cognition.2012.02.001)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p3.1),[§1](https://arxiv.org/html/2608.28610#S1.p4.1),[§2](https://arxiv.org/html/2608.28610#S2.p1.1)\.
- M\. Forbes, J\. D\. Hwang, V\. Shwartz, M\. Sap, and Y\. Choi \(2020\)Social chemistry 101: learning to reason about social and moral norms\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing,pp\. 653–670\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.48),[Link](https://aclanthology.org/2020.emnlp-main.48)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- M\. Fraiwan and N\. Khasawneh \(2023\)A review of ChatGPT applications in education, marketing, software engineering, and healthcare: benefits, drawbacks, and research directions\.External Links:2305\.00237,[Document](https://dx.doi.org/10.48550/arXiv.2305.00237),[Link](https://arxiv.org/abs/2305.00237)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p1.1)\.
- S\. Frechenet al\.\(2022\)Wait, did I do that? effects of previous decisions on moral decision\-making\.Journal of Behavioral Decision Making\.External Links:[Document](https://dx.doi.org/10.1002/bdm.2279),[Link](https://doi.org/10.1002/bdm.2279)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p3.1)\.
- Gemma Team \(2024\)Gemma 2: improving open language models at a practical size\.External Links:2408\.00118,[Document](https://dx.doi.org/10.48550/arXiv.2408.00118),[Link](https://arxiv.org/abs/2408.00118)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.\(2024\)The Llama 3 herd of models\.External Links:2407\.21783,[Document](https://dx.doi.org/10.48550/arXiv.2407.21783),[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Critch, J\. Li, D\. Song, and J\. Steinhardt \(2021\)Aligning AI with shared human values\.InInternational Conference on Learning Representations,External Links:2008\.02275,[Document](https://dx.doi.org/10.48550/arXiv.2008.02275),[Link](https://arxiv.org/abs/2008.02275)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier,et al\.\(2023\)Mistral 7b\.External Links:2310\.06825,[Document](https://dx.doi.org/10.48550/arXiv.2310.06825),[Link](https://arxiv.org/abs/2310.06825)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- L\. Jiang, J\. D\. Hwang, C\. Bhagavatula, R\. Le Bras, J\. Liang, J\. Dodge, K\. Sakaguchi, M\. Forbes, J\. Borchardt, S\. Gabriel, Y\. Tsvetkov, O\. Etzioni, M\. Sap, R\. Rini, and Y\. Choi \(2021\)Can machines learn morality? the Delphi experiment\.External Links:2110\.07574,[Document](https://dx.doi.org/10.48550/arXiv.2110.07574),[Link](https://arxiv.org/abs/2110.07574)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- J\. Jiao, S\. Afroogh, A\. Murali, K\. Chen, D\. Atkinson, and A\. Dhurandhar \(2025\)LLM ethics benchmark: a three\-dimensional assessment system for evaluating moral reasoning in large language models\.Scientific Reports15,pp\. 34642\.External Links:[Document](https://dx.doi.org/10.1038/s41598-025-18489-7),[Link](https://doi.org/10.1038/s41598-025-18489-7)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- Z\. Jin, S\. Levine, F\. Gonzalez Adauto, O\. Kamal, M\. Sap, M\. Sachan, R\. Mihalcea, J\. Tenenbaum, and B\. Schölkopf \(2022\)When to make exceptions: exploring language models as accounts of human moral judgment\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 28458–28473\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/b654d6150630a5ba5df7a55621390daf-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- S\. Krügel, A\. Ostermaier, and M\. Uhl \(2023\)ChatGPT’s inconsistent moral advice influences users’ judgment\.Scientific Reports13,pp\. 4569\.External Links:[Document](https://dx.doi.org/10.1038/s41598-023-31341-0),[Link](https://doi.org/10.1038/s41598-023-31341-0)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p1.1),[§4](https://arxiv.org/html/2608.28610#S4.SS0.SSS0.Px4.p2.1)\.
- B\. B\. O’Connor, K\. Lee, D\. Campbell, and L\. Young \(2022\)Moral psychology from the lab to the wild: relief registries as a paradigm for studying real\-world altruism\.PLOS ONE17\(6\),pp\. e0269469\.External Links:[Document](https://dx.doi.org/10.1371/journal.pone.0269469),[Link](https://doi.org/10.1371/journal.pone.0269469)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- OpenAI \(2024\)GPT\-4o system card\.Note:OpenAI system cardExternal Links:[Link](https://openai.com/index/gpt-4o-system-card/)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- G\. Piatti, Z\. Jin, M\. Kleiman\-Weiner, B\. Schölkopf, M\. Sachan, and R\. Mihalcea \(2024\)Cooperate or collapse: emergence of sustainable cooperation in a society of LLM agents\.InThe Thirty\-Eighth Annual Conference on Neural Information Processing Systems,External Links:2404\.16698,[Document](https://dx.doi.org/10.48550/arXiv.2404.16698),[Link](https://arxiv.org/abs/2404.16698)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- Z\. Qin, H\. Wang, D\. Liu, Z\. Song, C\. Fan, Z\. Lv, J\. Wu, Z\. Lei, Z\. Tu, D\. Chu, X\. Yu, and D\. Sui \(2024\)UNO arena for evaluating sequential decision\-making capability of large language models\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 7630–7645\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.435),[Link](https://aclanthology.org/2024.emnlp-main.435)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- P\. Röttger, H\. Kirk, B\. Vidgen, G\. Attanasio, F\. Bianchi, and D\. Hovy \(2024\)XSTest: a test suite for identifying exaggerated safety behaviours in large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Mexico City, Mexico,pp\. 5377–5400\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.301),[Link](https://aclanthology.org/2024.naacl-long.301)Cited by:[§4](https://arxiv.org/html/2608.28610#S4.SS0.SSS0.Px4.p2.1)\.
- N\. Scherrer, C\. Shi, A\. Feder, and D\. M\. Blei \(2023\)Evaluating the moral beliefs encoded in LLMs\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 51778–51809\.External Links:[Link](https://papers.nips.cc/paper_files/paper/2023/hash/a2cf225ba392627529efef14dc857e22-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1),[§4](https://arxiv.org/html/2608.28610#S4.SS0.SSS0.Px4.p2.1)\.
- B\. Sel, P\. Shanmugasundaram, M\. Kachuee, K\. Zhou, R\. Jia, and M\. Jin \(2024\)Skin\-in\-the\-game: decision making via multi\-stakeholder alignment in LLMs\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Note:ACL 2024 long paperExternal Links:2405\.12933,[Document](https://dx.doi.org/10.48550/arXiv.2405.12933),[Link](https://arxiv.org/abs/2405.12933)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- K\. Takemoto \(2024\)The moral machine experiment on large language models\.Royal Society Open Science11\(2\),pp\. 231393\.External Links:[Document](https://dx.doi.org/10.1098/rsos.231393),[Link](https://doi.org/10.1098/rsos.231393)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.
- E\. Tennant, S\. Hailes, and M\. Musolesi \(2025\)Moral alignment for LLM agents\.InThe Thirteenth International Conference on Learning Representations,External Links:2410\.01639,[Document](https://dx.doi.org/10.48550/arXiv.2410.01639),[Link](https://openreview.net/forum?id=MeGDmZjUXy)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- Y\. Wu, Q\. Sheng, D\. Wang, G\. Yang, Y\. Sun, Z\. Wang, Y\. Bu, and J\. Cao \(2025\)The staircase of ethics: probing LLM value priorities through multi\-step induction to complex moral dilemmas\.External Links:2505\.18154,[Document](https://dx.doi.org/10.48550/arXiv.2505.18154),[Link](https://arxiv.org/abs/2505.18154)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p4.1)\.
- A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei,et al\.\(2024\)Qwen2\.5 technical report\.External Links:2412\.15115,[Document](https://dx.doi.org/10.48550/arXiv.2412.15115),[Link](https://arxiv.org/abs/2412.15115)Cited by:[§3](https://arxiv.org/html/2608.28610#S3.SS0.SSS0.Px1.p1.1)\.
- C\. Ziems, J\. Yu, Y\. Wang, A\. Halevy, and D\. Yang \(2023\)NormBank: a knowledge bank of situational social norms\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7756–7776\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.429),[Link](https://aclanthology.org/2023.acl-long.429)Cited by:[§1](https://arxiv.org/html/2608.28610#S1.p2.1)\.Similar Articles
Efficient LLM Reasoning via Variational Posterior Guidance with Efficiency Awareness
This paper introduces the VPG-EA framework, which uses variational inference and posterior guidance to improve the reasoning efficiency of large language models by addressing the 'overthinking' phenomenon in chain-of-thought generation.
Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas
The paper introduces VirtueMap, a framework that profiles large language models by evaluating their rankings of ethical dilemma responses through an Aristotelian virtue ethics lens, using a validated common-sense ground truth.
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
A comprehensive dual-aspect evaluation framework for large language models on Vietnamese legal text simplification, combining quantitative benchmarking (Accuracy, Readability, Consistency) with qualitative error analysis across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro, and Grok-1.
Evaluating Large Language Models in a Complex Hidden Role Game
This paper introduces an open-source framework to evaluate LLMs' reasoning, persuasion, and deception capabilities in the hidden role game Secret Hitler, finding that current models fail at sustained multi-turn manipulation while rule-based agents outperform them.
Sequential statistical inference for Large Language Models: Representation, validity, and monitoring
This paper argues for a sequential inference framework to enhance LLM trustworthiness by modeling interactions as dependent stochastic processes, ensuring validity under repeated use, and enabling online monitoring for behavioral shifts.