Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study
Summary
This paper explores whether valence features can reflect morality in natural language by analyzing human annotations of moral scenarios, finding significant correlations and achieving a Matthew's correlation coefficient of 0.764 for binary morality classification.
View Cached Full Text
Cached at: 07/24/26, 05:17 AM
# Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study
Source: [https://arxiv.org/html/2607.20461](https://arxiv.org/html/2607.20461)
Malika Bendechache, Louise McCormack, Elif Calik, Ramin Ranjbarzadeh, Dost Muhammad, Shokofeh Anari BozcheloeiIshita Singh
###### Abstract
Present implementations of artificial intelligence \(AI\) ethics do not adequately take feelings, or affect, into account\. If AI should be aligned with human ethics, it seems reasonable to thoroughly investigate the possibility of AI behaviour that mirrors virtuous human ethical conduct, where feelings play a role in the actions, judgements or statements one makes\. Furthermore, while prominent theories of normative ethics are often discussed in terms of their differences and shortcomings, Virtue, Consequentialist, and Kantian Deontological ethics all share a common feature of considering human feeling to some degree while the popular descriptive ethics theory, Moral Foundations Theory, positions feelings as central to many of its foundations\. Therefore, in the present paper, a data set of moral valence is proposed, consisting of 500 annotations by six human participants for both action/judgement and consequence moral valence, ranging from \-1 to 1 for text\-presented scenarios from the Commonsense Norm Bank data set\. The resulting valence features share significant relationships with multi\-class \(immoral/discretionary/moral\) and binary immoral/moral categories while additionally providing a noteworthy test set Matthew’s correlation coefficient of 0\.764 using regularised logistic regression for binary classification\. This provides early evidence of the usefulness of valence features for morality estimation of text, indicating that valenced consequences of responses for others can be considered toward more human morally\-aligned AI\. In the interest of promoting further affective\-moral computing research, this study’s annotations will be made available for research on request\.
## IIntroduction
Artificial intelligence \(AI\), can serve humanity in unprecedented ways where menial, routine, dangerous, or some types of knowledge work processes can be automated by AIs with varying levels of agency\. Assuming that the negative implications of deploying AI can be understood and improved upon or avoided, one could imagine a utopian future for society\. However, while tools leveraging these technologies have potential to offer individuals and societies great benefits, they come with risks, potentially affecting hundreds of millions of users\. Inadequate algorithmic accountability\[[30](https://arxiv.org/html/2607.20461#bib.bib80)\], responsibility gaps\[[29](https://arxiv.org/html/2607.20461#bib.bib23)\], and moral crumple zones\[[10](https://arxiv.org/html/2607.20461#bib.bib28)\], blaming the nearest human for complex system failures, are some potential issues\. There are further issues for, largely speaking, presently available amoral AI\[[22](https://arxiv.org/html/2607.20461#bib.bib5),[25](https://arxiv.org/html/2607.20461#bib.bib3),[5](https://arxiv.org/html/2607.20461#bib.bib47)\]and even a truly ethical AI, if possible, should still be subjected to careful governance\. However, work toward such an ethical AI is warranted to promote safe and equitable outcomes for those using AI or affected by its use\.
The idea that feelings play a role in moral conduct has been understood for some time\. Aristotle\[[1](https://arxiv.org/html/2607.20461#bib.bib39)\]wrote of emotional states, and pleasure and pain in the perceptions of virtuous persons\. The classical Utilitarianism of Bentham and Mill perhaps matches most closely with affect, where this theory proposes that a moral action is the one that produces the most pleasure or reduction in pain for oneself or others\[[32](https://arxiv.org/html/2607.20461#bib.bib41)\]\. Kant’s Deontology\[[21](https://arxiv.org/html/2607.20461#bib.bib40)\]describes feelings of respect for the moral law, and respect for oneself and love of others among some of the feelings that are “necessary conditions of rational moral agency”\[[14](https://arxiv.org/html/2607.20461#bib.bib43), p\. 170\]\. Modern philosophers agree the that feelings play a role in morality, either as a basis for moral consideration\[[35](https://arxiv.org/html/2607.20461#bib.bib45)\]or due to feelings associated with morality or lack thereof\[[33](https://arxiv.org/html/2607.20461#bib.bib46)\]\. However, present implementations of AI ethics have shown little to no moral affect/feeling incorporation\[[37](https://arxiv.org/html/2607.20461#bib.bib30)\]\. Motivated by this research opportunity, the following research question is addressed in this paper:
Do subjective valence ratings of moral judgements/actions and their consequences offer indicative features of the morality of text\-presented scenarios?
A review of prior art was conducted to inform this question \(Section[II](https://arxiv.org/html/2607.20461#S2)\)\. Following confirmation of a research opportunity for moral valence corpora development, moral action/judgement and consequence valence data were generated by human annotation of moral scenarios presented in text\. The source data used, and the annotation process is described in Section[III](https://arxiv.org/html/2607.20461#S3)\. This set offers the first continuous\-valued, multi\-temporal perception ratings for the valence, ranging from \-1, maximally unpleasant, through 0 to \+1 or maximally pleasant for moral situations\. This derivative data set from the Commonsense Norm Bank corpus\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\], hereafter Norm Bank, provides fine\-grained affective information for morally\-relevant text stimuli\. A goal of this work was to show the predictive usefulness of valence toward morality classification in text, where prudence dictates that immorality recognition should be an initial focus\. Thus, experiments for the exploration of the generated labels and analysis of their predictive usefulness for binary immoral/moral classification are described in Section[IV](https://arxiv.org/html/2607.20461#S4), with results in Section[V](https://arxiv.org/html/2607.20461#S5)\. This is followed by a discussion of the results \(Section[VI](https://arxiv.org/html/2607.20461#S6)\) and concluding remarks in Section[VII](https://arxiv.org/html/2607.20461#S7)\. The annotation data will be made available to researchers on request to the lead author\.
## IIRelated Work
Researchers have generated automatic morality recognition systems using bottom\-up, learning from examples\[[19](https://arxiv.org/html/2607.20461#bib.bib49),[4](https://arxiv.org/html/2607.20461#bib.bib48),[20](https://arxiv.org/html/2607.20461#bib.bib17),[38](https://arxiv.org/html/2607.20461#bib.bib50)\], top\-down, rule\-based\[[37](https://arxiv.org/html/2607.20461#bib.bib30)\], and hybrid approaches, combining both of the aforementioned methods\[[37](https://arxiv.org/html/2607.20461#bib.bib30),[20](https://arxiv.org/html/2607.20461#bib.bib17)\]\. The philosopher John Rawls\[[33](https://arxiv.org/html/2607.20461#bib.bib46), pp\. 40\-46\]proposed a hybrid approach to justify moral principles \(top\-down rules\) and one’s considered \(bottom\-up\) judgements of moral situations by adjustment of both to bring them into agreement, a process he called reflective equilibrium\. The present paper focuses on one side of reflective equilibrium, namely feature representation of moral content presented in text, toward more effective characterisation of bottom\-up, descriptive ethics\.
Annotations for moral content of text data commonly use Moral Foundations Theory \(MFT\) as a descriptive theory\[[40](https://arxiv.org/html/2607.20461#bib.bib54)\]\. Initiated by Haidt and Craig\[[17](https://arxiv.org/html/2607.20461#bib.bib51)\]for characterisation of the innate/intuitive ethics of human beings, the moral foundations are based on the presence of one or more of the virtues, care, fairness, loyalty, authority, and/or purity \(fairness was later broken intoequalityandproportionalitymoral foundations\[[2](https://arxiv.org/html/2607.20461#bib.bib53)\]\)\. Labelling and automatic recognition of vices is also pursued, for example the presence/absence of harm and/or care\[[19](https://arxiv.org/html/2607.20461#bib.bib49)\], while an additional non\-moral class can be further added\[[19](https://arxiv.org/html/2607.20461#bib.bib49),[38](https://arxiv.org/html/2607.20461#bib.bib50)\]\. Problems for the practical use of MFT in AI/ML research settings is the need to train annotators in the theory, while some foundations, notably loyalty and purity, have low base rates in online text\[[19](https://arxiv.org/html/2607.20461#bib.bib49),[38](https://arxiv.org/html/2607.20461#bib.bib50)\]\. Binary categorisation is another option for labelling the moral content of text\[[18](https://arxiv.org/html/2607.20461#bib.bib73),[20](https://arxiv.org/html/2607.20461#bib.bib17)\]\. This format is attractive due to its simplicity, and the fact that a potentially ambiguous separation of neutral and positive moral classes can logically be collapsed into a positive/moral class\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\]\. A problem with binary categorisation is a clear lack of fine\-grained information describing the moral phenomenon\. For example, both mass murder and reckless speeding are both immoral, however, one of these scenarios appears much more immoral than the other\.
Moral emotion has been characterised by orthogonal valence \(ranging from harm to help\) and agency \(ranging from agent to patient\) dimensions by Gray and Wegner\[[15](https://arxiv.org/html/2607.20461#bib.bib57)\]\. This characterisation of morally praiseworthy \(blameworthy\) actions could be advantageous to determine the moral \(immoral\) weight of an action/judgement\. Psychologists have measured the valence of moral stimuli numerically\[[23](https://arxiv.org/html/2607.20461#bib.bib55)\], and combined it with MFT\[[9](https://arxiv.org/html/2607.20461#bib.bib56)\], however, moral valence remains underexplored in AI ethics\. Valence brings forth external evaluation\[[3](https://arxiv.org/html/2607.20461#bib.bib58)\], facilitating internal positive/negative feelings based on, but not exclusive to, external world stimuli\. Valence measurement has the practical advantage of little to no training requirement; individuals must subjectively rate stimuli as negative, neutral, or positive, and if numerical ratings are to be obtained, specify the degree of valence, perhaps ranging from \-1 to 1, for example\.
Motivated by this research need, in this work continuous\-valued valence measurements were generated toward improved descriptive characterisation of the moral content of text\. Additionally, due to the truism that humans generally see themselves as existing over time, both action/judgement valence and consequence valence were captured in this work\. This serves as the main contribution of this work, as not only is moral valence underexplored in text, but, to the knowledge of the authors, this is the first annotation data set considering both the valence intrinsic to a moral action/judgement in addition to the consequences associated with that action/judgement\. The intention of the development of these descriptive features of morality is to augment present research, whether MFT, binary immoral/moral, or other morality description, classification, or prediction tasks\.
## IIIData Acquisition and Metadata
Prior to data acquisition the proposed data collection protocol was ethically approved by the Research Ethics Committee of the lead author’s institution\. A convenience sample of annotators were then contacted by email with project information to seek their voluntary participation in the study with six annotators consenting to participation\. All annotators were told that they would be provided with opportunities to collaborate on research paper writing for their participation in the study\.
### III\-ASource Corpus: Norm Bank
Norm Bank\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\]111Data source: https://github\.com/liweijiang/delphi, licensed under CC BY\-NC\-SA 4\.0222Licence: https://creativecommons\.org/licenses/by\-nc\-sa/4\.0/legalcode, was used as the source corpus for valence annotation in this work\. This data set consists of 1\.7M text scenarios unified from SocialChem\[[12](https://arxiv.org/html/2607.20461#bib.bib67)\], ETHICS\[[18](https://arxiv.org/html/2607.20461#bib.bib73)\], Moral Stories\[[11](https://arxiv.org/html/2607.20461#bib.bib68)\], Social Bias Frames\[[34](https://arxiv.org/html/2607.20461#bib.bib69)\]and SCRUPLES\[[27](https://arxiv.org/html/2607.20461#bib.bib70)\]sub\-corpora, with ratings for moral content in text provided as either freeform \(moral, discretionary, or immoral category rating\), or yes/no \(QA\-format target\), for example “yes, it is kind”\. Only freeform categorical rated scenarios from SocialChem, ETHICS, and Moral Stories sub\-corpora were sampled from Norm Bank in this work, hence, further details are provided on these data sets below\.
SocialChem\[[12](https://arxiv.org/html/2607.20461#bib.bib67)\]includes 971,620 instances, and is a large scale corpus of people’s ethical judgements and social norms on a wide range of everyday situations based on text from subreddits, the ROCStories corpus, and the Dear Abby advice column\. ETHICS\[[18](https://arxiv.org/html/2607.20461#bib.bib73)\]is composed of 20,948 instances and involves situations and judgements of everyday events and interpersonal relationships depicted in text for Justice, Deontology, Virtue, Utilitarianism, and Commonsense Ethics scenarios\. Moral Stories\[[11](https://arxiv.org/html/2607.20461#bib.bib68)\], a data set of 144,000 instances, is a text corpus of norms, intents, actions, and their consequences, with both normed/moral and divergent/immoral contextual scenarios present\.
### III\-BAnnotation Generation
Annotator subjects were presented with a paragraph describing valence and instructed to provide valence ratings in the range\[−1,1\]\[\-1,1\]for both an action/judgement, hereafter, action, and perceived consequences associated with that action presented to them in text\. Subjects were provided with 500 instances randomly sampled from the Norm Bank corpus, along with a simple Rshinyapplication\. The application was designed to run on their own laptop/PC and present moral situations one at a time for subjective valence rating generation\. As part of the annotation procedure/presentation, subjects were blinded from moral category label information of the presented scenarios\. An example annotation window for a text\-display moral situation is shown in Figure[1](https://arxiv.org/html/2607.20461#S3.F1)where sliders for the two continuous valence ratings that they were required to provide are shown\. Three places of decimal precision was used for the ratings, meaning that a total of2,0012,001unique valence values could be selected for each valence dimension\. Subjects were told that they could take multiple weeks to provide their ratings\. The total annotation time was estimated to be three hours per subject\.
Figure 1:Example application window where a participating subject could view a moral scenario presented in text and provide their ratings for both action and consequence valence in the range\[−1,1\]\[\-1,1\]using the appropriate slider\. After a selection is made and submitted, rating for the present scenario ends, and subjects are presented with the next one\. Annotation subjects all rated a total of 500 examples from Norm Bank\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\]\.
### III\-CMetadata
Annotations were provided by six subjects from May 2025 to February 2026\. All subjects were associated with an Irish university during the annotation period, either in a research/teaching or professional support capacity\. The annotation subjects’ regions of origin and biological sexes included the Middle East \(2M, 2F\), South Asia \(1F\), and the United Kingdom and Ireland \(1M\)\.
## IVExperiment Design
All experimental analyses outside of annotation generation were performed on a Dell 16 Plus 2\-in\-1 with Intel Core Ultra 7 8\-core processor\. The R programming language and interpreter \(version 4\.4\.1\) was used for the experiments with the random number generator seed set usingset\.seed\(123\)for random data sampling\.
### IV\-AData Partitioning & Gold Standard Generation
Because the line between discretionary \(neutral\) and positive \(moral\) labels is not as clear as that between either of those classes with negative \(immoral\) ratings\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\], the discretionary and moral classes were collapsed into one discretionary/moral class\. The discretionary/moral class, hereafter moral, was assigned a value of 1 while immoral rated scenarios were assigned the binary value 0\. Following this, data were split 80/20 into training and test partitions using stratified sampling based on the binary immoral/moral class label, the ultimate label for prediction in this work\. This resulted in 400 training examples and 100 test examples\. This was followed by individual annotator statistics generation to determine general characteristics for the obtained valence ratings, and calculation of Lin’s concordance correlation coefficient \(CCC\)\[[26](https://arxiv.org/html/2607.20461#bib.bib62)\]to determine annotator agreement with each other, and with gold standard values generated by simple averaging \(mean\) and weighted averaging based on annotator CCC values, the evaluator weighted estimator \(EWE\) gold standard\. The CCC, Equation[1](https://arxiv.org/html/2607.20461#S4.E1), provides information on both precision and accuracy in one metric, hence facilitating estimation of annotator agreement and therefore, gold standard label quality\.
CCC=2σabσa2\+σb2\+\(μa−μb\)2,\\text\{CCC\}=\\frac\{2\\sigma\_\{ab\}\}\{\\sigma^\{2\}\_\{a\}\+\\sigma^\{2\}\_\{b\}\+\(\\mu\_\{a\}\-\\mu\_\{b\}\)^\{2\}\}\\;,\(1\)whereσ\\sigmadenotes covariance,σ2\\sigma^\{2\}is the uncorrected sample variance, andμ\\muis the mean\.
As is frequent practice in affective computing\[[36](https://arxiv.org/html/2607.20461#bib.bib60)\], the generated raw labels can be weighted based on their correlation with a mean gold standard toward increased label quality\[[16](https://arxiv.org/html/2607.20461#bib.bib61)\]\. This weighting technique, the EWE, facilitates a downweighting of low confidence \(or overly noisy\) annotators for gold standard generation as follows
x^nEWE=1∑k=1Krk∑k=1Krkx^n,k,\\hat\{x\}\_\{n\}^\{\\text\{EWE\}\}=\\frac\{1\}\{\\sum\_\{k=1\}^\{K\}r\_\{k\}\}\\sum\_\{k=1\}^\{K\}\{r\_\{k\}\\hat\{x\}\_\{n,k\}\}\\;,\(2\)wherex^n,k\\hat\{x\}\_\{n,k\}is an annotation rating by annotation subjectkkfor instance examplenn\. For the present work, the CCC was chosen as the correlation weighting coefficientrrfollowing beneficial results obtained by Stappen et al\.\[[36](https://arxiv.org/html/2607.20461#bib.bib60)\]\. The EWE weighting coefficients were learned and applied on the training data for each annotator and were compared against a simple mean weighting for valence ratings\. The better of the two approaches in terms of average annotator agreement \(largerμCCC\\mu\\text\{CCC\}\) with respective generated gold standard ratings was further used in the later experimental steps\. EWE, if providing better performance, having learned coefficients applicationonlyon the test data to prevent data leakage\.
### IV\-BTraining Set Exploratory Analysis
Exploratory graphs were generated for the gold standard ratings for action and consequence valence ratings\. In addition, continuous with multi\-class, and continuous with binary, valence\-morality associations were assessed using analysis of variance \(ANOVA\) and Pearson’s correlation, respectively\. The intention of this was to explore the possibility of significant associations between the Norm Bank\-provided morality ratings and the generated valence ratings of the present work\.
### IV\-CBinary Immoral/Moral Classification using Action and Consequence Valence
Following the exploratory analysis, a majority class baseline classifier was first generated\. Then, a L2 \(ridge\) regularised logistic regression model was learned on the training data usingglmnet\[[13](https://arxiv.org/html/2607.20461#bib.bib66)\]andcaret\[[24](https://arxiv.org/html/2607.20461#bib.bib65)\]R software packages\. As is standard withglmnet, action and consequence annotations, now serving as input features, were centred to a mean of zero with unit variance prior to model training\. Candidateλ\\lambdaregularisation parameters were generated using five\-fold cross\-validation on the training set\. This was followed by stratified five\-fold cross\-validation and finalλ\\lambdaselection based on Matthew’s correlation coefficient \(MCC\) maximisation from candidate values in this validation setting\. Final model training on the whole training set for the selected lambda was then performed\. Results for the final model are provided for both cross\-validation and tests sets, using accuracy as an intuitive metric, and MCC which is more robust to class imbalance\[[8](https://arxiv.org/html/2607.20461#bib.bib63)\]\. Proposed by Matthew\[[28](https://arxiv.org/html/2607.20461#bib.bib64)\], the MCC can be written as
MCC=\(TP\)\(TN\)−\(FP\)\(FN\)\(TP\+FP\)\(TP\+FN\)\(TN\+FP\)\(TN\+FN\),\\begin\{array\}\[\]\{@\{\}l\}\\text\{MCC\}=\\\\ \\displaystyle\\frac\{\(TP\)\(TN\)\-\(FP\)\(FN\)\}\{\\sqrt\{\(TP\+FP\)\(TP\+FN\)\(TN\+FP\)\(TN\+FN\)\}\}\\\>,\\end\{array\}\(3\)whereTP,TN,FP,FNTP,TN,FP,FNdenote true positive, true negative, false positive, and false negative counts from the confusion matrix, respectively\. The MCC takes on values in\[−1,1\]\[\-1,1\], where a larger positive value is better andMCC=0\\text\{MCC\}=0, for example, indicating a classifier no better than a random guess\. Error analysis of false positives and false negatives was conducted on the final model’s test set performance as well\.
## VExperimental Results
### V\-AIndividual Annotations and Gold Standard Annotations Generation on the Training Set
Statistics for generated annotations from subjects S1\-S6 on the training set are provided in Table[I](https://arxiv.org/html/2607.20461#S5.T1)\. The subjects often have mean ratings close to zero/exactly neutral, while it can be seen that subjects S1 and S2 were less dispersed as measured by smaller standard deviations \[0\.2, 0\.4\]\. In general, subjects were quite dispersed in terms of standard deviations \[0\.6, 0\.9\], considering the scale of the data \[\-1, 1\]\. This demonstrates one of the difficulties of subjectively rating such data, where individuals bring different sets of biases and judgements, potentially based on their own values and generated understanding of the world\. Another interesting observation one can make from the generated valence labels is that a small amount of unique values were selected from the2,0012,001potential unique values available\. Exemplars being S4, who only selected 1% of unique values available for consequence valence, and S1 who, selecting the largest number of unique values, only selected 15% of available values for consequence valence\. Perhaps with a larger number of scenarios, user may select more diverse subjective ratings covering a larger space of moral affect\. A final observation, not shown in Table[I](https://arxiv.org/html/2607.20461#S5.T1)is that raters selected values close to extreme \-1 and 1 poles\. This tells us that while in many cases there were few unique values were selected, overall, the general annotation space \(negative, neutral, positive\) was used and particularly emotional/affectively strong moral situations are present in the data\.
TABLE I:Individual Annotation Statistics for Action and Consequence Valence: Mean \(Standard Deviation\), andUnique ValuesAverage and standard deviation values for the generated valence annotationsμCCC\\mu\\text\{CCC\}are provided in Table[II](https://arxiv.org/html/2607.20461#S5.T2)\. Average pairwise annotator agreement is low \(action valence CCC = 0\.260, consequence valence CCC = 0\.356\), suggesting discordance among the annotators\. This is to be expected to some degree due to personal values shaping the provided ratings in addition to the fact that some moral situations may lack objective truth\. Further, even if a clear objective truth is available, annotators still have to provide theirfeelingratings, for the given situation’s action and consequence valence\.μCCC\\mu\\text\{CCC\}values were always higher for consequence compared with action valence, suggesting that labelling this phenomenon can yield higher quality valence characterisation of moral situations\. The comparison of mean and EWE valence ratings with the original annotator labels suggest that the EWE weighting scheme is the best approach for aggregation of the individual annotator’s ratings for both action and consequence valence\. Table[II](https://arxiv.org/html/2607.20461#S5.T2)shows absolute improvements of 0\.023 and 0\.004μCCC\\mu\\text\{CCC\}for action and consequence valence, respectively\. Annotator standard deviations are larger for the EWE weighted gold standard valence annotations due to low confidence annotator downweighting and hence more dispersion when comparing the annotators to EWE gold standard\.
TABLE II:Average \(Standard Deviation\) Annotator CCC for Unqiue Annotator\-Annotator Pairs \(μCCCpairwise\\mu\{\\text\{CCC\}\}\_\{pairwise\}\), All Annotators with the Mean Gold Standard \(μCCCmean\\mu\{\\text\{CCC\}\}\_\{mean\}\), and All Annotators with the EWE\-weighted Gold Standard \(μCCCEWE\\mu\{\\text\{CCC\}\}\_\{EWE\}\)
### V\-BTraining Set Exploratory Analysis
Based on the previous experimental result, the EWE gold standard was used for valence annotations henceforth\. These valence gold standard annotations are shown in Figure[2](https://arxiv.org/html/2607.20461#S5.F2), where action and consequence valence appear highly correlated with each other\. The Pearson’s correlation coefficient \(PCC\) calculated for the presented action and consequence valence data was 0\.976, suggesting strong association, or redundancy, between these measures of moral valence\. In terms of class\-wise distributions of action and consequence valence, clusters appear to be present in the data, with discretionary and moral classes appearing in general to be represented by positive valence values while the immoral class is generally represented by negative valence ratings\. The training set multi\-class moral label distribution included discretionary 43\.25% of the time, immoral for 33\.50%, and moral for 23\.25% of sample morality labels\.
Figure 2:Training data set action and consequence valence ratings \(Pearson’s correlation coefficient = 0\.976 \[t=88\.995,p<0\.001t=88\.995,p<0\.001\]\) for moral scenarios rated as discretionary \(light greyblue, label proprtion = 43\.25%\), immoral \(red, 33\.50%\), and moral \(green, 23\.25%\)\.Levene’s test indicated homogeneity of variance for both action \(p=0\.995p=0\.995\) and consequence valence \(p=0\.459p=0\.459\) across class groups\. Following this, one\-way ANOVA was conducted along with residual plot checks and Shapiro\-Wilk statistic rejection of Gaussian distributed residuals for action \(p=0\.043p=0\.043\) and consequence \(p=0\.017p=0\.017\) valence models\. The residual plots showed only a small amount of outliers, hence, robustness checks were performed including Kruskal\-Wallis test, Welch’s ANOVA, and permutation ANOVA \(with random permutationsB=10,000B=10,000\)\. Agreement was found in the robustness checks both in terms ofppvalue significance and approximate effect size estimates,η2\\eta^\{2\}\. Due to these results, in addition to the large group sizes present \(immoral = 134, discretionary = 173, & moral = 93 examples\), and indicated homogeneity of variance, the standard one\-way ANOVA results are reported along with 95% confidence intervals for effect sizes\.
This analysis resulted in statistically significant associations between the multicategory morality ratings and action valence \(F=198\.100,p<0\.001,η2=0\.50\[0\.45,1\.00\]F=198\.100,p<0\.001,\\eta^\{2\}=0\.50\\,\[0\.45,1\.00\]\), and consequence valence \(F=218\.100,p<0\.001,η2=0\.52\[0\.47,1\.00\]F=218\.100,p<0\.001,\\eta^\{2\}=0\.52\\,\[0\.47,1\.00\]\)\. The observed large effect sizes forη2\\eta^\{2\}indicate that both action and consequence valence share strong associations with the multicategory morality ratings\. Unfortunately, due to the large CIs observed, the estimation of this effect size is uncertain\. These early exploratory results appear promising, however, and it is expected that the small number of outliers observed in residual QQ\-plots contributed to this estimation uncertainty\.
Following the ANOVA analysis, Tukey’s honest significant difference \(HSD\) test was performed to determine pairwise mean differences between groups, the results of which can be seen in Table[III](https://arxiv.org/html/2607.20461#S5.T3)\. The results of this analysis show significant differences in valence ratings, for both action and consequence valence, when comparing the immoral class ratings to that of both discretionary and moral classes\. Additionally, Table[III](https://arxiv.org/html/2607.20461#S5.T3)shows that discriminating between the moral and discretionary classes based on valence ratings is difficult\. This is evidenced by small mean differences between these groups’ ratings, that in both cases of action and consequence valence did not reach statistical significance at thep=0\.05p=0\.05level\. This corroborates visual clustering\-type behaviour observed in Figure[2](https://arxiv.org/html/2607.20461#S5.F2)for these classes and further confirms the difficulty of discriminating between them, based on moral valence\.
TABLE III:Tukey’s Honest Significant Difference \(HSD\) Test Results for Action Valence \(AV\) and Consequence Valence \(CV\) Differences Between Immoral, Discretionary, & Moral ClassesFinally, both action and consequence valence had noteworthy positive PCC values with the binary immoral/moral class label, with action valence PCC = 0\.703 \(t=−19\.717,p<0\.001t=\-19\.717,p<0\.001\) and consequence valence PCC = 0\.720 \(t=−20\.671,p<0\.001t=\-20\.671,p<0\.001\)\. These results show that both measurements can be strong predictors of binary immoral/moral ratings, with consequence valence being a marginally better predictor\. Further, these results bolster that of Table[II](https://arxiv.org/html/2607.20461#S5.T2)where consequence valence was indicated as a higher quality measure \(higherμCCC\\mu\\text\{CCC\}\) compared to that of action valence\.
### V\-CBinary Immoral/Moral Classification using Action and Consequence Valence
For this task, baseline scores were calculated using the majority class \(moral\) guess classifier on the training set\. This resulted in a zero information rate model accuracy of 66\.50% along with a MCC of 0\. The L2 regularised logistic regressionλ\\lambdavalue learned from cross\-validation was 0\.104 and this was used for learning the final logistic regression classifier weights\. The baseline, cross\-validation, and test set accuracy % and MCC obtained from this experimentation are provided in Table[IV](https://arxiv.org/html/2607.20461#S5.T4)\. It can be seen that very strong prediction performance was obtained for the simple L2 regularised logistic regression model containing only two input features, where the baseline score was beaten by a large margin, both in terms of accuracy %, and more importantly MCC\. Absolute improvements of 0\.695 \(five\-fold cross\-validation\) and 0\.764 \(test set\) were obtained compared with the majority class baseline classifier MCC\. This provides early evidence that valence can provide effective prediction of binary immoral/moral ratings of text\.
TABLE IV:Action and Consequence Valence Binary Immoral/Moral Logistic Regression Classification Results Including Training Majority Class Prediction Only Baseline, Five\-fold Cross\-validation Performance \(Standard Deviation\), and Final Test Set Performance Accuracy % & Matthew’s Correlation Coefficient \(MCC\)The logistic regression model performed very well for both classes based on valence feature input alone, where high true positive \(0\.863\) and true negative \(0\.963\) rates were obtained\. The most problematic error produced on the test set was where the model predicted that an immoral scenario was moral \(×1\\times 1\)\. Predicting immoral text as moral \(false positive\) facilitates text display to users that has more harm potential than blocking users from seeing moral text \(false negative\)\. Hence, this problematic case was evaluated, both for the action and consequence valence ratings serving as input to the model, along with the text stimulus that those ratings were based on\. The valence input for this case was \-0\.626 for action valence and \-0\.566 for consequence valence \(EWE gold standard\) while the text input that these ratings were based on was “telling your friend’s business to someone else when you are worried about them, if sharing private information about a friend with another friend”\. This item’s ground\-truth label was originallydiscretionarybut was later collapsed into the moral class\. This therefore indicates annotator error compared with ground\-truth, perhaps due to it being a discretionary case \(labelled moral for binary prediction\)\. General error analysis was conducted from 10 test set false negative errors\. An interesting result from this analysis was that five out of ten error cases had consequence valence ratings close to zero\. This indicates an opportunity for the development of borderline, hard\-to\-predict examples in the future, so that models are more capable in reducing these observed high prevalence errors\.
## VIDiscussion
The text material presented to annotators from Norm Bank\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\]was often morally, and hence emotionally/affectively, evocative\. It has been shown that emotional experience correlates with moral choice in previous work\[[7](https://arxiv.org/html/2607.20461#bib.bib76)\]\. Therefore, it is not surprising that the valence annotations from this study share significant associations with the Norm Bank moral category labels\. Agreement was additionally found in the results for previously mentioned difficulty in differentiation of discretionary and moral classes\[[20](https://arxiv.org/html/2607.20461#bib.bib17)\]\. Differences in means for both action and consequence valence were very small between moral and discretionary class labels and did not reach statistical significance\. Therefore, valence measurements cannot offer much in terms of distinguishing between these moral classes\. However, both action and consequence valence were indicative of immoral vs the other classes\. This class was distinguishable from either discretionary or moral classes with large group mean differences observed, all reaching statistical significance\. Further, strong positive correlations between action valence and consequence valence with binary immoral/moral labels were obtained\. From the evaluations undertaken, consequence valence appeared to be more certain across annotators while additionally sharing stronger relationships with moral category labels\. This shows that action valence may be a candidate for removal from the proposed moral valence features due to its inferiority when compared with consequence valence\. Further, for moral valence, it appears slightly more favourable to think of valence in terms of the consequences that follow a valent action/judgement, rather than the intrinsic valence of an action/judgement\.
The binary morality classification results obtained on the data using action and consequence valence are strong for the relatively simple modelling approach taken\. Additionally, emerging evidence from affective computing suggests that frontier LLMs have capabilities of emotion prediction\[[39](https://arxiv.org/html/2607.20461#bib.bib77)\]and cognitive emotional self\-appraisal\[[6](https://arxiv.org/html/2607.20461#bib.bib75)\], the simulation of subjective understanding of an LLM’s own simulated emotion\. It seems reasonable then, that despite a lack of awareness of this ability unless explicitly prompted, many LLMs may have a capability for, albeit imperfect, moral estimation of their responses/actions\. The results of the present paper suggest, in particular if morality is of primary concern, that valence inference could be conducted by an AI system on their own generated responses prior to action execution or display to a human\. While valence prediction has traditionally been performed for estimation of a human’s valence \(this will be illegal without consent under the AI Act\), such an affective\-moral computing effort provides another use for this well\-studied dimension of human affect\. Any implementation of such a feature would of course serve as a descriptive moral feature component as part of broader system implementation and should be subject to human oversight\[[31](https://arxiv.org/html/2607.20461#bib.bib82)\]\. Due to the impracticalities of humans monitoring all human\-AI interactions, such a system could offer a second\-best estimate, compared with human judgement, intended to make practically feasible/scalable but not diminish, human compliance workload\.
## VIIConclusion and Future Work
The proposed valence features and experiments presented were intended to discover if subjective valence ratings of moral judgements/actions and their consequences offer indicative features of morality in text\. Based on the experimental results obtained, subjective valence features of text appear to share strong relationships with both multi\-class and binary moral category\-rated text\. Consequence moral valence was shown as the more important of the two proposed valence features\. Further, the prediction results obtained were strong, indicating morality estimation based on moral valence is practicable\. While further experimental evaluation of the proposed affective\-moral features is required, this work provides early evidence of the usefulness of moral valence features toward morality estimation of text\.
Some important study limitations are now worth mentioning\. The number of annotations provided in this work is small while only one descriptive moral phenomenon was studied\. Therefore, a reasonable first step to advance the present work is to extend the size of the labelled set\. To this end, further manual annotation could be performed along with data augmentation and weak supervision\. The generated descriptive moral features of the proposed work, while promising, offer an incomplete view of descriptive ethics, perhaps consequentialist\-leaning in nature\. Hence, they can, and should, be combined with other descriptive measures such as MFT to offer a more comprehensive characterisation of bottom\-up, descriptive ethics\. In addition, for ultimate deployment, a complete implementation of reflective equilibrium with the inclusion of a normative rule\-based system should be pursued while such a system should remain human\-aligned and auditable\. This leads to further future work including alignment evaluation of human\- compared with LLM\-generated moral valence ratings along with multi\-theory descriptive & normative framework implementation and proposals for audit criteria for effective human oversight\. It is hoped that this work advances affective\-moral computing efforts that can inform future AI ethics system development\.
## References
- \[1\]Aristotle, H\. Tredennick, and J\. Barnes\(2004\)The Nicomachean Ethics\.Penguin Classics,London\(English\)\.External Links:ISBN 978\-0\-14\-044949\-5Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1)\.
- \[2\]M\. Atari, J\. Haidt, J\. Graham, S\. Koleva, S\. T\. Stevens, and M\. Dehghani\(2023\)Morality beyond the WEIRD: How the nomological network of morality varies across cultures\.Journal of Personality and Social Psychology125\(5\),pp\. 1157–1188\.Note:Place: US Publisher: American Psychological AssociationExternal Links:ISSN 1939\-1315,[Document](https://dx.doi.org/10.1037/pspp0000470)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p2.1)\.
- \[3\]L\. F\. Barrett\(2006\-02\)Valence is a basic building block of emotional life\.Journal of Research in Personality40\(1\),pp\. 35–55\.External Links:ISSN 0092\-6566,[Link](https://www.sciencedirect.com/science/article/pii/S0092656605000590),[Document](https://dx.doi.org/10.1016/j.jrp.2005.08.006)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p3.1)\.
- \[4\]M\. G\. Beiró, J\. D’Ignazi, V\. Perez Bustos, M\. F\. Prado, and K\. Kalimeri\(2023\-04\)Moral Narratives Around the Vaccination Debate on Facebook\.InProceedings of the ACM Web Conference 2023,WWW ’23,New York, NY, USA,pp\. 4134–4141\.External Links:ISBN 978\-1\-4503\-9416\-1,[Link](https://dl.acm.org/doi/10.1145/3543507.3583865),[Document](https://dx.doi.org/10.1145/3543507.3583865)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p1.1)\.
- \[5\]M\. Bhat and D\. Long\(2025\-10\)Emotional Plausibility vs\. Emotional Truth: Designing Against Affective Misinformation in Conversational AI\.Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society8\(1\),pp\. 430–444\(en\)\.External Links:ISSN 3065\-8365,[Link](https://ojs.aaai.org/index.php/AIES/article/view/36561),[Document](https://dx.doi.org/10.1609/aies.v8i1.36561)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[6\]S\. Bhattacharyya, L\. Craig, T\. Dilliraj, J\. Li, and J\. Z\. Wang\(2025\-08\)Do Machines Think Emotionally? Cognitive Appraisal Analysis of Large Language Models\.arXiv\.Note:arXiv:2508\.05880 \[cs\]External Links:[Link](http://arxiv.org/abs/2508.05880),[Document](https://dx.doi.org/10.48550/arXiv.2508.05880)Cited by:[§VI](https://arxiv.org/html/2607.20461#S6.p2.1)\.
- \[7\]M\. Carmona\-Perera, C\. Martí\-García, M\. Pérez\-García, and A\. Verdejo\-García\(2013\-09\)Valence of emotions and moral decision\-making: increased pleasantness to pleasant images and decreased unpleasantness to unpleasant images are associated with utilitarian choices in healthy adults\.Frontiers in Human Neuroscience7,pp\. 626\.External Links:ISSN 1662\-5161,[Link](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3783947/),[Document](https://dx.doi.org/10.3389/fnhum.2013.00626)Cited by:[§VI](https://arxiv.org/html/2607.20461#S6.p1.1)\.
- \[8\]D\. Chicco and G\. Jurman\(2023\-02\)The Matthews correlation coefficient \(MCC\) should replace the ROC AUC as the standard metric for assessing binary classification\.BioData Mining16,pp\. 4\.External Links:ISSN 1756\-0381,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC9938573/),[Document](https://dx.doi.org/10.1186/s13040-023-00322-4)Cited by:[§IV\-C](https://arxiv.org/html/2607.20461#S4.SS3.p1.2)\.
- \[9\]D\. L\. Crone, S\. Bode, C\. Murawski, and S\. M\. Laham\(2018\-01\)The Socio\-Moral Image Database \(SMID\): A novel stimulus set for the study of social, moral and affective processes\.PLoS ONE13\(1\),pp\. e0190954\.External Links:ISSN 1932\-6203,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC5783374/),[Document](https://dx.doi.org/10.1371/journal.pone.0190954)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p3.1)\.
- \[10\]M\. C\. Elish\(2019\-03\)Moral Crumple Zones: Cautionary Tales in Human\-Robot Interaction\.Engaging Science, Technology, and Society5,pp\. 40–60\(en\)\.External Links:ISSN 2413\-8053,[Link](https://estsjournal.org/index.php/ests/article/view/260),[Document](https://dx.doi.org/10.17351/ests2019.260)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[11\]D\. Emelin, R\. Le Bras, J\. D\. Hwang, M\. Forbes, and Y\. Choi\(2021\-11\)Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 698–718\.External Links:[Link](https://aclanthology.org/2021.emnlp-main.54/),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.54)Cited by:[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1),[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p2.1)\.
- \[12\]M\. Forbes, J\. D\. Hwang, V\. Shwartz, M\. Sap, and Y\. Choi\(2020\-11\)Social Chemistry 101: Learning to Reason about Social and Moral Norms\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 653–670\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.48/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.48)Cited by:[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1),[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p2.1)\.
- \[13\]J\. H\. Friedman, T\. Hastie, and R\. Tibshirani\(2010\-02\)Regularization Paths for Generalized Linear Models via Coordinate Descent\.Journal of Statistical Software33,pp\. 1–22\(en\)\.External Links:ISSN 1548\-7660,[Link](https://doi.org/10.18637/jss.v033.i01),[Document](https://dx.doi.org/10.18637/jss.v033.i01)Cited by:[§IV\-C](https://arxiv.org/html/2607.20461#S4.SS3.p1.2)\.
- \[14\]I\. Geiger\(2011\)Kant on the Affective Moods of Morality\.InPhilosophy’s Moods: The Affective Grounds of Thinking,H\. Kenaan and I\. Ferber \(Eds\.\),pp\. 159–172\(en\)\.External Links:ISBN 978\-94\-007\-1503\-5,[Link](https://doi.org/10.1007/978-94-007-1503-5%5C_11),[Document](https://dx.doi.org/10.1007/978-94-007-1503-5%5F11)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1)\.
- \[15\]K\. Gray and D\. M\. Wegner\(2011\-07\)Dimensions of Moral Emotions\.Emotion Review3\(3\),pp\. 258–260\(EN\)\.Note:Publisher: SAGE PublicationsExternal Links:ISSN 1754\-0739,[Link](https://doi.org/10.1177/1754073911402388),[Document](https://dx.doi.org/10.1177/1754073911402388)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p3.1)\.
- \[16\]M\. Grimm and K\. Kroschel\(2005\)Evaluation of natural emotions using self assessment manikins\.InIEEE Workshop on Automatic Speech Recognition and Understanding, 2005\.,San Juan, Puerto Rico,pp\. 381–385\(en\)\.External Links:ISBN 978\-0\-7803\-9479\-7,[Link](http://ieeexplore.ieee.org/document/1566530/),[Document](https://dx.doi.org/10.1109/ASRU.2005.1566530)Cited by:[§IV\-A](https://arxiv.org/html/2607.20461#S4.SS1.p2.6)\.
- \[17\]J\. Haidt and C\. Joseph\(2004\)Intuitive Ethics: How Innately Prepared Intuitions Generate Culturally Variable Virtues\.Daedalus133\(4\),pp\. 55–66\.Note:Publisher: The MIT PressExternal Links:ISSN 0011\-5266,[Link](https://www.jstor.org/stable/20027945)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p2.1)\.
- \[18\]D\. Hendrycks, C\. Burns, S\. Basart, A\. Critch, J\. Li, D\. Song, and J\. Steinhardt\(2021\)Aligning ai with shared human values\.Proceedings of the International Conference on Learning Representations \(ICLR\)\.Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p2.1),[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1),[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p2.1)\.
- \[19\]J\. Hoover, G\. Portillo\-Wightman, L\. Yeh, S\. Havaldar, A\. M\. Davani, Y\. Lin, B\. Kennedy, M\. Atari, Z\. Kamel, M\. Mendlen, G\. Moreno, C\. Park, T\. E\. Chang, J\. Chin, C\. Leong, J\. Y\. Leung, A\. Mirinjian, and M\. Dehghani\(2020\-11\)Moral Foundations Twitter Corpus: A Collection of 35k Tweets Annotated for Moral Sentiment\.Social Psychological and Personality Science11\(8\),pp\. 1057–1071\(EN\)\.Note:Publisher: SAGE Publications IncExternal Links:ISSN 1948\-5506,[Link](https://doi.org/10.1177/1948550619876629),[Document](https://dx.doi.org/10.1177/1948550619876629)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p1.1),[§II](https://arxiv.org/html/2607.20461#S2.p2.1)\.
- \[20\]L\. Jiang, J\. D\. Hwang, C\. Bhagavatula, R\. L\. Bras, J\. T\. Liang, S\. Levine, J\. Dodge, K\. Sakaguchi, M\. Forbes, J\. Hessel, J\. Borchardt, T\. Sorensen, S\. Gabriel, Y\. Tsvetkov, O\. Etzioni, M\. Sap, R\. Rini, and Y\. Choi\(2025\-01\)Investigating machine moral judgement through the Delphi experiment\.Nature Machine Intelligence7\(1\),pp\. 145–160\(en\)\.Note:Publisher: Nature Publishing GroupExternal Links:ISSN 2522\-5839,[Link](https://www.nature.com/articles/s42256-024-00969-6),[Document](https://dx.doi.org/10.1038/s42256-024-00969-6)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1),[§II](https://arxiv.org/html/2607.20461#S2.p1.1),[§II](https://arxiv.org/html/2607.20461#S2.p2.1),[Figure 1](https://arxiv.org/html/2607.20461#S3.F1),[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1),[§IV\-A](https://arxiv.org/html/2607.20461#S4.SS1.p1.4),[§VI](https://arxiv.org/html/2607.20461#S6.p1.1)\.
- \[21\]I\. Kant\(1797\)The metaphysics of morals\.Cambridge University Press,New York\.Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1)\.
- \[22\]B\. Kim and J\. Lee\(2024\-11\)The mental health implications of artificial intelligence adoption: the crucial role of self\-efficacy\.Humanities and Social Sciences Communications11\(1\),pp\. 1561\(en\)\.Note:Publisher: PalgraveExternal Links:ISSN 2662\-9992,[Link](https://www.nature.com/articles/s41599-024-04018-w),[Document](https://dx.doi.org/10.1057/s41599-024-04018-w)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[23\]K\. M\. Knutson, F\. Krueger, M\. Koenigs, A\. Hawley, J\. R\. Escobedo, V\. Vasudeva, R\. Adolphs, and J\. Grafman\(2010\-12\)Behavioral norms for condensed moral vignettes\.Social Cognitive and Affective Neuroscience5\(4\),pp\. 378–384\.External Links:ISSN 1749\-5016,[Link](https://doi.org/10.1093/scan/nsq005),[Document](https://dx.doi.org/10.1093/scan/nsq005)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p3.1)\.
- \[24\]M\. Kuhn\(2008\-11\)Building Predictive Models in R Using the caret Package\.Journal of Statistical Software28,pp\. 1–26\(en\)\.External Links:ISSN 1548\-7660,[Link](https://doi.org/10.18637/jss.v028.i05),[Document](https://dx.doi.org/10.18637/jss.v028.i05)Cited by:[§IV\-C](https://arxiv.org/html/2607.20461#S4.SS3.p1.2)\.
- \[25\]H\. \(\. Lee, A\. Sarkar, L\. Tankelevitch, I\. Drosos, S\. Rintel, R\. Banks, and N\. Wilson\(2025\-04\)The Impact of Generative AI on Critical Thinking: Self\-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers\.InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems,CHI ’25,New York, NY, USA,pp\. 1–22\.External Links:ISBN 979\-8\-4007\-1394\-1,[Link](https://dl.acm.org/doi/10.1145/3706598.3713778),[Document](https://dx.doi.org/10.1145/3706598.3713778)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[26\]L\. I\. Lin\(1989\)A Concordance Correlation Coefficient to Evaluate Reproducibility\.Biometrics45\(1\),pp\. 255–268\.Note:Publisher: \[Wiley, International Biometric Society\]External Links:ISSN 0006\-341X,[Link](https://www.jstor.org/stable/2532051),[Document](https://dx.doi.org/10.2307/2532051)Cited by:[§IV\-A](https://arxiv.org/html/2607.20461#S4.SS1.p1.4)\.
- \[27\]N\. Lourie, R\. L\. Bras, and Y\. Choi\(2021\-05\)SCRUPLES: A Corpus of Community Ethical Judgments on 32,000 Real\-Life Anecdotes\.Proceedings of the AAAI Conference on Artificial Intelligence35\(15\),pp\. 13470–13479\(en\)\.External Links:ISSN 2374\-3468,[Link](https://ojs.aaai.org/index.php/AAAI/article/view/17589),[Document](https://dx.doi.org/10.1609/aaai.v35i15.17589)Cited by:[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1)\.
- \[28\]B\. W\. Matthews\(1975\-10\)Comparison of the predicted and observed secondary structure of T4 phage lysozyme\.Biochimica et Biophysica Acta \(BBA\) \- Protein Structure405\(2\),pp\. 442–451\.External Links:ISSN 0005\-2795,[Link](https://www.sciencedirect.com/science/article/pii/0005279575901099),[Document](https://dx.doi.org/10.1016/0005-2795%2875%2990109-9)Cited by:[§IV\-C](https://arxiv.org/html/2607.20461#S4.SS3.p1.2)\.
- \[29\]A\. Matthias\(2004\-09\)The responsibility gap: Ascribing responsibility for the actions of learning automata\.Ethics and Information Technology6\(3\),pp\. 175–183\(en\)\.External Links:ISSN 1572\-8439,[Link](https://doi.org/10.1007/s10676-004-3422-1),[Document](https://dx.doi.org/10.1007/s10676-004-3422-1)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[30\]L\. McCormack and M\. Bendechache\(2025\-06\)A comprehensive survey and classification of evaluation criteria for trustworthy artificial intelligence\.AI and Ethics5\(3\),pp\. 1973–1994\(en\)\.External Links:ISSN 2730\-5961,[Link](https://doi.org/10.1007/s43681-024-00590-8),[Document](https://dx.doi.org/10.1007/s43681-024-00590-8)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p1.1)\.
- \[31\]L\. McCormack and M\. Bendechache\(2026\-06\)The Trustworthy AI Maturity Model \(TAIMM\): Integrating ethics and regulation across the AI lifecycle\.Journal of Responsible Technology26,pp\. 100156\.External Links:ISSN 2666\-6596,[Link](https://www.sciencedirect.com/science/article/pii/S2666659626000090),[Document](https://dx.doi.org/10.1016/j.jrt.2026.100156)Cited by:[§VI](https://arxiv.org/html/2607.20461#S6.p2.1)\.
- \[32\]J\. S\. Mill\(2001\)Utilitarianism\.2nd ed\. edition,Hackett Publishing Company,Indianapolis/Cambridge\.Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1)\.
- \[33\]J\. Rawls\(1999\)A theory of justice\.The Belknap Press of Harvard University Press,Cambridge, Mass\.Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1),[§II](https://arxiv.org/html/2607.20461#S2.p1.1)\.
- \[34\]M\. Sap, S\. Gabriel, L\. Qin, D\. Jurafsky, N\. A\. Smith, and Y\. Choi\(2020\-07\)Social Bias Frames: Reasoning about Social and Power Implications of Language\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 5477–5490\.External Links:[Link](https://aclanthology.org/2020.acl-main.486/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.486)Cited by:[§III\-A](https://arxiv.org/html/2607.20461#S3.SS1.p1.1)\.
- \[35\]P\. Singer\(1993\)Practical ethics, 2nd edition\.Cambridge University Press\.Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1)\.
- \[36\]L\. Stappen, L\. Schumann, B\. Sertolli, A\. Baird, B\. Weigell, E\. Cambria, and B\. W\. Schuller\(2021\-10\)MuSe\-Toolbox: The Multimodal Sentiment Analysis Continuous Annotation Fusion and Discrete Class Transformation Toolbox\.InProceedings of the 2nd on Multimodal Sentiment Analysis Challenge,MuSe ’21,New York, NY, USA,pp\. 75–82\.External Links:ISBN 978\-1\-4503\-8678\-4,[Link](https://dl.acm.org/doi/10.1145/3475957.3484451),[Document](https://dx.doi.org/10.1145/3475957.3484451)Cited by:[§IV\-A](https://arxiv.org/html/2607.20461#S4.SS1.p2.5),[§IV\-A](https://arxiv.org/html/2607.20461#S4.SS1.p2.6)\.
- \[37\]S\. Tolmeijer, M\. Kneer, C\. Sarasua, M\. Christen, and A\. Bernstein\(2021\-12\)Implementations in Machine Ethics: A Survey\.ACM Comput\. Surv\.53\(6\),pp\. 132:1–132:38\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3419633),[Document](https://dx.doi.org/10.1145/3419633)Cited by:[§I](https://arxiv.org/html/2607.20461#S1.p2.1),[§II](https://arxiv.org/html/2607.20461#S2.p1.1)\.
- \[38\]J\. Trager, A\. S\. Ziabari, E\. Rahmati, A\. M\. Davani, P\. Golazizian, F\. Karimi\-Malekabadi, A\. Omrani, Z\. Li, B\. Kennedy, N\. K\. Reimer, M\. Reyes, K\. Cheng, M\. Wei, C\. Merrifield, A\. Khosravi, E\. Alvarez, and M\. Dehghani\(2025\-10\)The Moral Foundations Reddit Corpus\.arXiv\.Note:arXiv:2208\.05545 \[cs\]External Links:[Link](http://arxiv.org/abs/2208.05545),[Document](https://dx.doi.org/10.48550/arXiv.2208.05545)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p1.1),[§II](https://arxiv.org/html/2607.20461#S2.p2.1)\.
- \[39\]N\. Yongsatianchot, T\. Thejll\-Madsen, and S\. Marsella\(2023\-09\)What’s Next in Affective Modeling? Large Language Models\.In2023 11th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos \(ACIIW\),pp\. 1–7\.External Links:[Link](https://ieeexplore.ieee.org/document/10388124),[Document](https://dx.doi.org/10.1109/ACIIW59127.2023.10388124)Cited by:[§VI](https://arxiv.org/html/2607.20461#S6.p2.1)\.
- \[40\]L\. Zangari, C\. M\. Greco, D\. Picca, and A\. Tagarelli\(2025\-08\)A survey on moral foundation theory and pre\-trained language models: current advances and challenges\.AI & SOCIETY40\(6\),pp\. 4973–4998\(en\)\.External Links:ISSN 1435\-5655,[Link](https://doi.org/10.1007/s00146-025-02225-w),[Document](https://dx.doi.org/10.1007/s00146-025-02225-w)Cited by:[§II](https://arxiv.org/html/2607.20461#S2.p2.1)\.Similar Articles
Do Emotions Influence Moral Judgment in Large Language Models?
University of Cincinnati researchers show that adding positive or negative emotions to prompts can flip LLMs’ moral acceptability judgments in ~20% of cases, revealing an emotion-driven alignment gap with humans.
Negative Before Positive: Asymmetric Valence Processing in Large Language Models
This paper investigates how large language models process emotional valence through mechanistic interpretability. Using activation patching and steering on three open-source LLMs, the authors find that negative valence is localized to early layers while positive valence peaks in mid-to-late layers, and they validate this through topic-controlled flip tests.
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
This paper introduces a psychometric battery to separate framing artifacts from genuine moral judgment in LLMs, finding that frontier models have a coherent internal moral scale but display a yes/no bias that is purely a surface-level artifact of answer order and wording, not a real disposition to reject.
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
A systematic study on detecting Schwartz values in political text, comparing context lengths, model sizes, and retrieval-augmented generation methods. Results show that full-document context improves supervised models but not zero-shot LLMs, while retrieved moral knowledge consistently helps via early fusion.
Accounting for Context: Shaping Moral Credences for Value Alignment
This paper argues that aggregating moral evaluations for AI value alignment must account for contextual factors, showing that ignoring context can lead to violations of the weak Pareto principle, analogous to Simpson's paradox.