Bounded Morality: Defining the Space of Moral Computation
Summary
This paper extends Herbert Simon's bounded rationality to moral cognition, formalizing a tradeoff between moral breadth and moral depth for finite agents. It argues that ethical theories are locally optimal strategies within a constrained space, with implications for AI alignment.
View Cached Full Text
Cached at: 07/02/26, 05:40 AM
# Bounded Morality: Defining the Space of Moral Computation Source: [https://arxiv.org/abs/2607.00002](https://arxiv.org/abs/2607.00002) [View PDF](https://arxiv.org/pdf/2607.00002) > Abstract:Moral cognition has traditionally been modeled as adherence to fixed ethical theories\-\-deontology, consequentialism, virtue ethics\-\-implemented as static rules or value functions\. We propose Bounded Morality, a formal framework for analyzing the computational demands of moral problems faced by finite agents\. Extending Herbert Simon's notion of bounded rationality, we formalize moral situations along two orthogonal dimensions: moral breadth, the scope of entities treated as morally relevant, and moral depth, the inferential integration required to evaluate their interactions\. Limited resources impose an unavoidable tradeoff between these dimensions, defining a feasible space of moral computation\. Within this space, ethical theories correspond to locally efficient strategies adapted to different demand regimes rather than competing accounts of moral truth\. The framework yields a formal notion of moral regret and moral progress under constraint, and implies that moral alignment in artificial systems depends on the scaling and allocation of moral reasoning capacity rather than on direct imitation of human judgments\. ## Submission history From: Max Kanwal \[[view email](https://arxiv.org/show-email/6b7b7b6b/2607.00002)\] **\[v1\]**Wed, 1 Apr 2026 17:35:43 UTC \(85 KB\)
Similar Articles
Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning
This paper uses mechanistic interpretability to audit ethical reasoning in LLaMA 3.1-8B-Instruct, finding a 'Situational Anchor Effect' where domain-specific representations dominate moral computation, and proposing 'Mechanistic Alignment' as a research program.
Accounting for Context: Shaping Moral Credences for Value Alignment
This paper argues that aggregating moral evaluations for AI value alignment must account for contextual factors, showing that ignoring context can lead to violations of the weak Pareto principle, analogous to Simpson's paradox.
Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs
This paper introduces Moral Trolley Arena, a benchmark to evaluate how LLMs compose multiple moral signals within a single option, finding that composite judgments are compressed rather than additive.
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
This paper introduces MCLASH, a multilingual moral decision-making benchmark, and MET, a theory-grounded prompting method for culture-aware moral reasoning, along with MET-D, a self-distillation training method that improves multilingual moral reasoning across different model families and sizes.
Beyond Goodhart's Law: A Dynamic Benchmark for Evaluating Compliance in Multi-Agent Systems
This paper introduces MAC-Bench, a dynamic adversarial benchmark for evaluating procedural compliance in multi-agent systems. It proposes the SERV pipeline to generate contamination-free scenarios and new metrics like Compliance-Weighted Success Rate (CSR) and Machiavellian Gap (MG).