momdp

Tag

Cards List
#momdp

Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning

arXiv cs.LG · 2026-06-26 Cached

This paper introduces a novel preference-conditioned Bellman operator based on Chebyshev scalarization to compute deterministic Pareto-optimal policies for Multi-Objective Markov Decision Processes, proving its convergence and effectiveness in capturing the entire Pareto frontier.

0 favorites 0 likes
← Back to home

Submit Feedback