preference-based-rl

Tag

Cards List
#preference-based-rl

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

arXiv cs.LG · 2026-08-05 Cached

Introduces SP3O, a novel reward-model-free, critic-free, gradient-based preference-based RL algorithm that leverages segment-level preferences, demonstrating improved performance in robotic control and LLM fine-tuning, especially for long-horizon tasks.

0 favorites 0 likes
#preference-based-rl

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

arXiv cs.AI · 2026-08-03 Cached

This paper introduces LEMUR, a framework that combines multi-objective reinforcement learning with preference-based learning from multiple human feedback to learn Pareto-optimal policies without predefined reward functions.

0 favorites 0 likes
#preference-based-rl

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF

arXiv cs.AI · 2026-07-22 Cached

S2T-RLHF proposes a sentence-to-token reward decomposition framework that improves training stability and robustness in preference-based RLHF by assigning sequence-level rewards at sentence granularity, avoiding the instability of overly fine-grained token-level refinement.

0 favorites 0 likes
← Back to home

Submit Feedback