Tag
This paper critiques the use of pairwise comparisons for learning human preferences, arguing that internal pluralism (multiple conflicting priorities) undermines the standard approach. It proposes a formal model and suggests that allowing indecision can improve learning efficiency.
Trace2Policy extracts human-readable decision rules from expert behavior traces and iteratively refines them via error-driven skill refinement, outperforming pure LLM baselines on compliance-sensitive tasks in logistics.