Tag
The paper proposes the Plan, Learn, Adapt (PLA) framework for personalized on-device itinerary generation, combining feasibility-guaranteed combinatorial planning with human preference learning via a Bradley-Terry reward model. In deployment, it achieved a 91% increase in itinerary completion rates with low latency, outperforming frontier LLMs in feasibility.
PerceptionRubrics introduces a rubric-based evaluation framework for multimodal models that uses atomic auditing and gated scoring to better align benchmark scores with human perception, revealing reliability gaps and open-closed stratification.
This paper introduces CLIPR, a framework that learns transferable latent user preferences from minimal conversational input to improve human-aligned decision making in LLMs.