Tag
This paper proposes separating generation from selection in explainable-recommendation systems to reduce serving costs, using a frozen candidate pool of explanations and a small CPU-resident selector. It benchmarks offline-pool selectors and finds that pairwise learning-to-rank outperforms single-action RL formulations like PPO, GRPO, and DPO in terms of F1 scores.