Tag
This paper presents a diagnostic protocol for selecting rewards and policies in delayed-feedback contextual bandits, arguing that standard offline evaluation can mislead and validating the approach on benchmarks and a deployed push system.
The paper introduces a behavioral alignment framework for personalized LLM judges in recommendation evaluation, addressing bidirectional rationalization where off-the-shelf LLMs argue both for and against user engagement on the same item. Their fine-tuned and preference-optimized approach achieves a 32.19% Macro-F1 lift over zero-shot and matches production feature-engineered baselines.
The paper argues that detecting an average effect of an acquired LLM-derived signal is not the same as learning per-instance acquisition policies, and establishes a reward-SNR floor (ρ* ≈ 2.8/√N) below which offline routing is impossible. It introduces Structured Hypothesis Embeddings (SHE) and shows across three datasets that learned per-example acquisition collapses below this floor, recommending design-time regime gates instead.