offline-evaluation

Tag

Cards List
#offline-evaluation

When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection in Delayed-Feedback Contextual Bandits

arXiv cs.LG · yesterday Cached

This paper presents a diagnostic protocol for selecting rewards and policies in delayed-feedback contextual bandits, arguing that standard offline evaluation can mislead and validating the approach on benchmarks and a deployed push system.

0 favorites 0 likes
#offline-evaluation

From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation

arXiv cs.AI · yesterday Cached

The paper introduces a behavioral alignment framework for personalized LLM judges in recommendation evaluation, addressing bidirectional rationalization where off-the-shelf LLMs argue both for and against user engagement on the same item. Their fine-tuned and preference-optimized approach achieves a 32.19% Macro-F1 lift over zero-shot and matches production feature-engineered baselines.

0 favorites 0 likes
#offline-evaluation

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

arXiv cs.LG · 2d ago Cached

The paper argues that detecting an average effect of an acquired LLM-derived signal is not the same as learning per-instance acquisition policies, and establishes a reward-SNR floor (ρ* ≈ 2.8/√N) below which offline routing is impossible. It introduces Structured Hypothesis Embeddings (SHE) and shows across three datasets that learned per-example acquisition collapses below this floor, recommending design-time regime gates instead.

0 favorites 0 likes
← Back to home

Submit Feedback