linucb

Tag

Cards List
#linucb

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv cs.CL · 2026-08-10 Cached

This paper proposes a novel contextual bandit algorithm that explicitly models reward decay for progressive content refinement in LLMs, using EM to estimate arm-specific and decay parameters. Experiments on Sentiment Reversal and GSM8K show significant gains over strong baselines.

0 favorites 0 likes
← Back to home

Submit Feedback