linucb

标签

Cards List
#linucb

Progressive Content Refinement with Decaying Reward Joint LinUCB

arXiv cs.CL · 2026-08-10 缓存

This paper proposes a novel contextual bandit algorithm that explicitly models reward decay for progressive content refinement in LLMs, using EM to estimate arm-specific and decay parameters. Experiments on Sentiment Reversal and GSM8K show significant gains over strong baselines.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈