pomo

标签

Cards List
#pomo

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

arXiv cs.LG · 2026-08-04 缓存

This paper presents a narrow extension to Leader Reward training for neural combinatorial optimization, replacing the binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget K. Tests on TSP-100 show modest improvements in Best-of-8 cost under independent sampling, though the authors make no universal superiority claims.

0 人收藏 0 人点赞
← 返回首页

提交意见反馈