pomo

Tag

Cards List
#pomo

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

arXiv cs.LG · 2026-08-04 Cached

This paper presents a narrow extension to Leader Reward training for neural combinatorial optimization, replacing the binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget K. Tests on TSP-100 show modest improvements in Best-of-8 cost under independent sampling, though the authors make no universal superiority claims.

0 favorites 0 likes
← Back to home

Submit Feedback