Tag
This paper presents a narrow extension to Leader Reward training for neural combinatorial optimization, replacing the binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget K. Tests on TSP-100 show modest improvements in Best-of-8 cost under independent sampling, though the authors make no universal superiority claims.
This paper introduces CALVER, a training-free symbolic verifier that scores structured causal reasoning traces against Pearl's criteria to select the best candidate, outperforming plurality voting and other selection methods on causal reasoning benchmarks where multiple valid answers exist.