标签
This paper presents a narrow extension to Leader Reward training for neural combinatorial optimization, replacing the binary leader/non-leader distinction with a stabilized rank signal indexed by a sampling budget K. Tests on TSP-100 show modest improvements in Best-of-8 cost under independent sampling, though the authors make no universal superiority claims.
本文介绍了CALVER,一种无需训练的符号验证器,它根据Pearl的因果准则对结构化因果推理轨迹进行评分,以选出最佳候选答案;在存在多个有效答案的因果推理基准上,其表现优于plurality投票和其他选择方法。