Tag
This preregistered reproduction study validates that the shape of chain-of-thought entropy trajectories predicts large language model answer correctness, while the total entropy drop is inconsistent across settings, and explores final-step entropy as an improved metric.