Tag
The paper introduces a controlled benchmark for evaluating Large Language Models' robustness in step-level mathematical verification, revealing significant performance degradation on perturbed solution traces.