Tag
Introduces TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving transformations of mathematical formulas. The best tested model achieves only 60.73% accuracy, showing that theorem knowledge is fragile under representation changes.