Tag
Introduces TREAT, a benchmark for evaluating whether large language models can recover known theorem identities from equivalence-preserving transformations of mathematical formulas. The best tested model achieves only 60.73% accuracy, showing that theorem knowledge is fragile under representation changes.
This paper introduces Hybrid Ontology Matching (HOM), unifying equivalence and subsumption discovery, and proposes AgentMap, an LLM-based multi-agent framework for joint ontology matching. Experiments show promising results on hybrid, equivalence-only, and subsumption-only settings.