标签
介绍了TREAT,这是一个基准测试,用于评估大型语言模型是否能从数学公式的等价保持变换中恢复已知定理身份。测试中表现最好的模型准确率仅为60.73%,表明定理知识在表示变化下是脆弱的。
This paper introduces Hybrid Ontology Matching (HOM), unifying equivalence and subsumption discovery, and proposes AgentMap, an LLM-based multi-agent framework for joint ontology matching. Experiments show promising results on hybrid, equivalence-only, and subsumption-only settings.