Tag
Introduces Spider 2.0-AIFunc, a benchmark of 465 instances across 125 real-world databases for evaluating AI-native SQL queries that use cloud platform AI functions. Evaluates ten state-of-the-art models, finding proprietary models reach 67-70% accuracy while open-source models lag behind.
This paper presents a semantic-layer-mediated NL2SQL agent that decouples intent from physical execution by reasoning over a curated semantic model, achieving 94.15% execution accuracy on the Spider2-snow benchmark.