Tag
Introduces ReliableTableQA, a framework for training LLMs to annotate statistical reliability of tabular QA results, showing that a small SFT set is sufficient and GRPO only helps when SFT is under-trained.
The article introduces DataGovBench, a benchmark derived from governmental open data, designed to evaluate LLMs on real-world data analysis tasks including table question answering and insight discovery. Experiments show current LLMs still underperform in complex data analytics scenarios.