Tag
The paper introduces StatFormBench, a benchmark for evaluating LLMs on statistical problem formulation, and finds that current models have significant limitations in classifying problems and identifying variables.