Tag
This paper introduces the Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to quantify how much LLM accuracy varies under different prompt wrappers. Through 140,000 generations across models and tasks, it shows that wrapper choice can drastically affect scores, with parseability failures being a key driver.