Tag
The paper quantifies how improvements in fact-verification scores are partitioned between answer accuracy and evidence quality, using trained DeBERTa checkpoints and LLMs across multiple benchmarks.