Tag
This paper proposes Evidence Sufficiency Boundary Training to enhance selective answering in grounded multi-hop QA systems by learning when to abstain from answering due to insufficient evidence, leading to improved boundary localization and reduced unsupported-answer rates on benchmarks like HotpotQA and MuSiQue.
FinRAG-12B is a 12B-parameter LLM optimized for retrieval-augmented generation in banking, featuring a unified training framework that improves answer quality, citation grounding, and calibrated refusal. The model outperforms GPT-4.1 in citation grounding and is deployed across over 40 financial institutions with significant cost and latency advantages.