Tag
This paper evaluates citation faithfulness in agentic scientific synthesis systems, showing that current verifiers are unreliable with unsupported-citation rates varying from 3% to 18% depending on strictness. It proposes a gold-anchored evaluation protocol and a deployable guard that uses split-conformal prediction to provide a distribution-free bound on truly unsupported citations.
This paper investigates what makes a small 4B parameter model's citations faithful when used as an on-device research agent, finding that exposure to more source text improves faithfulness while retrieval recall limits coverage.