attribution-evaluation

Tag

Cards List
#attribution-evaluation

Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis

arXiv cs.AI · 2026-07-24 Cached

This paper evaluates citation faithfulness in agentic scientific synthesis systems, showing that current verifiers are unreliable with unsupported-citation rates varying from 3% to 18% depending on strictness. It proposes a gold-anchored evaluation protocol and a deployable guard that uses split-conformal prediction to provide a distribution-free bound on truly unsupported citations.

0 favorites 0 likes
← Back to home

Submit Feedback