Tag
The paper introduces ClaimProbe, a claim-level audit framework to detect hallucinations and misattributions in deep research reports, and ClaimWriter, a hierarchical writer that reduces hallucinations by up to 4.5 times while improving fact recall.
This paper formalizes a claim-replay layer for AI evaluation artifacts and censuses evaluation units, finding that most stop before deterministic inference due to missing historical evidence or semantic grounding.