Tag
This paper formalizes a claim-replay layer for AI evaluation artifacts and censuses evaluation units, finding that most stop before deterministic inference due to missing historical evidence or semantic grounding.