Tag
This paper introduces Cargo, a framework for evaluating agentic AI systems that addresses reference-instance divergence by grounding factual judgments in live context and gating evaluation on retrieval confidence, along with Cargo-Bench for benchmarking.