Tag
This paper introduces Evidence-State Reliability (ESR) as an evaluation layer for multi-stage LLM pipelines, showing that structural conformance can improve while evidence-sensitive stage success deteriorates under controlled degradation.
Claim-Level Reliability Assessment (CLR) is a training-free framework that improves reasoning accuracy in LLMs by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance.
A paper analyzing AI agent reliability, accepted at ICML 2026, finds that even the latest frontier models (GPT 5.5, Gemini 3.1 Pro, Claude Opus 4.7) show only marginal reliability improvements over earlier versions, with low outcome consistency and persistent issues in agent scaffolding.