reliability-evaluation

Tag

Cards List
#reliability-evaluation

Evidence-State Reliability Under Controlled Degradation: Parser-Validity Divergence in a Multi-Stage LLM Pipeline

arXiv cs.CL · 2026-08-25 Cached

This paper introduces Evidence-State Reliability (ESR) as an evaluation layer for multi-stage LLM pipelines, showing that structural conformance can improve while evidence-sensitive stage success deteriorates under controlled degradation.

0 favorites 0 likes
#reliability-evaluation

Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

Hugging Face Daily Papers · 2026-08-12 Cached

Claim-Level Reliability Assessment (CLR) is a training-free framework that improves reasoning accuracy in LLMs by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance.

0 favorites 0 likes
#reliability-evaluation

@steverab: Very excited to share that our paper "Towards a Science of AI Agent Reliability" was accepted at ICML 2026! See you in …

X AI KOLs Timeline · 2026-06-05 Cached

A paper analyzing AI agent reliability, accepted at ICML 2026, finds that even the latest frontier models (GPT 5.5, Gemini 3.1 Pro, Claude Opus 4.7) show only marginal reliability improvements over earlier versions, with low outcome consistency and persistent issues in agent scaffolding.

0 favorites 0 likes
← Back to home

Submit Feedback