coverage

Tag

Cards List
#coverage

Aiki Alpha 3 Released - Performance Improved

Lobsters Hottest · 2026-08-30 Cached

Aiki Alpha 3 is released with substantial performance improvements, including cheaper runtime realization, adaptive number representations, enhanced profiling and coverage, and stricter library constraints.

0 favorites 0 likes
#coverage

From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning

arXiv cs.CL · 2026-08-14 Cached

This paper studies the trade-off between grounding and coverage in long-form hallucination reinforcement learning, proposing rubric-based rewards to represent required and optional information for questions. A soft combination of grounding, rubric coverage, and relevance yields the best balance between support and richness.

0 favorites 0 likes
#coverage

Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential Memory Problem

arXiv cs.LG · 2026-08-13 Cached

This paper proposes a formal framework for agent memory, defining memory as a basis, knowledge as its span, and optimality as capacity-constrained expected coverage, while framing continual memory as a sequential MDP and demonstrating the framework on Homer's Odyssey.

0 favorites 0 likes
#coverage

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation

Hugging Face Daily Papers · 2026-08-03 Cached

Introduces CAPEval, a caption evaluation framework that decouples coverage and precision, showing that coverage better predicts vision-language understanding performance while precision better predicts text-to-image generation performance.

0 favorites 0 likes
#coverage

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

arXiv cs.LG · 2026-06-30 Cached

This paper identifies the 'modal ceiling' and 'correlation ceiling' in test-time scaling for reasoning models, showing that beyond a few dozen samples, additional sampling does not improve selection accuracy and can even harm it, highlighting the identifiability gap between generating and recognizing correct answers.

0 favorites 0 likes
#coverage

Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle

Hugging Face Daily Papers · 2026-06-08 Cached

This paper identifies a blind spot in reference-free faithfulness metrics: they only measure precision (whether claims are supported) but not recall (coverage of relevant facts). The authors introduce a complete-oracle evaluation using Formula 1 telemetry and weather data, showing that high-precision models often have poor coverage, and propose a combined metric.

0 favorites 0 likes
← Back to home

Submit Feedback