Tag
Aiki Alpha 3 is released with substantial performance improvements, including cheaper runtime realization, adaptive number representations, enhanced profiling and coverage, and stricter library constraints.
This paper studies the trade-off between grounding and coverage in long-form hallucination reinforcement learning, proposing rubric-based rewards to represent required and optional information for questions. A soft combination of grounding, rubric coverage, and relevance yields the best balance between support and richness.
This paper proposes a formal framework for agent memory, defining memory as a basis, knowledge as its span, and optimality as capacity-constrained expected coverage, while framing continual memory as a sequential MDP and demonstrating the framework on Homer's Odyssey.
Introduces CAPEval, a caption evaluation framework that decouples coverage and precision, showing that coverage better predicts vision-language understanding performance while precision better predicts text-to-image generation performance.
This paper identifies the 'modal ceiling' and 'correlation ceiling' in test-time scaling for reasoning models, showing that beyond a few dozen samples, additional sampling does not improve selection accuracy and can even harm it, highlighting the identifiability gap between generating and recognizing correct answers.
This paper identifies a blind spot in reference-free faithfulness metrics: they only measure precision (whether claims are supported) but not recall (coverage of relevant facts). The authors introduce a complete-oracle evaluation using Formula 1 telemetry and weather data, showing that high-precision models often have poor coverage, and propose a combined metric.