scalable-evaluation

Tag

Cards List
#scalable-evaluation

LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation

arXiv cs.AI · 2026-07-29 Cached

This paper introduces LivingArena, an automated evaluation framework where LLMs probe each other's weaknesses by generating questions, enabling contamination-resistant and scalable assessment that adapts as models improve.

0 favorites 0 likes
#scalable-evaluation

Codifying the Judge: Scalable Evaluation via Program Distillation

arXiv cs.AI · 2026-07-28 Cached

This paper introduces PAJAMA, a system that distills LLM-as-a-judge decision logic into programmatic judges, eliminating per-sample API costs while matching the performance of a 13B LLM judge, and providing transparency and efficiency improvements for scalable evaluation.

0 favorites 0 likes
#scalable-evaluation

Measuring Intelligence Beyond Human Scale

arXiv cs.AI · 2026-07-09 Cached

This paper proposes a new paradigm for measuring intelligence beyond human capability using adversarial psychometric rating systems, where models generate challenges to separate other systems, enabling evaluation that scales with AI capabilities.

0 favorites 0 likes
#scalable-evaluation

Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms

Hugging Face Daily Papers · 2026-06-03 Cached

The paper proposes an attack-agnostic robustness metric based on the spectral norm of the Fisher Information Matrix, providing theoretical bounds and scalable evaluation methods for deep neural networks.

0 favorites 0 likes
← Back to home

Submit Feedback