Tag
This paper introduces LivingArena, an automated evaluation framework where LLMs probe each other's weaknesses by generating questions, enabling contamination-resistant and scalable assessment that adapts as models improve.
This paper introduces PAJAMA, a system that distills LLM-as-a-judge decision logic into programmatic judges, eliminating per-sample API costs while matching the performance of a 13B LLM judge, and providing transparency and efficiency improvements for scalable evaluation.
This paper proposes a new paradigm for measuring intelligence beyond human capability using adversarial psychometric rating systems, where models generate challenges to separate other systems, enabling evaluation that scales with AI capabilities.
The paper proposes an attack-agnostic robustness metric based on the spectral norm of the Fisher Information Matrix, providing theoretical bounds and scalable evaluation methods for deep neural networks.