Tag
This paper introduces LivingArena, an automated evaluation framework where LLMs probe each other's weaknesses by generating questions, enabling contamination-resistant and scalable assessment that adapts as models improve.
This paper introduces Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents designed to resist data contamination by reverse-engineering tasks from real commits and business scenarios, covering Code, Web, Office, and Security domains.