Tag
This paper audits offline root-cause-analysis benchmarks and finds that pooled leaderboards hide subsystem-specific winners, using pairwise comparisons on 778 cases across 11 subsystems. It releases a 320-line audit module for recomputing per-subsystem stability checks.
Papers with Code, a platform for AI research papers and code, has been rebuilt by Hugging Face after being acquired by Meta, providing a centralized hub for studies, code implementations, and task-specific leaderboards.
When several AI models were priced equally for a week, actual token usage revealed preference differences from leaderboard rankings, showing that coding and general chat have different top models and long context usage concentrated on two trusted models.
This paper argues that aggregate-score leaderboards for LLM agent benchmarks fail to capture deployment-relevant dimensions and show rank instability. It proposes ranking configurations by predictive validity—the correlation between in-sample and out-of-sample rank—and introduces a twelve-tier measurement apparatus along with falsifiable out-of-distribution criteria.
Papers Without Code now lets users browse evaluation results and leaderboards for closed-source models like GPT-5.5 and Opus 4.8, with a toggle in settings.
This article recommends a website called Sophon, which aggregates AI papers, models, benchmarks, leaderboards, and reinforcement learning environments. It provides real-time rankings, comparisons, and subscription features, and is hailed as the Bloomberg terminal for AI research.
This paper presents a qualitative study based on interviews with CS researchers, revealing a paradox of pragmatic skepticism where researchers distrust LLM leaderboard rankings yet continue to use them as rough guides. It finds that peer networks are primary for model selection, arena-based leaderboards are preferred, and cost transparency is the most demanded feature.
Niels from Hugging Face announces new features for the revived PapersWithCode platform, including multi-metric leaderboards, support for external papers, paper lineage, and more.
Niels from Hugging Face announces the revival of PapersWithCode as paperswithcode.co, a platform that parses high-impact AI papers at scale and automatically generates leaderboards and benchmarks, incorporating features like trending papers, domain categorization, and external paper support.
NielsRogge announces a revival of PapersWithCode, featuring SOTA per domain, leaderboards, and methods parsed at scale using AI agents.