leaderboards

Tag

Cards List
#leaderboards

Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks

arXiv cs.AI · 2026-06-30 Cached

This paper audits offline root-cause-analysis benchmarks and finds that pooled leaderboards hide subsystem-specific winners, using pairwise comparisons on 778 cases across 11 subsystems. It releases a 320-line audit module for recomputing per-subsystem stability checks.

0 favorites 0 likes
#leaderboards

@lucas_flatwhite: For those of you who follow AI research/agent trends, be sure to bookmark this..! Papers with Code https://paperswithco…

X AI KOLs Timeline · 2026-06-27 Cached

Papers with Code, a platform for AI research papers and code, has been rebuilt by Hugging Face after being acquired by Meta, providing a centralized hub for studies, code implementations, and task-specific leaderboards.

0 favorites 0 likes
#leaderboards

Some models got priced the same for a week, so I watched what people actually used

Reddit r/ArtificialInteligence · 2026-06-25

When several AI models were priced equally for a week, actual token usage revealed preference differences from leaderboard rankings, showing that coding and general chat have different top models and long context usage concentrated on two trusted models.

0 favorites 0 likes
#leaderboards

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Hugging Face Daily Papers · 2026-06-18 Cached

This paper argues that aggregate-score leaderboards for LLM agent benchmarks fail to capture deployment-relevant dimensions and show rank instability. It proposes ranking configurations by predictive validity—the correlation between in-sample and out-of-sample rank—and introduces a twelve-tier measurement apparatus along with falsifiable out-of-distribution criteria.

0 favorites 0 likes
#leaderboards

@NielsRogge: Introducing Papers Without Code Jokes aside, you can now also browse the evaluation results and leaderboards of closed-…

X AI KOLs Following · 2026-06-08 Cached

Papers Without Code now lets users browse evaluation results and leaderboards for closed-source models like GPT-5.5 and Opus 4.8, with a toggle in settings.

0 favorites 0 likes
#leaderboards

@vikingmute: Who created this amazing website? https://sophon.at It collects and displays all AI-related information and content: papers, newest models, benchmarks, leaderboards. Papers can be viewed online directly, very comprehensive. Also has a feed to subscribe for the latest news. And this…

X AI KOLs Timeline · 2026-06-05 Cached

This article recommends a website called Sophon, which aggregates AI papers, models, benchmarks, leaderboards, and reinforcement learning environments. It provides real-time rankings, comparisons, and subscription features, and is hailed as the Bloomberg terminal for AI research.

0 favorites 0 likes
#leaderboards

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

arXiv cs.CL · 2026-05-29 Cached

This paper presents a qualitative study based on interviews with CS researchers, revealing a paradox of pragmatic skepticism where researchers distrust LLM leaderboard rankings yet continue to use them as rough guides. It finds that peer networks are primary for model selection, arena-based leaderboards are preferred, and cost transparency is the most demanded feature.

0 favorites 0 likes
#leaderboards

PapersWithCode new features - week 1 [P]

Reddit r/MachineLearning · 2026-05-24

Niels from Hugging Face announces new features for the revived PapersWithCode platform, including multi-metric leaderboards, support for external papers, paper lineage, and more.

0 favorites 0 likes
#leaderboards

Reviving PapersWithCode (by Hugging Face) [P]

Reddit r/MachineLearning · 2026-05-18

Niels from Hugging Face announces the revival of PapersWithCode as paperswithcode.co, a platform that parses high-impact AI papers at scale and automatically generates leaderboards and benchmarks, incorporating features like trending papers, domain categorization, and external paper support.

0 favorites 0 likes
#leaderboards

@NielsRogge: Introducing a revival of PapersWithCode! As @ilyasut said, we're back to the "age of research". Hence, it's important t…

X AI KOLs Following · 2026-05-18 Cached

NielsRogge announces a revival of PapersWithCode, featuring SOTA per domain, leaderboards, and methods parsed at scale using AI agents.

0 favorites 0 likes
← Back to home

Submit Feedback