Benchmarking LLMs
Summary
A study or report on benchmarking large language models, likely comparing performance across various tasks.
Similar Articles
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
This paper proposes a dataset-centric meta-evaluation framework that audits LLM benchmarks at the sample level across five latent dimensions, exposing internal heterogeneity and enabling criterion-driven composition of benchmark subsets for targeted model evaluation.
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
RepBench is a benchmark-grounded data layer for representation probing in LLMs, built from 13,427 benchmark papers and 353 public datasets, yielding 46,149 probe texts across 94 capabilities. It evaluates capability vector transfer across models and readout methods.
Benchmarking Language Models for Statistical Problem Formulation
The paper introduces StatFormBench, a benchmark for evaluating LLMs on statistical problem formulation, and finds that current models have significant limitations in classifying problems and identifying variables.
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
CulturALL introduces a 2,610-sample benchmark across 14 languages and 51 regions to evaluate LLMs on real-world, culturally grounded tasks; top model scores only 44.48%, highlighting large room for improvement.
Benchmarking the Personalization Capabilities of Large Language Models
This paper introduces SDR-Bench, a benchmark for evaluating the personalization capabilities of large language models in a two-party Bayesian Persuasion framework, finding a consistent plateau across frontier LLMs and validating the framework with a field deployment.