Benchmarking LLMs
Summary
A study or report on benchmarking large language models, likely comparing performance across various tasks.
Similar Articles
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
CulturALL introduces a 2,610-sample benchmark across 14 languages and 51 regions to evaluate LLMs on real-world, culturally grounded tasks; top model scores only 44.48%, highlighting large room for improvement.
Benchmarking the Personalization Capabilities of Large Language Models
This paper introduces SDR-Bench, a benchmark for evaluating the personalization capabilities of large language models in a two-party Bayesian Persuasion framework, finding a consistent plateau across frontier LLMs and validating the framework with a field deployment.
Benchmarking Different Methods of LLM Confidence Estimation
This article benchmarks various blackbox and whitebox methods for LLM confidence estimation, including verbalized confidence, linguistic uncertainty, reasoning-length, P(Answer), P(True), and self-consistency, comparing their effectiveness for tasks like active learning and safety classification.
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
The article introduces DataGovBench, a benchmark derived from governmental open data, designed to evaluate LLMs on real-world data analysis tasks including table question answering and insight discovery. Experiments show current LLMs still underperform in complex data analytics scenarios.
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
This article introduces Magis-Bench, a benchmark for evaluating large language models on magistrate-level legal tasks such as judicial reasoning and sentence drafting, using data from Brazilian judicial exams.