GoBench: Evaluating LLMs on the game of Go [R]
Summary
GoBench is a benchmark for evaluating large language models on 9x9 Go games, demonstrating strong correlation with ARC-AGI and featuring a leaderboard with KataGo opponents from random to superhuman levels.
Similar Articles
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
XLGoBench introduces a synthetic benchmark of algorithmic tasks to detect cross-lingual skill gaps in LLMs, demonstrating persistent gaps across multiple state-of-the-art models.
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
This article introduces Magis-Bench, a benchmark for evaluating large language models on magistrate-level legal tasks such as judicial reasoning and sentence drafting, using data from Brazilian judicial exams.
RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
RepBench is a benchmark-grounded data layer for representation probing in LLMs, built from 13,427 benchmark papers and 353 public datasets, yielding 46,149 probe texts across 94 capabilities. It evaluates capability vector transfer across models and readout methods.
Benchmarking LLMs
A study or report on benchmarking large language models, likely comparing performance across various tasks.
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement
CollabBench is a new benchmark for evaluating and training LLM agents in cooperative games, featuring diverse player simulation and a collaborative training paradigm. Experiments show 19.5% higher efficiency and 24.4% improved affective performance over base models.