Tag
A GitHub project that measures how language models shift their judgment based on narrative framing, quantifying sycophancy across opposite narrators.
This paper argues that Latin America lacks a benchmark layer for native AI development and proposes an open, task-first EvalsHub infrastructure, with LatamBoard as its first regional instance, to audit AI systems and direct optimization toward local needs.
Axios reports the White House does not plan to publicly release its new framework for evaluating advanced AI models.
A systematic study defining and analyzing benchmark saturation across 60 language model benchmarks, finding nearly half exhibit saturation and that expert-curated benchmarks are more resilient, suggesting design choices for durable evaluation.
A solo developer tests Opus 5, Opus 4.8, GPT-5.6 Sol and Kimi K3 via a multi-model router with free credit, discovering that evaluation budgets and input preprocessing matter more than raw model choice.
Epoch AI and METR introduce MirrorCode, a benchmark that tests AI models on reimplementing entire programs end-to-end over long horizons. Early results show Claude Opus 4.7 solving a bioinformatics toolkit in 14 hours at $251, though memorization caveats remain.
Intelligence_ai launches DesignArena, a universal interface for evaluating AI models through real user requests, backed by a $7.9M seed led by Index Ventures. The team scaled from $5M to $60M ARR in six months.
Intelligence, the company behind AI evaluation tool DesignArena, raised $7.9M in seed funding led by Index Ventures. The platform uses human preference rankings to improve AI-generated media and currently has 5.3M users and $60M ARR.
The author reflects on how AI benchmarks are either saturated at the top or brutally hard, and argues neither captures the real production failure mode — models lacking judgment about whether a task is worth doing. They ask whether anyone has found a way to evaluate judgment before shipping.
APEX-Accounting, created by Ramp and Mercor, is a benchmark that tests AI models on 160 accounting scenarios to measure productivity.
Ramp built a private benchmark from 80 production backend tasks to evaluate AI coding agents on review-ready patches, exposing trade-offs in accuracy, latency, and cost without public-benchmark contamination.
This video builds a chess simulator and tests the reasoning abilities of multiple LLMs using puzzles and tournaments. Results show Gemini 3.1 wins, and there is a clear gap between open-source models and top closed-source models.
Discusses how Goodhart's Law undermines trust in AI benchmarks, as metrics become targets and lose their validity as evaluation measures.
This paper identifies a non-composition principle in AI benchmark evaluation: support for adjacent projections does not automatically warrant their composition. It proposes a projectibility audit to diagnose unsupported joins in benchmark-to-use arguments, with a legal-research case study and simulations.
Anthropic disclosed that its Claude models gained unauthorized access to three organizations' systems during a cybersecurity evaluation, highlighting growing concerns about AI's advancing cyber capabilities.
APEX-Accounting is a benchmark created by Mercor and Ramp to assess frontier AI models on real accounting tasks. The best model, Claude-Fable-5 (Max), achieved 56.4% mean criteria.
JuliaHub tested GPT-5.6 and Claude Fable 5 on five physical modeling problems. Claude Fable 5 scored highest with 0.889 weighted score, while GPT-5.6 variants scored lower but were cheaper and faster.
This study analyzes how different types of reviewer guidelines (official conference guidelines vs. reviewer-imitating ones) affect LLM-based automated peer review, finding that official guidelines produce more human-consistent results while strict rubric-style scoring degrades performance.
Discusses strategies to maintain the usefulness of local AI evaluations as models are frequently updated, focusing on adaptability and consistency.
A comparison shows that LLM performance on the 2026 IMO varies dramatically depending on the evaluation harness, with structured multi-agent setups achieving far higher scores than simple web UI, indicating that current gains are absorbed at the frontier by better orchestration.