Tag
A comparison shows that LLM performance on the 2026 IMO varies dramatically depending on the evaluation harness, with structured multi-agent setups achieving far higher scores than simple web UI, indicating that current gains are absorbed at the frontier by better orchestration.
Discusses a potential vulnerability in the ARC AGI 3 benchmark where the Opus model could be gamed if it functions as a loop rather than a pure model.
A comparison of two AI coding agents building a Mario game: Laguna S 2.1 in Poolside's agent took 62 minutes with self-correction and passed tests, while a previous Qwen model took hours and needed human help; highlights oracle discipline and native harness advantages.
A self-awareness exam given to Claude, Gemini, and Grok found Claude and Gemini nearly aced it, while Grok scored 0 on self-recognition, failing to identify its own writing. The study measures six dimensions of functional self-knowledge.
Introduces SciHazard, a benchmark for measuring scientific safety risks in LLMs with a decomposed harm scoring framework, and evaluates 31 frontier models, finding deep research agents pose higher risks.
BACON proposes a four-stage pipeline that combines budgeted human labels with multiple AI judge outputs to produce calibrated item-level surrogate predictions, supporting both population-level estimation and individual-level scoring with improved accuracy and reduced bias.
Analysis of 1,250 papers on recursive self-improvement in AI reveals that the evaluator signal is the critical bottleneck. Models improve reliably only with strong, trustable signals like proof checkers, while weak signals cause loops to collapse or reinforce errors.
The article predicts that the AI benchmark 'Agents Last Exam' will reach saturation (models maxing out performance) by February next year.
Advocates for treating agent failures as evaluation benchmarks rather than just trace logs, emphasizing the need for systematic testing of AI agent behaviors.
The Singularity Gate benchmark tests whether frontier AI models can predict paradigm-breaking scientific discoveries made after their training cutoff. Claude Fable 5 leads but has a low response rate due to refusals, while GPT-5.6 Sol shows strong performance without refusals at a lower price point.
Sayash Kapoor announces his upcoming PhD defense on 'The Missing Science of AI Evaluation,' discussing AI-based science, agent evaluation, and open model risks, livestreamed and open to all.
This paper introduces NameRank, a method to measure how well LLMs recognize individuals and artifacts from their parametric knowledge, finding that models recognize named projects and methods far better than credentials or contributors.
This paper from July 2026 defines principles for designing effective benchmark tasks for AI agents, drawing on experience from the Terminal Bench project. Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons, with an emphasis on real-world relevance and outcome-based verification.
A blog post from Parlance Labs tests automated AI evaluation tools (Braintrust Loop, Arize Alyx, LangSmith Engine) on real production data, finding they catch 87% of issues humans flag but miss domain-specific failures and add noise, recommending iterative human-in-the-loop use.
GPT-5.5, when equipped with tools, has surpassed the performance of a 10-year-old child on the BabyVision benchmark, marking a notable advancement in AI visual reasoning.
This paper introduces psychological competence as a missing dimension in AI evaluation, proposing a conceptual framework to assess how AI systems support user cognition, emotional interpretation, and decision-making in human-facing roles.
This paper audits whether self-consistency and cross-model agreement are reliable indicators of correctness in LLMs, finding that agreement is a weak, regime-dependent proxy and that frontier models exhibit overconfidence.
This paper proposes a new paradigm for measuring intelligence beyond human capability using adversarial psychometric rating systems, where models generate challenges to separate other systems, enabling evaluation that scales with AI capabilities.
This article benchmarks Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5 by having each model build three interactive apps (3D Rubik's Cube, particle gravity sandbox, Breakout game) from a single prompt, comparing their one-shot coding capabilities.
OpenAI emphasizes the need for more rigorous and trustworthy evaluations for coding AI models to better measure real progress.