Tag
This paper derives sharp theoretical limits for honest uncertainty in repeated evaluation under a hard budget, with applications to language model and agent benchmarking, demonstrating practical improvements in interval width and MSE.
This paper measures how AI agents contaminate stores they write to, identifying a threshold for error propagation and validating findings with synthetic and real-world Wikidata data.
The article explains how to use Jev, a model for structured decisions, as an efficient judge for evaluating AI agent responses, reducing latency and cost compared to traditional LLM judges.
This article evaluates Jev, a System One model from TypeSafe AI, as a new agent evaluator, showing it outperforms LLM judges in consistency, speed, and cost.
The article discusses a critical automation failure mode where actions succeed but responses time out, leading to duplicates, and advocates for using stable operation IDs and state checks to improve agent evaluation robustness.
The paper introduces CHASE, a counterfactual-guided method for harness evolution in AI agent evaluation that mitigates benchmark shortcuts, enhancing reliability through constraint generation and validity checks.
A comparison test between MiniCPM5-2B and Spark-X2.5-4B on a multi-step customer-service agent task shows MiniCPM5-2B completing the task accurately and efficiently, while Spark-X2.5-4B underperformed, highlighting the role of agency in task execution.
The article explains evaluation methods for AI agents, covering easy, hard, and advanced modes with practical advice on tools like Harbor and LangSmith for safe testing and self-improvement.
Clay relies on LangSmith's online evaluators to scale manual review for millions of agent runs per month and is testing its insights product to better understand agent behavior.
DAREBench is a benchmark designed for deployment-aware and reliable evaluation of AI models as agents, assessing multimodal tasks and execution forms, with results showing trade-offs in accuracy and cost across commercial and open-weight models.
Sierra AI open-sources hyper-𝜏-bench, a new benchmark that evaluates AI models' ability to construct customer-service agents, revealing limitations in autonomous builds and improvements with human assistance.
KC-Bench is a dynamic interactive benchmark for evaluating how LLM agents detect and resolve knowledge conflicts between user instructions, parametric knowledge, and environmental observations, with evaluations showing no model handles all conflict types reliably.
The paper introduces A-CEGIS, a framework that uses counterexample feedback to evaluate and improve multi-turn agent self-correction in natural-language-to-regex synthesis, achieving high success rates.
The paper introduces ClaimReceipt, a claim-relative receipt specification and verifier for verifying evidence sufficiency and coverage in agent evaluations, validated through experiments on historical records and prospective audits.
FM-Bench is a benchmark for evaluating long-horizon AI agents, revealing that short-term performance does not guarantee long-term success in simulated management tasks over 20 years.
EarlyEval introduces a lightweight framework that predicts LLM agent outcomes from intermediate behavior to reduce evaluation costs by halting runs early with high accuracy. It eliminates 13%-26% of agent steps and up to 44% input tokens across benchmarks like SWE-bench Verified, TerminalBench, and Toolathlon.
CommerceAgentBench is a new benchmark with 107 real-world e-commerce tasks designed to test whether AI agents can actually complete jobs, moving beyond traditional answer-based AI benchmarks.
Introduces HarnessOpt-Bench to measure recursive self-improvement in AI, evaluating 5 frontier models on 4 tasks and finding that model choice has a greater impact than coding harness choice.
UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior.
MerchantBench is a benchmark that assesses AI agents by having them manage a simulated online store for a year, revealing challenges with sustained performance and continuous action.