agent-evaluation

Tag

Cards List
#agent-evaluation

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

arXiv cs.AI ↗ · yesterday Cached

This paper derives sharp theoretical limits for honest uncertainty in repeated evaluation under a hard budget, with applications to language model and agent benchmarking, demonstrating practical improvements in interval width and MSE.

0 favorites 0 likes
#agent-evaluation

Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures

arXiv cs.CL ↗ · 3d ago Cached

This paper measures how AI agents contaminate stores they write to, identifying a threshold for error propagation and validating findings with synthetic and real-world Wikidata data.

0 favorites 0 likes
#agent-evaluation

@akshay_pachaar: https://x.com/akshay_pachaar/status/2102087107410002345

X AI KOLs Timeline ↗ · 4d ago Cached

The article explains how to use Jev, a model for structured decisions, as an efficient judge for evaluating AI agent responses, reducing latency and cost compared to traditional LLM judges.

0 favorites 0 likes
#agent-evaluation

@LangChain: We tested Jev against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could o…

X AI KOLs Timeline ↗ · 6d ago Cached

This article evaluates Jev, a System One model from TypeSafe AI, as a new agent evaluator, showing it outperforms LLM judges in consistency, speed, and cost.

0 favorites 0 likes
#agent-evaluation

The automation failure nobody budgets for: the action landed, but the timeout said it failed

Reddit r/AI_Agents ↗ · 2026-09-17

The article discusses a critical automation failure mode where actions succeed but responses time out, leading to duplicates, and advocates for using stable operation IDs and state checks to improve agent evaluation robustness.

0 favorites 0 likes
#agent-evaluation

Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper introduces CHASE, a counterfactual-guided method for harness evolution in AI agent evaluation that mitigates benchmark shortcuts, enhancing reliability through constraint generation and validity checks.

0 favorites 0 likes
#agent-evaluation

MiniCPM5-2B vs. Spark-X2.5-4B / -64% thinking, x1.5 speed while keeping the accuracy of xhigh

Reddit r/LocalLLaMA ↗ · 2026-09-16

A comparison test between MiniCPM5-2B and Spark-X2.5-4B on a multi-step customer-service agent task shows MiniCPM5-2B completing the task accurately and efficiently, while Spark-X2.5-4B underperformed, highlighting the role of agency in task execution.

0 favorites 0 likes
#agent-evaluation

@businessbarista: A few of my smartest friends in AI called me a "idiot" for not deeply understanding evals. So...I found the smartest pe…

X AI KOLs Timeline ↗ · 2026-09-14 Cached

The article explains evaluation methods for AI agents, covering easy, hard, and advanced modes with practical advice on tools like Harbor and LangSmith for safe testing and self-improvement.

0 favorites 0 likes
#agent-evaluation

@LangChain: Manual review works at small scale. At millions of agent runs a month, it falls apart, so @Clay relies on LangSmith's o…

X AI KOLs Following ↗ · 2026-09-10 Cached

Clay relies on LangSmith's online evaluators to scale manual review for millions of agent runs per month and is testing its insights product to better understand agent behavior.

0 favorites 0 likes
#agent-evaluation

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

arXiv cs.AI ↗ · 2026-09-10 Cached

DAREBench is a benchmark designed for deployment-aware and reliable evaluation of AI models as agents, assessing multimodal tasks and execution forms, with results showing trade-offs in accuracy and cost across commercial and open-weight models.

0 favorites 0 likes
#agent-evaluation

Hyper-𝜏-bench: Evaluating agents that build agents (4 minute read)

TLDR AI ↗ · 2026-09-09 Cached

Sierra AI open-sources hyper-𝜏-bench, a new benchmark that evaluates AI models' ability to construct customer-service agents, revealing limitations in autonomous builds and improvements with human assistance.

0 favorites 0 likes
#agent-evaluation

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

arXiv cs.AI ↗ · 2026-09-04 Cached

KC-Bench is a dynamic interactive benchmark for evaluating how LLM agents detect and resolve knowledge conflicts between user instructions, parametric knowledge, and environmental observations, with evaluations showing no model handles all conflict types reliably.

0 favorites 0 likes
#agent-evaluation

Counterexamples as Feedback for Agent Self-Correction

arXiv cs.CL ↗ · 2026-09-04 Cached

The paper introduces A-CEGIS, a framework that uses counterexample feedback to evaluate and improve multi-turn agent self-correction in natural-language-to-regex synthesis, achieving high success rates.

0 favorites 0 likes
#agent-evaluation

ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations

arXiv cs.AI ↗ · 2026-09-03 Cached

The paper introduces ClaimReceipt, a claim-relative receipt specification and verifier for verifying evidence sufficiency and coverage in agent evaluations, validated through experiments on historical records and prospective audits.

0 favorites 0 likes
#agent-evaluation

@rohanpaul_ai: Current agent benchmarks may be ending before the real failures start. FM-Bench shows that a model can look strong afte…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

FM-Bench is a benchmark for evaluating long-horizon AI agents, revealing that short-term performance does not guarantee long-term success in simulated management tasks over 20 years.

0 favorites 0 likes
#agent-evaluation

EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction

Hugging Face Daily Papers ↗ · 2026-09-02 Cached

EarlyEval introduces a lightweight framework that predicts LLM agent outcomes from intermediate behavior to reduce evaluation costs by halting runs early with high accuracy. It eliminates 13%-26% of agent steps and up to 44% input tokens across benchmarks like SWE-bench Verified, TerminalBench, and Toolathlon.

0 favorites 0 likes
#agent-evaluation

@ApexAIHighlight: Most AI benchmarks test whether a model can give the right answer. @Accio_official is testing something far harder: Can…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

CommerceAgentBench is a new benchmark with 107 real-world e-commerce tasks designed to test whether AI agents can actually complete jobs, moving beyond traditional answer-based AI benchmarks.

0 favorites 0 likes
#agent-evaluation

Can AI Improve Itself? RSI Might Be the Answer [R]

Reddit r/MachineLearning ↗ · 2026-08-27

Introduces HarnessOpt-Bench to measure recursive self-improvement in AI, evaluating 5 frontier models on 4 tasks and finding that model choice has a greater impact than coding harness choice.

0 favorites 0 likes
#agent-evaluation

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

Hugging Face Daily Papers ↗ · 2026-08-27 Cached

UrbanGround evaluates whether multimodal language model agents can sustain reliable navigation and spatial reasoning in a realistic 3D city replica, revealing that local perceptual skills fail to compose into extended goal-directed behavior.

0 favorites 0 likes
#agent-evaluation

@rohanpaul_ai: Most agent benchmarks end after one task, but running a store doesn't, and that's where these agents come apart. Mercha…

X AI KOLs Timeline ↗ · 2026-08-23 Cached

MerchantBench is a benchmark that assesses AI agents by having them manage a simulated online store for a year, revealing challenges with sustained performance and continuous action.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback