@SoHarshhh: Really happy to share that “ToolFailBench” got accepted at two ICML 2026 workshops, FAGEN and AIWILD. Most benchmarks e…
Summary
ToolFailBench, a diagnostic benchmark for tool-using agents, has been accepted at two ICML 2026 workshops, FAGEN and AIWILD.
View Cached Full Text
Cached at: 06/01/26, 11:20 AM
Really happy to share that “ToolFailBench” got accepted at two ICML 2026 workshops, FAGEN and AIWILD.
Most benchmarks evaluate tool-using agents with a single aggregate success rate, but that number can’t explain why a model actually fails. ToolFailBench is a diagnostic https://t.co/UCKA2H29Aw
Similar Articles
@Lyubh22: Coding benchmarks are saturating. AI4Research is the next frontier. Thrilled to see our MLS-Bench (https://mls-bench.co…
Announcing MLS-Bench, the first AI4Research benchmark to gain broad community adoption, testing AI agents on 140 executable tasks across 12 domains to propose modular ML improvements. The post includes leaderboard scores for models like Claude Opus 4.6 and GPT-5.4.
Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability
Introduces ToolBench-X, a benchmark for evaluating large language model agents under various tool-environment reliability hazards, revealing a substantial gap in performance compared to clean environments.
ToolGate: An Executable Acceptance Pipeline for Tool-Dependent Scientific Benchmark Construction
ToolGate is an executable pipeline for constructing scientific benchmarks that validates generated tasks through executable scripts, random no-tool screening, and tool-using agents, reducing manual labor and enhancing audibility.
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
This paper introduces EvalDetectBench, an open benchmark and pipeline for measuring evaluation awareness in frontier language models, addressing biases in existing methods to improve AI safety assessments.
TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agents
TOBench is a new benchmark for evaluating AI agents on real-world, task-oriented tool use with multimodal inputs and closed-loop verification. Experiments show top models like Qwen 3.5 Plus achieve only 41% success, far below the 94% human benchmark, highlighting a significant gap.