agent-benchmarks

Tag

Cards List
#agent-benchmarks

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

arXiv cs.AI · 3h ago Cached

This paper extends AssetOpsBench with a Smart Grid Transformer asset class and introduces ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation that achieves an 8x runtime improvement while maintaining scenario quality.

0 favorites 0 likes
#agent-benchmarks

An open-weight, MIT trillion-param model (Ant's Ring-2.6) reportedly matches the closed frontier on reasoning + agent benchmarks. Does "open" catching up actually change the trajectory?

Reddit r/singularity · 2026-07-14

The article reports that Ant's Ring-2.6, an open-weight trillion-parameter model under MIT license, reportedly matches closed frontier models on reasoning and agent benchmarks, raising questions about the impact of open models catching up.

0 favorites 0 likes
#agent-benchmarks

Dissecting model behavior through agent trajectories

arXiv cs.AI · 2026-06-17 Cached

This paper introduces the Simple Strands Agent (SSA), a minimal harness designed to reduce the intent-execution gap between AI models and their agentic behavior, and analyzes 138k trajectories across various model families to reveal fine-grained behavioral differences.

0 favorites 0 likes
#agent-benchmarks

Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

Hugging Face Daily Papers · 2026-06-08 Cached

Researchers propose an adversarial hacker-fixer loop using LLM agents to automatically patch brittle verifiers in agent benchmarks, reducing attack success rates from 62% to 0% on KernelBench and demonstrating that weaker defenders can neutralize much stronger attackers.

0 favorites 0 likes
#agent-benchmarks

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

arXiv cs.AI · 2026-05-27 Cached

Anchor is a task-generation pipeline that addresses artifact drift in AI agent benchmarks by jointly producing instructions, environments, solutions, and verifiers from a single constraint optimization specification, yielding consistent and auditable evaluation tasks for enterprise workflows. The paper introduces ERP-Bench, a benchmark of 300 long-horizon tasks in a production ERP system, showing that frontier models satisfy explicit constraints in 26.1% of trials but reach optimal solutions in only 17.4%.

0 favorites 0 likes
#agent-benchmarks

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Hugging Face Daily Papers · 2026-05-27 Cached

TASTE is an automated method for generating challenging agent benchmarks with broader tool-use coverage by evolving tool sequences through adaptive contrastive n-gram modeling and iterative difficulty refinement. The resulting τ^c-Bench reveals that models nearly saturating existing benchmarks suffer severe performance drops, indicating saturation rather than robust skill.

0 favorites 0 likes
#agent-benchmarks

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

arXiv cs.CL · 2026-05-22 Cached

SynAE is a framework for evaluating the quality of synthetic data used in tool-calling agent evaluations, assessing validity, fidelity, and diversity across multiple axes. It addresses challenges of insufficient or sensitive real data by providing metrics to guide synthetic data generation.

0 favorites 0 likes
#agent-benchmarks

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Hugging Face Daily Papers · 2026-05-20 Cached

This paper introduces temporal semantic caching and MCP workflow optimizations for agentic plan-execute pipelines, achieving up to 30.6x speedup on cache hits and 1.67x overall speedup on the AssetOpsBench industrial benchmark.

0 favorites 0 likes
#agent-benchmarks

Interactive Evaluation Requires a Design Science

Hugging Face Daily Papers · 2026-05-18 Cached

This position paper argues that interactive AI evaluation should be treated as a design science paradigm, proposing a two-axis taxonomy and reporting standards for assessing dynamic system behavior through trajectories.

0 favorites 0 likes
← Back to home

Submit Feedback