Tag
A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.
EdgeGen is a synthetic task generation framework that creates database-grounded edge-case tasks to improve tool-calling agents through fine-tuning and harness optimization, demonstrating consistent performance improvements.
This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.
This paper studies agent harness optimization to improve LLM tool agents without retraining, focusing on prompts and tool-boundary middleware, and introduces a protocol and the PRISM optimizer for measurable gains.
This paper introduces HarnessEvo to decompose LLM agent harnesses into separately-evolvable slots, revealing that optimization value is localized in specific components like reflection/control, and that uniform budget-splitting is sub-optimal, advocating for targeted budget concentration.
The paper proposes WHALE, an alternating optimization method for jointly training model weights and harness code in AI agents, achieving significant performance improvements across search QA, math reasoning, and chess puzzles.
The article introduces HarnessOpt-Bench, a benchmark for measuring how LLMs can improve other AI agents' harnesses, and presents findings from 5 frontier models, showing that model choice has a greater impact than harness choice.
AutoSaddler is a Microsoft tool that automatically optimizes LLM-agent harnesses by diagnosing execution traces and applying structured updates to prompts, tools, and middleware, demonstrating significant benchmark improvements.
AutoSaddler is an automatic harness optimization framework that improves LLM agent performance on long-horizon tasks by iteratively updating harnesses using failure signals, achieving substantial gains on benchmarks like GAIA2 and SWE-Bench.
Task-CoEvolve is a novel approach for efficient LLM agent harness optimization that adaptively selects validation tasks to reduce evaluation costs while maintaining performance, achieving an 80% reduction in evaluations on benchmarks.
Introduces SBCO, a self-supervised verifier-grounded harness optimizer for planning agents that improves agent outputs via approximate block coordinate ascent, matching or exceeding self-modifying baselines with far less compute.
HarnessOpt-Bench is a benchmark for evaluating LLMs' ability to optimize the harness—the prompts, tools, control flow, memory, and orchestration code—around a target agent, using a fixed evaluation budget. Experiments with five frontier LLMs show that optimizer models separate more than the coding harnesses they act through, with substantial room for improvement.
A blog post describes how automatic harness optimization enabled DeepSeek V4 Pro to achieve Sonnet 4.6 performance on the Legal Agent Benchmark at one-seventh the cost.
Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.
Retrospective Harness Optimization (RHO) is a self-supervised method that improves LLM agent performance using only past trajectories, achieving a 78% pass rate on SWE-Bench Pro without external grading.