harness-optimization

Tag

Cards List
#harness-optimization

A 35B model beat a 120B one on my coding agent, 95% vs 53%. Build your own benchmark.

Reddit r/AI_Agents ↗ · 2d ago

A developer built a custom benchmark for coding AI agents and found that a 35B-parameter model outperformed a 120B-parameter one when the harness was optimized, highlighting the importance of tailored evaluation over generic specifications.

0 favorites 0 likes
#harness-optimization

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

EdgeGen is a synthetic task generation framework that creates database-grounded edge-case tasks to improve tool-calling agents through fine-tuning and harness optimization, demonstrating consistent performance improvements.

0 favorites 0 likes
#harness-optimization

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Hugging Face Daily Papers ↗ · 2026-09-21 Cached

This paper introduces Regularized Recursive Self-Improvement (RRSI) for AI agent harnesses, which applies regularization to prevent overfitting during recursive evolution, demonstrating performance gains on multiple benchmarks.

0 favorites 0 likes
#harness-optimization

Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses

arXiv cs.AI ↗ · 2026-09-10 Cached

This paper studies agent harness optimization to improve LLM tool agents without retraining, focusing on prompts and tool-boundary middleware, and introduces a protocol and the PRISM optimizer for measurable gains.

0 favorites 0 likes
#harness-optimization

Where Does Harness-Optimization Value Live? Localized Gains and the Budget-Splitting Trap in Self-Evolving LLM Agents

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper introduces HarnessEvo to decompose LLM agent harnesses into separately-evolvable slots, revealing that optimization value is localized in specific components like reflection/control, and that uniform budget-splitting is sub-optimal, advocating for targeted budget concentration.

0 favorites 0 likes
#harness-optimization

WHALE: A Simple Recipe for Joint Harness-Weight Optimization

arXiv cs.LG ↗ · 2026-09-02 Cached

The paper proposes WHALE, an alternating optimization method for jointly training model weights and harness code in AI agents, achieving significant performance improvements across search QA, math reasoning, and chess puzzles.

0 favorites 0 likes
#harness-optimization

Can an AI make other AIs better? We benchmarked 5 frontier LLMs at rewriting other agents' harnesses, scored on a test set they never see (HarnessOpt-Bench, arXiv + MIT code)

Reddit r/artificial ↗ · 2026-08-27

The article introduces HarnessOpt-Bench, a benchmark for measuring how LLMs can improve other AI agents' harnesses, and presents findings from 5 frontier models, showing that model choice has a greater impact than harness choice.

0 favorites 0 likes
#harness-optimization

Automatic Harness Optimization (GitHub Repo)

TLDR AI ↗ · 2026-08-27 Cached

AutoSaddler is a Microsoft tool that automatically optimizes LLM-agent harnesses by diagnosing execution traces and applying structured updates to prompts, tools, and middleware, demonstrating significant benchmark improvements.

0 favorites 0 likes
#harness-optimization

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

Hugging Face Daily Papers ↗ · 2026-08-24 Cached

AutoSaddler is an automatic harness optimization framework that improves LLM agent performance on long-horizon tasks by iteratively updating harnesses using failure signals, achieving substantial gains on benchmarks like GAIA2 and SWE-Bench.

0 favorites 0 likes
#harness-optimization

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

arXiv cs.CL ↗ · 2026-08-21 Cached

Task-CoEvolve is a novel approach for efficient LLM agent harness optimization that adaptively selects validation tasks to reduce evaluation costs while maintaining performance, achieving an 80% reduction in evaluations on benchmarks.

0 favorites 0 likes
#harness-optimization

SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

arXiv cs.AI ↗ · 2026-08-12 Cached

Introduces SBCO, a self-supervised verifier-grounded harness optimizer for planning agents that improves agent outputs via approximate block coordinate ascent, matching or exceeding self-modifying baselines with far less compute.

0 favorites 0 likes
#harness-optimization

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

HarnessOpt-Bench is a benchmark for evaluating LLMs' ability to optimize the harness—the prompts, tools, control flow, memory, and orchestration code—around a target agent, using a fixed evaluation budget. Experiments with five frontier LLMs show that optimizer models separate more than the coding harnesses they act through, with substantial room for improvement.

0 favorites 0 likes
#harness-optimization

@joelniklaus: New blog post on harness optimization. We hit Sonnet 4.6 performance with a 7x cost improvement. Fable 5 was the first …

X AI KOLs Following ↗ · 2026-07-01 Cached

A blog post describes how automatic harness optimization enabled DeepSeek V4 Pro to achieve Sonnet 4.6 performance on the Legal Agent Benchmark at one-seventh the cost.

0 favorites 0 likes
#harness-optimization

Self-Harness: Harnesses That Improve Themselves

Hacker News Top ↗ · 2026-06-22 Cached

Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.

0 favorites 0 likes
#harness-optimization

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

Hugging Face Daily Papers ↗ · 2026-06-04 Cached

Retrospective Harness Optimization (RHO) is a self-supervised method that improves LLM agent performance using only past trajectories, achieving a 78% pass rate on SWE-Bench Pro without external grading.

0 favorites 0 likes
← Back to home

Submit Feedback