harness-optimization

Tag

Cards List
#harness-optimization

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Hugging Face Daily Papers · 4d ago Cached

HarnessOpt-Bench is a benchmark for evaluating LLMs' ability to optimize the harness—the prompts, tools, control flow, memory, and orchestration code—around a target agent, using a fixed evaluation budget. Experiments with five frontier LLMs show that optimizer models separate more than the coding harnesses they act through, with substantial room for improvement.

0 favorites 0 likes
#harness-optimization

@joelniklaus: New blog post on harness optimization. We hit Sonnet 4.6 performance with a 7x cost improvement. Fable 5 was the first …

X AI KOLs Following · 2026-07-01 Cached

A blog post describes how automatic harness optimization enabled DeepSeek V4 Pro to achieve Sonnet 4.6 performance on the Legal Agent Benchmark at one-seventh the cost.

0 favorites 0 likes
#harness-optimization

Self-Harness: Harnesses That Improve Themselves

Hacker News Top · 2026-06-22 Cached

Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.

0 favorites 0 likes
#harness-optimization

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

Hugging Face Daily Papers · 2026-06-04 Cached

Retrospective Harness Optimization (RHO) is a self-supervised method that improves LLM agent performance using only past trajectories, achieving a 78% pass rate on SWE-Bench Pro without external grading.

0 favorites 0 likes
← Back to home

Submit Feedback