Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
Summary
FrontierHarness Eval benchmarks nine software engineering harnesses on a single model, showing that cost per successful task can vary by up to 17x depending on the harness used.
View Cached Full Text
Cached at: 09/02/26, 05:51 PM
Similar Articles
Same Model, Different Harness: Different Coding-Agent Results
This paper investigates how changing the harness configuration in a coding agent impacts performance on coding benchmarks when the model remains fixed. The study shows that a treatment harness, which shortens older tool results to manage context, improves task completion rates, especially under tight context constraints.
@LangChain: .@FactoryAI CTO @enoreyes ran the numbers. Same code review task, wildly different price depending on the harness Eno o…
LangChain shares analysis by FactoryAI CTO Eno Reyes on how the same code review task has wildly different prices depending on the harness used, arguing a good model-agnostic harness can improve any model.
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
HarnessOpt-Bench is a benchmark for evaluating LLMs' ability to optimize the harness—the prompts, tools, control flow, memory, and orchestration code—around a target agent, using a fixed evaluation budget. Experiments with five frontier LLMs show that optimizer models separate more than the coding harnesses they act through, with substantial room for improvement.
@rohit4verse: 2 months ago, I wrote "The Harness Is Everything" 1.3M views. Last week's Life-Harness paper: 116 of 126 model-environm…
The Life-Harness paper shows that patching the evaluation harness alone, without modifying the model, improved performance in 116 of 126 setups, achieving an 88.5% mean lift across 18 backbones.
Anyone interested in building a harness-only benchmark?
A proposal to create a community-driven benchmark specifically for evaluating LLM harnesses, with a leaderboard measuring harness performance on diverse real-world tasks.