VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Hugging Face Daily Papers Papers

Summary

VeriHarness turns an LLM generator into an agentic verifier—using a workspace, evidence tools, reusable verification skills, a disagreement resolver, and a consensus challenger—to select reliable outputs for long-horizon tasks without reference answers. It achieves top selection scores across five benchmarks, gains of 6.2–6.4 points over single rollouts with Gemini 3.5 Flash and Claude Opus 4.8, and releases ~26,000 rollouts costing over $100,000.

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
Original Article
View Cached Full Text

Cached at: 10/05/26, 08:46 AM

Paper page - VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

Source: https://huggingface.co/papers/2610.00972

Abstract

AsLLMagentsundertakeincreasinglycomplex,long-horizontasks,verifyingtheiroutputsbecomesincreasinglychallenging.Westudyhowverificationcapabilitycanbestrengthenedwithafixedbasemodel,withoutaccesstoreferenceanswersorgradingrubricsattesttime.Repeatedsamplingyieldsmultiplerolloutsthatcancontaincomplementarycorrectclaims,butweneedareliableverificationmechanismtodeterminewhichclaimstotrust.Wefirstfindthatdisagreementoftenexposescorrectalternatives,whileconsensuscanconcealerrors.TheseobservationsmotivateVeriHarness,whichturnstheunderlyingLLMageneratorusesintoanagenticverifierbygivingitaworkspace,evidencetools,andreusableverificationskills.Adisagreementresolvercheckscompetingclaimsagainstenvironmentalevidence,whileaconsensuschallengertestssharedclaimsandsearchesforomittedrequirements.Theirfindingsguidetheselectionandrevisionofthefinalartifact.Acrossfivelong-horizonworkspacebenchmarksandtwofrontiermodels,VeriHarnessachievesthehighestselectionscoresamongtheevaluatedbaselines.Evidence-backedrevisionfurtherimprovesaverageperformance,bringinggainsoverasinglerolloutto6.2pointswithGemini3.5Flashand6.4pointswithClaudeOpus4.8.Wefurthershowthatverificationskillscanself-improvefromfailurefeedback,demonstratingVeriHarnessasanovelandcriticalapproachforscalinglong-horizonagenticverification.Wereleasethefullpoolofapproximately26,000rolloutsacrossallfivebenchmarksandbothmodels,producedatacostofover$100,000,tosupportfutureresearchonagenticverification.

View arXiv pageView PDFProject pageGitHubAdd to collection

Get this paper in your agent:

hf papers read 2610\.00972

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.00972 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.00972 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.00972 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

arXiv cs.AI

The paper introduces VERSE, a verified self-evolving optimizer that improves both the executor agent's harness and its own prompts, tools, and procedures using execution-based verification, achieving 42.3% and 37.7% accuracy on held-out SWE-rebench tasks versus 39.2% and 29.3% for the strongest baselines.

AgentV-RL: Scaling Reward Modeling with Agentic Verifier

arXiv cs.CL

AgentV-RL introduces an Agentic Verifier framework that enhances reward modeling through bidirectional verification with forward and backward agents augmented with tools, achieving 25.2% improvement over state-of-the-art ORMs. The approach addresses error propagation and grounding issues in verifiers for complex reasoning tasks through multi-turn deliberative processes combined with reinforcement learning.

Building an Advanced Agentic Harness

Hacker News Top

A technical blog post that walks through building a production-grade agentic harness around a basic LLM loop, covering typed tools, plan DAGs, tiered memory, verification hierarchies, budgets, and tracing.

best of the best agentic harnesses do this…

Reddit r/AI_Agents

The author shares insights on building effective agent harnesses: the best ones minimize LLM reliance for trivial tasks and reserve LLMs for complex reasoning, distinguishing genuine harnesses from simple wrappers.

Agentic Proving for Program Verification

arXiv cs.AI

This paper evaluates Claude Code in an agentic proving framework on the Clever benchmark for program verification, achieving over 98% success in specification generation and end-to-end verification, revealing that existing benchmarks may be insufficient for evaluating modern agentic provers.