harveyai/harvey-labs
Summary
Harvey AI released Harvey LAB, an open-source benchmark for evaluating LLM agents on realistic legal work, featuring a dataset of 1,671 tasks across 24+ practice areas and an execution harness.
View Cached Full Text
Cached at: 08/09/26, 01:54 PM
harveyai/harvey-labs
Source: https://github.com/harveyai/harvey-labs
Legal Agent Benchmark (LAB): An open-source benchmark for evaluating agents on real legal work.
Harvey LAB is an open-source project aimed at benchmarking LLM agents’ abilities to perform legal work in realistic environments.
LAB consists of two parts: a dataset of tasks containing agent instructions, documents, and rubrics as well as an execution harness for running and evaluating agents against those tasks.
LAB is an ongoing project and we expect to consistently add to and refine the task set and execution harness.
Read the announcement post: Introducing Harvey’s Legal Agent Benchmark
Getting Started
Start with the full walkthrough in docs/tutorial.md — it takes one realistic M&A data-room assignment end to end: setup, task inspection, agent run, scoring, report review, and comparison dashboards.
Additional Documentation
| Guide | Description |
|---|---|
| Architecture | Task model, harness, tools, adapters, reports, and sweeps |
| Evaluation Methodology | All-pass rubric scoring and LLM judge behavior |
| Contributing | Add tasks, model adapters, evaluation improvements, and docs |
Citation
If you use Harvey LAB in your research, please cite it as:
@misc{harveylab2026,
title = {Harvey LAB: The Legal Agent Benchmark},
author = {{Harvey AI}},
year = {2026},
version = {v1.0},
url = {https://github.com/harveyai/harvey-labs/tree/v1.0},
note = {Announcement: \url{https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark}}
}
Similar Articles
Introducing Harvey's Legal Agent Benchmark (12 minute read)
Harvey releases the Legal Agent Benchmark (LAB), an open-source tool designed to evaluate AI agents on long-horizon legal tasks to help law firms measure ROI and track progress.
@gabepereyra: Harvey partnered with @appliedcompute to train a legal agent. We optimized each part of the agent stack, including the …
Harvey partnered with Applied Compute to train a legal agent, optimizing the agent stack and post-training the GLM-5.1 model using reward signals from their Legal Agent Benchmark.
Customizing models for legal professionals
Harvey, a generative AI platform for legal professionals, partnered with OpenAI to create a custom-trained case law model that reduces hallucinations and improves reasoning for complex legal tasks like document drafting and contract analysis. The custom model, trained on 10 billion tokens of U.S. case law, achieved 97% lawyer preference over standard foundation models.
@harvey: Inference-time model routing based on legal practice area dramatically improves agent performance. So do other types of…
Harvey is experimenting with inference-time model routing and blended intelligence techniques to improve agent performance in legal tasks, highlighting both promise and risks.
@FinanceYF5: The lesson from Harvey is that in the AI era, companies are no longer selling software, but intelligence. 1/ Not legal software. Harvey CEO Winston Weinberg said: "Every company will eventually sell intelligence." On the surface, Harvey is a legal AI, but what it really sells is a lawyer's judgment, research, review...
Harvey CEO Winston Weinberg proposes that in the AI era, companies no longer sell software, but intelligence. Using Harvey legal AI as an example, he explains that what it truly sells is a lawyer's judgment and research capabilities.