harvey-labs

An open-source benchmark that measures how well AI agents handle realistic legal work. It provides tasks and evaluation for improving legal agents.

Share on XLicense: MIT

Overview

Harvey LAB, short for Legal Agent Benchmark, is an open-source project from Harvey AI for measuring how well LLM agents handle realistic legal work. It has two parts: a dataset of tasks with instructions, documents and rubrics, and an execution harness that runs agents on those tasks and scores them. A tutorial walks through one M&A data-room assignment from setup to comparison dashboards.

Key features

  • Task dataset with instructions, documents and rubrics
  • Execution harness for running and evaluating agents
  • All-pass rubric scoring with an LLM judge
  • Reports and comparison dashboards
  • Adapters for adding models and tasks

Best for

Researchers and teams building or comparing AI agents for legal work who need a shared test set. The README says the task set and harness are still being extended.

Upstream
harveyai/harvey-labs
Fork on GitHub
Guo-astro/harvey-labs
Upstream stars
1.4k
Category
AI agents and LLM tools
License
MIT
Forked
2026-09-28
Sync status
In syncLast synced 2026-09-29