GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Artificial Analysis Launches Harvey LAB-AA: 120 Private Legal Tasks Test AI Agents on Real Deliverables

Artificial Analysis has published the Harvey LAB-AA Benchmark Leaderboard, an implementation of Harvey’s Legal Agent Benchmark (LAB) running AI agents against a private dataset of 120 real legal tasks spanning 24 practice areas. The benchmark was made available on artificialanalysis.ai as of this week.

Harvey LAB is not a question-answering test. Each task gives an agent a partner-level instruction and a folder of case documents in a sandboxed environment. The agent must read the materials, work across them, and produce a legal deliverable: a memo, a disclosure schedule, a deposition summary, or equivalent output. A single LLM rubric judge scores each deliverable criterion by criterion against a task-specific rubric.

The scoring methodology reports two numbers: the all-pass rate — the share of tasks where the agent satisfies every criterion, with no partial credit — and the criterion pass rate, the share of individual rubric criteria satisfied across all tasks. The gap between these two numbers is diagnostic. A model that clears 85% of individual criteria but only completes 30% of tasks entirely has a different failure mode than one that completes 60% of tasks cleanly. Law firms care about the finished work product; the criterion pass rate explains where agents break down.

Why This Benchmark Is Different

Most public legal AI benchmarks test on legal questions with retrievable answers. Harvey LAB tests on legal work with no single correct answer — the output is a drafted document, and the rubric reflects what a supervising partner would check. That distinction matters because the legal AI market has largely been selling on retrieval accuracy and hallucination rate. Harvey LAB shifts the evaluation surface to drafting quality, structural completeness, and substantive legal reasoning across a full task.

The 24 practice areas in Harvey’s dataset include transactional and litigation work, making it a broader coverage test than the benchmarks most vendors run internally. The 120-task size is smaller than public benchmarks like SWE-bench Verified (500 tasks) but the tasks are harder, longer, and more diverse — each requires the agent to read multiple documents and produce a polished deliverable, not a patch to a codebase.

Artificial Analysis runs the benchmark using its Stirrup agent harness, the same infrastructure used for AA-Briefcase and the AA Legal Index. That consistency means Harvey LAB-AA results can be compared against performance on the AA Legal Index without controlling for different evaluation frameworks.

Context

The Harvey LAB-AA launch follows two closely related moves by Artificial Analysis: the AA-Briefcase benchmark (91 private knowledge work tasks, published in May 2026) and the AA Legal Index overhaul that made agentic execution 25% of the legal AI composite score. Harvey LAB-AA represents the most task-authentic legal evaluation of the three, because the tasks come from Harvey’s actual production dataset — not synthetic scenarios designed for a benchmark.

Harvey is the legal AI company with the deepest document-based evaluation history, having graded its own system’s outputs against partner-drafted rubrics since early 2024. Licensing that task dataset to Artificial Analysis for an open leaderboard is a notable shift toward external verifiability: Harvey’s claims about its own model performance can now be checked by third parties running other systems against the same tasks.

The leaderboard is live at artificialanalysis.ai/evaluations/harvey-lab-aa. Model scores and specific rankings were not yet visible at publication time; Artificial Analysis typically populates leaderboards iteratively as vendor submissions clear validation.