GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Stanford Blind Study: AI Wins 75% of Law Professor Comparisons — and Gets Flagged as Harmful 3.4x Less Often

A blind study from Stanford Law School published this week found that law professors overwhelmingly prefer AI-generated answers over responses from fellow instructors when evaluating student support for contract law.

The numbers are stark: AI won 75% of 2,885 head-to-head comparisons, judged by the same 16 law professors who also wrote the competing human answers.

Study Design

Professor Julian Nyarko, who co-chairs Stanford’s Law AI Initiative, designed the study to avoid the usual failure mode of AI benchmarks in professional domains: testing against known fact-lookup rather than defensible argument.

Contract law is specifically an area where the correct answer is often not a fact — it is a reasoned position built from rules, exceptions, and competing interpretations. The professors wrote 40 representative questions a student might bring to office hours, wrote their own answers, then blindly evaluated anonymized responses.

The AI systems tested included commercial tutoring systems and Google’s NotebookLM. The professors did not know which responses came from AI and which from peers.

The Harm Gap

The most significant finding is not the win rate — it is the harm rate.

When professors flagged a response as pedagogically harmful (misleading, incorrect, or likely to cause a student to misunderstand the law), they flagged AI responses 3.5% of the time. They flagged peer professor responses 12% of the time.

That is a 3.4x difference. Professors writing under time pressure, without structured review, produced answers that were three times more likely to be flagged as harmful by their own colleagues than answers from an AI system.

Even when AI responses were affected by context window limitations and incomplete reasoning, professors still frequently preferred them to human alternatives.

What This Means

Legal education has been the slow-moving exception to AI adoption in professional services. The resistance has two components: concern about accuracy, and concern about whether AI can handle judgment-based reasoning rather than fact retrieval.

This study challenges both. On accuracy, AI responses were flagged as harmful less than a third as often as human ones. On reasoning, the win rate held even in contract law questions that require structured legal argument, not lookup.

The study’s context is legal education specifically — student tutoring for contract law courses — not legal practice, where regulatory compliance, liability, and privilege create a separate set of constraints. But the implications for legal education platforms, bar prep, and law school pedagogy are direct.

Nyarko’s lab recently launched liftlab, a Stanford initiative focused on AI’s role in the legal profession. The study is titled “Law Professors Prefer AI Over Peer Answers” and is published through Stanford Law School.