GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Artificial Analysis Legal Index Makes Agentic Execution 25% of Legal AI Score

Artificial Analysis has launched a Legal Index for ranking AI models on legal work, and the weights say more than the headline category does. Legal knowledge carries 30% of the score. Agentic execution carries 25%.

That makes the index less a bar exam proxy than a workflow test. The remaining weights are non-hallucination at 15%, long-context reading at 15%, and reasoning at 15%. In practical terms, a model cannot win the index by memorising statutes if it drops facts, mishandles long documents, or cannot execute a multi-step legal task.

The Weighting

CapabilityWeight
Legal knowledge30%
Agentic execution25%
Non-hallucination15%
Long-context reading15%
Reasoning15%

The underlying evaluations include AA-Omniscience Law Accuracy for legal knowledge, GDPval-AA v2 for agentic execution, AA-Omniscience Non-Hallucination, long-context retrieval, and HLE-style reasoning.

The important design choice is that legal AI is being treated as a composite production capability, not a narrow Q&A task. For lawyers and compliance teams, the product risk usually sits in the workflow: missed constraints, invented citations, weak document handling, or brittle execution across a chain of steps.

Why It Matters

Legal AI vendors have had an easy time showing demos where a model answers a single question or drafts a polished first pass. That does not tell buyers whether the same model can hold a record straight across a 200-page matter file, avoid unsupported claims, and produce usable work product without a human rebuilding the reasoning chain.

The Legal Index pushes the market toward that harder question. A 30% knowledge weighting still rewards domain fluency, but the other 70% tests whether the model behaves like software that can be trusted inside a professional workflow.

The timing also matters. Legal is one of the first verticals where frontier labs are building native workflows rather than just selling chat access. Claude has moved into Microsoft Word with tracked changes. Thomson Reuters is connecting CoCounsel to Claude through MCP. Specialist legal benchmarks are now catching up to the product surface.

The Catch

The index is useful because it combines independent model runs across several dimensions. It is also still an index. The weightings encode assumptions about legal work, and different buyers will care about different failure modes: litigation review, compliance monitoring, contract redlining, and regulatory advice do not fail in the same way.

Still, this is the right direction. Legal AI will not be decided by which model sounds most lawyerly. It will be decided by which model can execute under document pressure, cite without hallucinating, and keep its state intact while the work gets boring.