Artificial Analysis Legal Index Makes Agentic Execution 25% of Legal AI Score
Artificial Analysis has launched a Legal Index for ranking AI models on legal work, and the weights say more than the headline category does. Legal knowledge carries 30% of the score. Agentic execution carries 25%.
That makes the index less a bar exam proxy than a workflow test. The remaining weights are non-hallucination at 15%, long-context reading at 15%, and reasoning at 15%. In practical terms, a model cannot win the index by memorising statutes if it drops facts, mishandles long documents, or cannot execute a multi-step legal task.
The Weighting
| Capability | Weight |
|---|---|
| Legal knowledge | 30% |
| Agentic execution | 25% |
| Non-hallucination | 15% |
| Long-context reading | 15% |
| Reasoning | 15% |
The underlying evaluations include AA-Omniscience Law Accuracy for legal knowledge, GDPval-AA v2 for agentic execution, AA-Omniscience Non-Hallucination, long-context retrieval, and HLE-style reasoning.
The important design choice is that legal AI is being treated as a composite production capability, not a narrow Q&A task. For lawyers and compliance teams, the product risk usually sits in the workflow: missed constraints, invented citations, weak document handling, or brittle execution across a chain of steps.
Why It Matters
Legal AI vendors have had an easy time showing demos where a model answers a single question or drafts a polished first pass. That does not tell buyers whether the same model can hold a record straight across a 200-page matter file, avoid unsupported claims, and produce usable work product without a human rebuilding the reasoning chain.
The Legal Index pushes the market toward that harder question. A 30% knowledge weighting still rewards domain fluency, but the other 70% tests whether the model behaves like software that can be trusted inside a professional workflow.
The timing also matters. Legal is one of the first verticals where frontier labs are building native workflows rather than just selling chat access. Claude has moved into Microsoft Word with tracked changes. Thomson Reuters is connecting CoCounsel to Claude through MCP. Specialist legal benchmarks are now catching up to the product surface.
The Catch
The index is useful because it combines independent model runs across several dimensions. It is also still an index. The weightings encode assumptions about legal work, and different buyers will care about different failure modes: litigation review, compliance monitoring, contract redlining, and regulatory advice do not fail in the same way.
Still, this is the right direction. Legal AI will not be decided by which model sounds most lawyerly. It will be decided by which model can execute under document pressure, cite without hallucinating, and keep its state intact while the work gets boring.