GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Artificial Analysis Endpoint Accuracy Index: Same Open-Weight Model, Different Scores by Provider

Artificial Analysis has launched the Endpoint Accuracy Index, a benchmark designed to answer a question inference buyers rarely get to ask directly: when a provider runs the same open-weight model, how much of the original accuracy do they actually deliver?

The answer, as it turns out, varies. Providers quantize weights, write custom kernels, and tune their inference stacks to hit target latency and cost. Sometimes they ship bugs. The Endpoint Accuracy Index puts a number on all of that.

How It Works

Each serverless endpoint is tested against a self-hosted reference deployment of the official weights. A score of 100% means the endpoint’s results fall within the 95% confidence interval of the reference. Below 100% means accuracy loss from provider-side changes.

Three evaluation areas, equally weighted:

  • Tool calling — BFCL-500, a 500-question subset of the Berkeley Function Calling Leaderboard, run with 3 repeats. Overweights the hardest categories: multi-turn, parallel calls, irrelevance detection.
  • Scientific reasoning — HLE-250, a 250-question subset of Humanity’s Last Exam, 10 repeats per question.
  • Long context recall — AA-LCR-25, 25 questions, 10 repeats. Tests the long-context fidelity that matters most for agentic retrieval.

The three-area structure was designed to catch the failure modes that matter to practitioners: a provider can pass on tool calling while silently degrading at long context, or vice versa.

Initial Coverage

Coverage launches with three models:

  • GLM-5.2 — Z.ai’s open-weight flagship, now one of the top open-source performers on the AA Intelligence Index at 51
  • gpt-oss-120b — OpenAI’s 120B open-source model
  • DeepSeek V4 Pro — DeepSeek’s open-weight frontier entry at 80.6% SWE-bench Verified

Kimi K3 accuracy coverage launches next. Kimi K3 currently sits at AA Intelligence Index 57 and is among the most widely served open-weight models on third-party inference APIs, which makes provider-side accuracy variation especially consequential.

Why It Matters Now

The open-weight inference market has fragmented quickly. Most developers pick providers on price and speed — the two metrics that have historically been easiest to measure. Accuracy is harder to compare without a standardized test against a known reference.

That creates a systematic selection problem. A provider running a quantized model at 90% reference accuracy and 2x throughput looks better on every visible metric until you notice the quality regression in production. The Endpoint Accuracy Index makes that regression visible before deployment.

The framing is precise: Artificial Analysis benchmarks the endpoint, not the model. The underlying weights are treated as ground truth. What varies is the provider’s implementation.

What Comes Next

AA describes Kimi K3 as coming soon for this index. Given that K3 is already benchmarked on the Intelligence Index, the extension is likely already in progress. The methodology — three areas, 10 repeats on the hardest questions — is designed to be consistent across any open-weight model with a canonical reference deployment, so expansion to additional models should be incremental.

For developers choosing between providers on a high-stakes model, this index is the first structured way to answer the question the speed and price charts can’t.