GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Cohere Releases Command A+: Open Weights, 86% Non-Hallucination, First Model in 14 Months

Cohere has released Command A+, an open-weights language model scoring 37 on the Artificial Analysis Intelligence Index. The release puts Command A+ in the same tier as Claude 4.5 Haiku, NVIDIA Nemotron 3 Super, and Gemini 3.1 Flash-Lite — capable mid-range models positioned for enterprise deployment rather than frontier reasoning.

It is the company’s first significant model release in over 14 months, following the original Command A.

Where It Leads

Command A+‘s clearest differentiation is on knowledge reliability. It scores 86% on AA-Omniscience Non-Hallucination, approximately three percentage points ahead of the next-best model at this intelligence tier. The trade-off is intentional: its AA-Omniscience Accuracy is 9%, which means the model tends to decline when uncertain rather than confabulate. For legal, financial, and compliance deployments — Cohere’s core enterprise markets — that calibration is the right one.

On output speed, Command A+ produces approximately 281 tokens per second on Cohere’s API, faster than Claude 4.5 Haiku, GPT-5.4 nano, and Grok 4.3, while still trailing Gemini 3.1 Flash-Lite Preview at 304 tokens per second.

Where It Trails

Command A+ has measurable gaps on the hardest benchmarks:

BenchmarkCommand A+Tier Context
HLE~11%Below Haiku class average
GPQA Diamond76%Below Haiku class average
Terminal-Bench Hard~25%Well below coding-optimised models
SciCode~38%Below scientific model class
MMMU-Pro (vision)63%Between Haiku (59%) and GPT-5.4 nano (65%)

The Terminal-Bench Hard gap is the most commercially relevant. At 25%, Command A+ is not a coding agent deployment option, even for tasks within its intelligence tier. That limit places it squarely in content, retrieval, and document-processing workflows rather than developer tools.

The Open Weights Angle

Command A+ is released with open weights — a significant shift for Cohere, which has historically positioned its models as API-only, enterprise-licensed products. Opening the weights is a competitive response to an environment where Qwen3.6-35B, Mistral Medium 3.5, and Llama 4 Scout have raised the baseline expectations for what open-weights models should offer.

For Cohere’s enterprise customers, the open-weights release enables self-hosted deployment, which has become a hard requirement for regulated-industry customers in financial services and healthcare across Europe — markets that Cohere targeted directly through its acquisition of Aleph Alpha earlier this year.

Context: Cohere After Aleph Alpha

Cohere’s $20 billion acquisition of Aleph Alpha gave it direct access to European government and financial services contracts, a dedicated on-premises deployment stack, and 30 researchers specialised in domain-specific AI. Command A+ is the first new Cohere model to enter those markets as a self-hosted option. Whether the combined entity can hold that enterprise position against DeepSeek V4 — which offers a substantially higher intelligence index score under an MIT licence — will determine how durable this release is commercially.