GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

DeepSeek V4.1 Flash Tops LiveBench Agentic Coding at 77.3% — More Than 10 Points Ahead of Fable 5.1

DeepSeek V4.1 Flash has entered the September 2026 LiveBench leaderboard with the highest Agentic Coding score of any measured model: 77.3. The closest competitor is Claude Fable 5.1 at 66.1 — an 11.2-point gap. V4.1 Flash also sits fifth overall at 81.1, behind Fable 5.1 (83.4), Fable 5 (83.0), GPT-6 Astra (82.2), and Muse Spark 1.3 (81.6), but leads the specific column that measures autonomous software task completion.

The Agentic Coding Column

LiveBench’s Agentic Coding category tests models on tasks that require multi-step tool use and code execution: navigating repositories, implementing features with dependencies, running tests and interpreting failures. It is one of the most demanding LiveBench columns to lead, because it captures the failure modes that matter most for production agent use — planning, error recovery, and state management across turns.

The September 2026 Agentic Coding ranking:

ModelAgentic CodingOverallCost/Task
DeepSeek V4.1 Flash (max)77.381.1$0.029
Claude Fable 5.1 (max)66.183.4$1.21
Muse Spark 1.3 (xhigh)64.181.6$0.22
GPT-6 Astra (max)57.382.2$0.74
Claude Fable 5 (max)62.283.0$1.44

V4.1 Flash’s 77.3 represents a 17.3-point lead over the model at the bottom of this comparison. In Agentic Coding, higher overall intelligence scores do not uniformly translate to better performance: Fable 5.1 leads overall but sits 11.2 points behind V4.1 Flash on this specific column.

Cost Differential Is Extreme

At $0.029 per successful task on LiveBench, V4.1 Flash costs approximately 42 times less than Claude Fable 5.1 ($1.21) and 25 times less than Muse Spark 1.3 ($0.22). For agentic workloads running at scale — CI pipelines, automated code review, continuous test generation — the cost difference is operationally significant independent of the benchmark gap.

Open Weights

DeepSeek V4.1 Flash is an open-weight model. The agentic coding lead, combined with open availability and sub-three-cent task cost, positions it as a credible base for self-hosted agent deployments where Fable 5.1 or GPT-6 Astra API costs would be prohibitive.

DeepSeek V4.1 Flash Max was added to Arena’s Code Arena WebDev leaderboard on September 10, providing a separate battle-based ranking context that will update as user preferences accumulate.

Caveats

LiveBench’s Agentic Coding column has not historically correlated perfectly with SWE-bench Verified results. Models that top LiveBench Agentic Coding have sometimes posted lower SWE-bench scores than models with worse LiveBench rankings. DeepSeek V4.1 Flash does not yet have a published SWE-bench Verified result; its position on this site’s ticker composite is null pending primary benchmark data.

The 11-point gap over Fable 5.1 is large enough to survive most measurement variation. Whether it reflects a genuine capability advantage or a strong match between V4.1 Flash’s architecture and LiveBench’s specific task distribution is a question that SWE-bench Verified data — when it arrives — will partially answer.