IBM Releases Granite 4.2 8B: Dense Reasoning at $0.10 Input per Million Tokens, 12-Language Support
IBM shipped Granite 4.2 8B on August 31, 2026 — a dense reasoning model built for production agentic pipelines that need multi-step logic without the overhead of a mixture-of-experts architecture.
Specs at a glance
- Context window: 131,072 tokens
- Pricing: $0.10/M input, $0.15/M output (via OpenRouter)
- Inference modes: Full reasoning, low-effort, and non-thinking (selectable per request)
- Languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, Chinese
What it’s built for
Granite 4.2 8B is aimed at the same workload tier as its predecessor — reasoning-heavy agentic chains where operators want a small, fast model that can handle structured multi-step problems without routing to a 70B+ model. IBM positions it against math, code, multilingual dialogue, and autonomous workflow tasks.
The selectable inference modes give operators a meaningful cost lever: the non-thinking mode strips chain-of-thought overhead for latency-sensitive tasks, while full reasoning mode preserves depth for harder problems. That kind of per-request control has become standard practice in production deployments since OpenAI popularised effort tiers with the o-series.
Pricing in context
At $0.10/$0.15 per million tokens, Granite 4.2 8B is significantly cheaper than mid-tier frontier models. For comparison, IBM’s own prior release — Granite 4.1 8B — demonstrated that the 8B dense class could match IBM’s own 32B MoE on key benchmarks with lower deployment cost. The 4.2 generation extends that efficiency push.
The sub-penny-per-thousand-token input price puts Granite 4.2 squarely in the commodity reasoning tier alongside Google’s Gemini Flash Lite and other sub-$0.50/M options. The difference IBM is betting on: stronger multilingual coverage (12 languages, including Arabic and Czech) and explicit agentic workflow tuning, rather than raw benchmark scores.
Competitive position
IBM is not chasing the frontier benchmark tables with Granite 4.2. It’s competing in the enterprise deployment tier — regulated industries, on-prem or private-cloud inference, and cost-controlled agentic pipelines where model provenance and licensing clarity matter. The model is available through OpenRouter from a single provider, keeping the distribution footprint tight.
What’s notable is how much capability has compressed into the 8B weight class. Two years ago, multi-step mathematical reasoning and multilingual code generation at this level required a 70B model. That compression continues to shift what “budget model” actually means.