State of AI Token Pricing: September 2026 Report
Based on empirical tracking of 30 production models across 9 provider ecosystems, this benchmark analyzes token price deflation, the universal adoption of KV caching, and the emerging bifurcation between high-deliberation reasoning and low-cost utility inference.
1. Macro Trend: 44% Annualized Deflation in Input Tokens
Over the past 12 months, the volume-weighted average price for standard input tokens has declined by 44.2%. Advances in FP8/FP4 quantization, speculative decoding, and dedicated inferencing silicon (NVIDIA Blackwell, Google TPU v5e/v6p) have dramatically lowered the cost per prefill token for cloud providers.
Lightweight models such as GPT-5.6 Luna ($0.20/$1.20), Gemini 2.5 Flash ($0.075/$0.30), DeepSeek V4.1 Flash ($0.14/$0.28), and Qwen 2.5 Coder 32B ($0.08/$0.16) now deliver capability that surpasses 2024 flagships at less than one-twentieth of the price.
2. The Prompt Caching Revolution: 82% Average Discount
In 2024, prompt caching was a novelty offered by only one provider. As of September 2026, 8 out of 9 major providers support native prefix caching:
- Anthropic: 90% read discount on 1,024+ token prefixes.
- Google: Free cache reads on Gemini 3.8 Flash ($0.00/1M).
- DeepSeek: 90% prefix discount ($0.014/1M on V4.1 Flash).
- OpenAI: 50% to 75% automated prefix caching.
- Cohere: 90% read discount on enterprise RAG prefixes.
3. The Bifurcation of Frontier Intelligence
The industry is no longer competing along a single price curve. Instead, pricing has bifurcated into two distinct markets:
4. Key Recommendations for Engineering Teams
- Enforce Two-Tier Routing: Direct 80% of routine traffic to utility models; route only complex logic to the deliberation tier.
- Lock Prompt Prefixes: Freeze system prompts to guarantee cache hit rates >75%.
- Leverage Asynchronous Batches: Run bulk jobs via Batch APIs for an immediate 50% discount.
Explore Full Audited Pricing Registry
Filter all 30 models with live rate comparisons.