AI Model Pricing Intelligence & Data Registry
Every price displayed on TokenMath is actively monitored, cross-verified against official provider documentation, and timestamped. Explore our centralized dataset, audit log, and verification methodology.
Frontier AI Model Pricing Matrix
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M | Context | Verified Date | Official Source |
|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | 200k | September 8, 2026 | Anthropic Official Pricing |
| Claude Opus 5 | Anthropic | $5.00 | $0.50 | $25.00 | 200k | September 8, 2026 | Anthropic Official Pricing |
| Claude Sonnet 5 | Anthropic | $3.00 | $0.20 | $15.00 | 200k | September 8, 2026 | Anthropic Official Pricing |
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 | 200k | September 8, 2026 | Anthropic Official Pricing |
| GPT-6 Astra | OpenAI | $10.00 | $5.00 | $50.00 | 256k | September 8, 2026 | OpenAI Official Pricing |
| GPT-5.6 Sol | OpenAI | $4.00 | $2.00 | $20.00 | 200k | September 8, 2026 | OpenAI Official Pricing |
| GPT-5.6 Terra | OpenAI | $2.00 | $1.00 | $12.00 | 128k | September 8, 2026 | OpenAI Official Pricing |
| GPT-5.6 Luna | OpenAI | $0.20 | $0.10 | $1.20 | 128k | September 8, 2026 | OpenAI Official Pricing |
| OpenAI o3 | OpenAI | $2.00 | $1.00 | $8.00 | 200k | September 8, 2026 | OpenAI Official Pricing |
| OpenAI o3-pro | OpenAI | $20.00 | $10.00 | $80.00 | 256k | September 8, 2026 | OpenAI Official Pricing |
| OpenAI o4-mini | OpenAI | $1.10 | $0.55 | $4.40 | 128k | September 8, 2026 | OpenAI Official Pricing |
| Gemini 3.8 Flash | $0.75 | $0.0750 | $3.75 | 1.05M | September 8, 2026 | Google Cloud / AI Studio Pricing | |
| Gemini 3.1 Pro | $2.00 | $0.20 | $12.00 | 2.10M | September 8, 2026 | Google Cloud / AI Studio Pricing | |
| Gemini 2.5 Flash-Lite | $0.10 | $0.0250 | $0.40 | 1.05M | September 8, 2026 | Google Cloud / AI Studio Pricing | |
| DeepSeek-V4.1-Flash | DeepSeek | $0.20 | $0.0050 | $0.80 | 1.05M | September 8, 2026 | DeepSeek Official API Pricing |
| Llama 4 Scout (109B) | Meta / Open | $0.30 | None | $0.60 | 10M | September 8, 2026 | Meta Llama on Together AI / Groq |
| Llama 4 Maverick (400B) | Meta / Open | $0.80 | None | $1.60 | 1M | September 8, 2026 | Meta Llama on Together AI / Fireworks |
| Mistral Large 3 | Mistral AI | $0.50 | $0.10 | $1.50 | 128k | September 8, 2026 | Mistral AI Official Pricing |
| Mistral Small 4 | Mistral AI | $0.15 | $0.0300 | $0.60 | 128k | September 8, 2026 | Mistral AI Official Pricing |
Recent Pricing Changes & Market Movements
Released at $3.00/1M input and $15.00/1M output, with prompt caching write at $3.75/1M and read at $0.30/1M.
Introduced frontier omni tier at $10.00/1M input and $50.00/1M output with 50% caching discount.
Updated to $0.27/1M input and $1.10/1M output; disk-cached prompt hits discounted to $0.028/1M.
Baseline 1M context pricing confirmed at $0.75/1M input and $3.75/1M output.
General availability pricing at $15.00/1M input and $60.00/1M reasoning output.
How TokenMath Verifies AI Pricing
TokenMath is built on transparency. We never scrape unverified aggregator blogs or fabricate rates. Here is our exact verification standard:
1. Primary Source Verification Only
All pricing values are sourced directly from official provider developer documentation (e.g. openai.com/api/pricing, anthropic.com/pricing, ai.google.dev/pricing, or direct cloud billing tables). Every single model links to its canonical source.
2. Weekly Audit Frequency
Our engineering team reviews active model pricing schedules weekly. In addition, automated web hooks and community pull requests alert us immediately when frontier providers announce price drops or new model revisions.
3. Exact vs Estimated Distinction
Token prices for text inference are 100% exact math based on official per-million token rates. For vision and multimodal inputs, calculations adhere strictly to the provider's published patch tiling algorithms (such as OpenAI's 512px tile formula and Anthropic's pixels/750 rule).
4. Handling Tiered & Batch Nuances
Where providers offer Batch API discounts (50% off for 24h turnaround) or prompt caching reductions (50% to 90% off), we model both standard synchronous rates and discounted tiers explicitly without obscuring the true baseline cost.