{T}
TokenMath.net
Entity Intelligence Hub9 Active AI Ecosystems · 30 Verified Models

AI Provider Directory & Pricing Policies

Every major AI provider implements fundamentally different billing rules for prompt caching, asynchronous batches, long-context thresholds, and multi-turn reasoning. Explore each provider's architecture below to optimize your API unit economics.

O

OpenAI

San Francisco, CA, USA
8 models

Pioneering frontier intelligence, reasoning models (o-series), and multimodal omni APIs

Input Rates:$0.20 – $20.00/1M
Cache: Automatic prefix caching on >= 1,024 tokens (50% to 75% discount off standard input).
A

Anthropic

San Francisco, CA, USA
4 models

Enterprise-grade frontier models engineered for coding, complex agents, and steerability

Input Rates:$1.00 – $10.00/1M
Cache: 90% read discount with 1.25x write surcharge on 1,024+ token prefixes (5-minute auto-refreshing TTL).
G

Google Cloud / Gemini

Mountain View, CA, USA
4 models

Million-token context native multimodal processing and high-throughput production inference

Input Rates:$0.075 – $2.00/1M
Cache: Gemini 3.8 Flash offers 100% discount on cache hits ($0.00/1M). Storage incurs $0.50/1M tokens/hour promo fee.
D

DeepSeek

Hangzhou, China
2 models

Ultra-cost-efficient Mixture-of-Experts (MoE) architecture with dynamic off-peak scheduling

Input Rates:$0.14 – $0.44/1M
Cache: Automatic prefix caching discounts hits by up to 90% ($0.014/1M on V4.1 Flash).
M

Mistral AI

Paris, France
3 models

European frontier foundation models, specialized code generators, and multilingual reasoning

Input Rates:$0.15 – $0.50/1M
Cache: Prompt caching supported on Large 3 ($0.15/1M cached) and Small 4 ($0.015/1M cached).
M

Meta Llama

Menlo Park, CA, USA
3 models

The open-weights standard powering global AI research, serverless inference, and self-hosted deployments

Input Rates:$0.18 – $3.50/1M
Cache:
x

xAI

San Francisco, CA, USA
2 models

Real-time knowledge, mathematical rigor, and high-throughput multimodal intelligence

Input Rates:$0.20 – $2.00/1M
Cache:
A

Alibaba Cloud / Qwen

Hangzhou, China
2 models

State-of-the-art open multilingual and specialized coding models with massive vocabulary BPE

Input Rates:$0.08 – $0.35/1M
Cache:
C

Cohere

Toronto, Canada
2 models

Enterprise search, retrieval-augmented generation (RAG), and multi-step tool orchestration

Input Rates:$0.15 – $2.50/1M
Cache: 90% read discount on cached documents and prompt prefixes ($0.25/1M on Command R+).