API Specs & DirectoryUpdated September 2026

AI Model Directory & Token Pricing

Browse complete token specifications, input/cached/output costs, context window depths, and provider rate limit quotas across all 16 supported frontier models.

reasoning

Next generation intelligence for long-running autonomous agents

Input / M$10.00
Output / M$50.00
Cache Read / Write$0.25 / $12.50
Context Limit200k
Tier 1 Limit:100 RPM · 100k TPM
AnthropicClaude Opus 5
flagship

Ideal for complex agentic coding, deep domain synthesis, and enterprise work

Input / M$5.00
Output / M$25.00
Cache Read / Write$0.50 / $6.25
Context Limit200k
Tier 1 Limit:100 RPM · 100k TPM
flagship

High-performance frontier model for production coding and complex agents

Input / M$3.00
Output / M$15.00
Cache Read / Write$0.20 / $2.50
Context Limit200k
Tier 1 Limit:1,000 RPM · 200k TPM
budget

Fastest, most cost-efficient model for sub-second high-volume tasks

Input / M$1.00
Output / M$5.00
Cache Read / Write$0.10 / $1.25
Context Limit200k
Tier 1 Limit:2,000 RPM · 500k TPM
flagship

OpenAI frontier flagship omni model for autonomous tasks and deep cross-modal reasoning

Input / M$10.00
Output / M$50.00
Cached Input$5.00
Context Limit256k
Tier 1 Limit:500 RPM · 50k TPM
flagship

High-capability general intelligence workhorse with promotional 2026 pricing

Input / M$4.00
Output / M$20.00
Cached Input$2.00
Context Limit200k
Tier 1 Limit:1,000 RPM · 200k TPM
balanced

Balanced production model delivering top-tier performance at everyday rates

Input / M$2.00
Output / M$12.00
Cached Input$1.00
Context Limit128k
Tier 1 Limit:2,000 RPM · 500k TPM
budget

Ultra-low-latency budget model engineered for high-frequency micro-tasks

Input / M$0.20
Output / M$1.20
Cached Input$0.10
Context Limit128k
Tier 1 Limit:5,000 RPM · 2M TPM
OpenAIOpenAI o3
reasoning

Next-generation reasoning model optimized for mathematical rigor and complex software engineering

Input / M$2.00
Output / M$8.00
Cached Input$1.00
Context Limit200k
Tier 1 Limit:500 RPM · 100k TPM
reasoning

Maximum compute reasoning model for PhD-level research and formal theorem proving

Input / M$20.00
Output / M$80.00
Cached Input$10.00
Context Limit256k
Tier 1 Limit:100 RPM · 30k TPM
reasoning

High-speed reasoning model specialized for STEM, competitive code, and analytical pipelines

Input / M$1.10
Output / M$4.40
Cached Input$0.55
Context Limit128k
Tier 1 Limit:1,000 RPM · 500k TPM
budget

Google frontier Flash model with native 1M context, sub-second latency, and multimodal video support

Input / M$0.75
Output / M$3.75
Cached Input$0.07
Context Limit1.05M
Tier 1 Limit:2,000 RPM · 4M TPM
flagship

Frontier 2M context Pro model with state-of-the-art multimodal reasoning and coding depth

Input / M$2.00
Output / M$12.00
Cached Input$0.20
Context Limit2.10M
Tier 1 Limit:1,000 RPM · 2M TPM

Ultra-economical 1M context model for automated high-volume bulk pipelines

Input / M$0.10
Output / M$0.40
Cached Input$0.03
Context Limit1.05M
Tier 1 Limit:4,000 RPM · 4M TPM
balanced

Flagship 1M context multimodal architecture with ultra-cheap prefix caching ($0.003/M)

Input / M$0.20
Output / M$0.80
Cached Input$0.01
Context Limit1.05M
Tier 1 Limit:200 RPM · 2M TPM
balanced

Meta open architecture with extreme 10 Million token context window and efficient 17B active MoE

Input / M$0.30
Output / M$0.60
Cached InputNone
Context Limit10M
Tier 1 Limit:1,000 RPM · 1M TPM
flagship

Meta flagship open-weights MoE model designed for high-reasoning multimodal dialogue

Input / M$0.80
Output / M$1.60
Cached InputNone
Context Limit1M
Tier 1 Limit:600 RPM · 600k TPM
Mistral AIMistral Large 3
flagship

Mistral flagship reasoning and coding model with aggressive 2026 pricing ($0.50/M in)

Input / M$0.50
Output / M$1.50
Cached Input$0.10
Context Limit128k
Tier 1 Limit:300 RPM · 1M TPM
Mistral AIMistral Small 4
budget

High-speed compact model delivering low-latency inference at minimal compute cost

Input / M$0.15
Output / M$0.60
Cached Input$0.03
Context Limit128k
Tier 1 Limit:600 RPM · 2M TPM