AI Model Directory & Token Pricing
Browse complete token specifications, input/cached/output costs, context window depths, and provider rate limit quotas across all 16 supported frontier models.
Next generation intelligence for long-running autonomous agents
Ideal for complex agentic coding, deep domain synthesis, and enterprise work
High-performance frontier model for production coding and complex agents
Fastest, most cost-efficient model for sub-second high-volume tasks
OpenAI frontier flagship omni model for autonomous tasks and deep cross-modal reasoning
High-capability general intelligence workhorse with promotional 2026 pricing
Balanced production model delivering top-tier performance at everyday rates
Ultra-low-latency budget model engineered for high-frequency micro-tasks
Next-generation reasoning model optimized for mathematical rigor and complex software engineering
Maximum compute reasoning model for PhD-level research and formal theorem proving
High-speed reasoning model specialized for STEM, competitive code, and analytical pipelines
Google frontier Flash model with native 1M context, sub-second latency, and multimodal video support
Frontier 2M context Pro model with state-of-the-art multimodal reasoning and coding depth
Ultra-economical 1M context model for automated high-volume bulk pipelines
Flagship 1M context multimodal architecture with ultra-cheap prefix caching ($0.003/M)
Meta open architecture with extreme 10 Million token context window and efficient 17B active MoE
Meta flagship open-weights MoE model designed for high-reasoning multimodal dialogue
Mistral flagship reasoning and coding model with aggressive 2026 pricing ($0.50/M in)
High-speed compact model delivering low-latency inference at minimal compute cost