Fast mid-tier reasoning successor to o4-mini for STEM, code, and agentic pipelines
o5-mini| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.00240 | $0.00180 |
| Document Summarization | 10,000 | 2,000 | $0.0144 | $0.00840 |
| Codebase & Context Analysis | 100,000 | 20,000 | $0.1440 | $0.0840 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $1.12 | $0.5200 |
Tier 1 (free tier) reference limits — your account's actual quota may be higher.
Ideal for production workloads demanding reasoning capabilities, deep context depth (200k tokens), and reliability from OpenAI. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For OpenAI o5-mini, 1 million input tokens costs $0.80, while 1 million output tokens costs $3.20. If using prompt caching, repetitive input prefixes are discounted to $0.20 per million.
OpenAI o5-mini features a maximum context window of 200,000 tokens (~150,000 words), with a maximum output limit of 100,000 tokens per completion.
OpenAI o5-mini utilizes the OpenAI o200k_base (200k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. OpenAI o5-mini supports prompt caching with a cached input rate of $0.20/1M (saving up to 75% on repeated input context).