Foundational deliberation model for complex multi-step science, coding, and mathematical reasoning
o1-2024-12-17| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.0450 | $0.0375 |
| Document Summarization | 10,000 | 2,000 | $0.2700 | $0.1950 |
| Codebase & Context Analysis | 100,000 | 20,000 | $2.70 | $1.95 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $21.00 | $13.50 |
Ideal for production workloads demanding reasoning capabilities, deep context depth (200k tokens), and reliability from OpenAI. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For OpenAI o1, 1 million input tokens costs $15.00, while 1 million output tokens costs $60.00. If using prompt caching, repetitive input prefixes are discounted to $7.50 per million.
OpenAI o1 features a maximum context window of 200,000 tokens (~150,000 words), with a maximum output limit of 100,000 tokens per completion.
OpenAI o1 utilizes the OpenAI o200k_base (200k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. OpenAI o1 supports prompt caching with a cached input rate of $7.50/1M (saving up to 50% on repeated input context).