Home/Models/OpenAI o4-mini
Verified against OpenAI Official Pricing
OpenAIreasoningGenerally Available

OpenAI o4-mini Token Counter & Cost Calculator

High-speed reasoning model specialized for STEM, competitive code, and analytical pipelines

Standard Input / 1M
$1.10
Prompt tokens
Cached Input / 1M
$0.55
Prompt caching rate
Output / 1M
$4.40
Generation completion
Context Window
128k
Max out: 65.5k

Interactive Cost & Token Simulator for OpenAI o4-mini

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Prompt Caching Ratio (50%)$0.55/M cached rate
Cost / Request$0.00558
Daily Spend (5,000 reqs)$27.9125
Monthly Run-Rate (30d)$837.38

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00330$0.00275
Document Summarization10,0002,000$0.0198$0.0143
Codebase & Context Analysis100,00020,000$0.1980$0.1430
Batch Corpus Processing1,000,000100,000$1.54$0.9900
Technical Architecture & Pricing VerificationVerified: September 8, 2026
Model ArchitectureOpenAI o4-mini
Provider OrganizationOpenAI
Tokenizer EncodingOpenAI o200k_base (200k vocabulary)
API ConnectivityOpenAI Chat Completions API with reasoning tokens support.
Batch API 50% DiscountSupported (50% off input & output)
Tier 1 Rate Quota1,000 RPM · 500k TPM
Blended 3:1 Cost / 1M Tokens$1.93
Fast, cost-effective reasoning model. Delivers strong coding and math at 1/10th the cost of o3.OpenAI Official Pricing

Verified Pricing History

August 2026: $1.1/M input · $4.4/M output ($0.55/M cached)
Successor to o1-mini at lower cost and higher throughput.

When to Choose OpenAI o4-mini

Ideal for production workloads demanding reasoning capabilities, deep context depth (128k tokens), and reliability from OpenAI. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About OpenAI o4-mini

How much does 1 million tokens cost with OpenAI o4-mini?

For OpenAI o4-mini, 1 million input tokens costs $1.10, while 1 million output tokens costs $4.40. If using prompt caching, repetitive input prefixes are discounted to $0.55 per million.

What is the context window for OpenAI o4-mini?

OpenAI o4-mini features a maximum context window of 128,000 tokens (~96,000 words), with a maximum output limit of 65,536 tokens per completion.

Which tokenizer does OpenAI o4-mini use?

OpenAI o4-mini utilizes the OpenAI o200k_base (200k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does OpenAI o4-mini support prompt caching discounts?

Yes. OpenAI o4-mini supports prompt caching with a cached input rate of $0.55/1M (saving up to 50% on repeated input context).