Generally available flagship with 1M native context and tiered >200k long-context pricing
gemini-3.8-pro-001| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.0130 | $0.00940 |
| Document Summarization | 10,000 | 2,000 | $0.0760 | $0.0400 |
| Codebase & Context Analysis | 100,000 | 20,000 | $0.7600 | $0.4000 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $5.80 | $2.20 |
Tier 1 (free tier) reference limits — your account's actual quota may be higher.
Ideal for production workloads demanding flagship capabilities, deep context depth (1.05M tokens), and reliability from Google. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For Gemini 3.8 Pro, 1 million input tokens costs $4.00, while 1 million output tokens costs $18.00. If using prompt caching, repetitive input prefixes are discounted to $0.40 per million.
Gemini 3.8 Pro features a maximum context window of 1,048,576 tokens (~786,432 words), with a maximum output limit of 65,536 tokens per completion.
Gemini 3.8 Pro utilizes the Google Gemini SentencePiece (256k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. Gemini 3.8 Pro supports prompt caching with a cached input rate of $0.40/1M (saving up to 90% on repeated input context).