Enterprise-grade foundation model optimized for conversational RAG, tool use, and citations
command-r-plus-08-2024| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.00750 | $0.00525 |
| Document Summarization | 10,000 | 2,000 | $0.0450 | $0.0225 |
| Codebase & Context Analysis | 100,000 | 20,000 | $0.4500 | $0.2250 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $3.50 | $1.25 |
Ideal for production workloads demanding flagship capabilities, deep context depth (128k tokens), and reliability from Cohere. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For Command R+ (08-2024), 1 million input tokens costs $2.50, while 1 million output tokens costs $10.00. If using prompt caching, repetitive input prefixes are discounted to $0.25 per million.
Command R+ (08-2024) features a maximum context window of 128,000 tokens (~96,000 words), with a maximum output limit of 4,096 tokens per completion.
Command R+ (08-2024) utilizes the Cohere BPE (~256k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. Command R+ (08-2024) supports prompt caching with a cached input rate of $0.25/1M (saving up to 90% on repeated input context).