Home/Models/Llama 4 Maverick (400B)
Verified against Meta Llama on Together AI / Fireworks
Meta / OpenflagshipGenerally Available

Llama 4 Maverick (400B) Token Counter & Cost Calculator

Meta flagship open-weights MoE model designed for high-reasoning multimodal dialogue

Standard Input / 1M
$0.80
Prompt tokens
Cached Input / 1M
N/A
Not supported
Output / 1M
$1.60
Generation completion
Context Window
1M
Max out: 16.4k

Interactive Cost & Token Simulator for Llama 4 Maverick (400B)

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Cost / Request$0.00328
Daily Spend (5,000 reqs)$16.40
Monthly Run-Rate (30d)$492.00

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00160N/A
Document Summarization10,0002,000$0.0112N/A
Codebase & Context Analysis100,00020,000$0.1120N/A
Batch Corpus Processing1,000,000100,000$0.9600N/A
Technical Architecture & Pricing VerificationVerified: September 8, 2026
Model ArchitectureLlama 4 Maverick (400B)
Provider OrganizationMeta / Open
Tokenizer EncodingMeta Llama 3/4 Tiktoken (128k vocabulary)
API ConnectivityManaged inference API and downloadable weights for FP8/INT4 cluster deployment.
Batch API 50% DiscountNot Available
Tier 1 Rate Quota600 RPM · 600k TPM
Blended 3:1 Cost / 1M Tokens$1.00
Frontier 400B dense/MoE open-weight architecture rivaling Claude Sonnet 5.Meta Llama on Together AI / Fireworks

Verified Pricing History

August 2026: $1.8/M input · $3.6/M output ($0.9/M cached)
Llama 4 Maverick 400B initial hosted endpoint availability.

When to Choose Llama 4 Maverick (400B)

Ideal for production workloads demanding flagship capabilities, deep context depth (1M tokens), and reliability from Meta / Open. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About Llama 4 Maverick (400B)

How much does 1 million tokens cost with Llama 4 Maverick (400B)?

For Llama 4 Maverick (400B), 1 million input tokens costs $0.80, while 1 million output tokens costs $1.60. If using prompt caching, repetitive input prefixes are discounted to $0.80 per million.

What is the context window for Llama 4 Maverick (400B)?

Llama 4 Maverick (400B) features a maximum context window of 1,000,000 tokens (~750,000 words), with a maximum output limit of 16,384 tokens per completion.

Which tokenizer does Llama 4 Maverick (400B) use?

Llama 4 Maverick (400B) utilizes the Meta Llama 3/4 Tiktoken (128k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does Llama 4 Maverick (400B) support prompt caching discounts?

No, Llama 4 Maverick (400B) does not currently advertise prompt caching discounts on its standard API tier.