{T}
TokenMath.net
Verified against Together AI / Meta Official Pricing
Meta / TogetherbalancedGenerally AvailableActive Production

Llama 4 Vivas Token Counter & Cost Calculator

Open-weights balanced model with native long-context support across Together and first-party serving

Standard Input / 1M
$0.25
Prompt tokens
Cached Input / 1M
$0.06
Prompt caching rate
Output / 1M
$0.85
Generation completion
Context Window
1M
Max out: 128k
API & Developer Integration Details
generally available
API Model Stringmeta-llama/Llama-4-Vivas-70B-Instruct
Provider & FamilyMeta / Together (llama3)
Prompt Caching Rules1,024 min tokens (76% off)
Long-Context TiersStandard flat rate throughout full window
First-party audit timestamp: 2026-09-23 04:43 UTCOfficial Documentation

Interactive Cost & Token Simulator for Llama 4 Vivas

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Prompt Caching Ratio (50%)$0.06/M cached rate
Cost / Request$0.00107
Daily Spend (5,000 reqs)$5.3375
Monthly Run-Rate (30d)$160.13

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00068$0.00048
Document Summarization10,0002,000$0.00420$0.00230
Codebase & Context Analysis100,00020,000$0.0420$0.0230
Batch Corpus Processing1,000,000100,000$0.3350$0.1450
Technical Architecture & Pricing VerificationVerified: 2026-09-23 04:43 UTC
Model ArchitectureLlama 4 Vivas
Provider OrganizationMeta / Together
Tokenizer EncodingMeta Llama 3 Tokenizer (128k vocabulary)
API ConnectivityTogether AI & Meta serving endpoints with automatic prefix caching.
Batch API 50% DiscountSupported (50% off input & output)
Tier 1 Rate Quota (Free-Tier Reference)500 RPM · 1M TPM

Tier 1 (free tier) reference limits — your account's actual quota may be higher.

Blended 3:1 Cost / 1M Tokens$0.40
Balance-focused successor in the Llama 4 family with automatic prefix caching. Cache writes incur a 1.25x input fee ($0.3125/1M) and cached reads receive a 76% discount ($0.06/1M).Together AI / Meta Official Pricing

Verified Pricing History

September 2026: $0.25/M input · $0.85/M output ($0.06/M cached)
Llama 4 Vivas launch with automatic prefix caching.

When to Choose Llama 4 Vivas

Ideal for production workloads demanding balanced capabilities, deep context depth (1M tokens), and reliability from Meta / Together. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About Llama 4 Vivas

How much does 1 million tokens cost with Llama 4 Vivas?

For Llama 4 Vivas, 1 million input tokens costs $0.25, while 1 million output tokens costs $0.85. If using prompt caching, repetitive input prefixes are discounted to $0.06 per million.

What is the context window for Llama 4 Vivas?

Llama 4 Vivas features a maximum context window of 1,000,000 tokens (~750,000 words), with a maximum output limit of 131,072 tokens per completion.

Which tokenizer does Llama 4 Vivas use?

Llama 4 Vivas utilizes the Meta Llama 3 Tokenizer (128k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does Llama 4 Vivas support prompt caching discounts?

Yes. Llama 4 Vivas supports prompt caching with a cached input rate of $0.06/1M (saving up to 76% on repeated input context).