Open-weights balanced model with native long-context support across Together and first-party serving
meta-llama/Llama-4-Vivas-70B-Instruct| Scenario | Input Tokens | Output Tokens | Uncached Cost | With Prompt Caching |
|---|---|---|---|---|
| Short Chat Query | 1,000 | 500 | $0.00068 | $0.00048 |
| Document Summarization | 10,000 | 2,000 | $0.00420 | $0.00230 |
| Codebase & Context Analysis | 100,000 | 20,000 | $0.0420 | $0.0230 |
| Batch Corpus Processing | 1,000,000 | 100,000 | $0.3350 | $0.1450 |
Tier 1 (free tier) reference limits — your account's actual quota may be higher.
Ideal for production workloads demanding balanced capabilities, deep context depth (1M tokens), and reliability from Meta / Together. Excellent when predictable tokenomics and prompt caching support are paramount.
If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.
For Llama 4 Vivas, 1 million input tokens costs $0.25, while 1 million output tokens costs $0.85. If using prompt caching, repetitive input prefixes are discounted to $0.06 per million.
Llama 4 Vivas features a maximum context window of 1,000,000 tokens (~750,000 words), with a maximum output limit of 131,072 tokens per completion.
Llama 4 Vivas utilizes the Meta Llama 3 Tokenizer (128k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.
Yes. Llama 4 Vivas supports prompt caching with a cached input rate of $0.06/1M (saving up to 76% on repeated input context).