Official Provider HubMenlo Park, CA, USA
Meta Llama API Pricing & Token Economics
Meta's Llama family (Llama 4 Scout, Llama 3.3 70B, and Llama 3.1 405B) represents the premier open-weights architecture in AI. Hosted by serverless cloud partners such as Together AI, Groq, and AWS Bedrock, Llama models provide ultra-competitive per-token rates without proprietary vendor lock-in.
Prompt Caching Policy
Batch / Asynchronous Rules
Standard API execution rules. Refer to individual endpoint rate limits for asynchronous processing.
Special Billing Conditions
Zero write fees on standard prefix evaluation; automatic cache invalidation.
Active Meta Llama Models(3 models verified)
All 30 Models →| Model | Context | Input / 1M | Cached / 1M | Output / 1M | Batch Discount | Actions |
|---|---|---|---|---|---|---|
Llama 4 Scout (17B Active / 109B MoE) — Together AI meta-llama/Llama-4-Scout-17B-16E-Instruct | 328k | $0.18 | — | $0.59 | N/A | Specs → |
Llama 4 Maverick (400B) — Together AI meta-llama/llama-4-maverick-instruct | 1M | $0.80 | — | $1.60 | N/A | Specs → |
Llama 3.1 405B — Together AI meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo | 128k | $3.50 | — | $3.50 | N/A | Specs → |
Calculate your monthly Meta Llama bill
Simulate prompt caching hit rates, batch discounts, and output token costs with zero server logging.