{T}
TokenMath.net
Verified against Together AI Serverless Pricing
MetaflagshipGenerally Available

Llama 3.1 405B — Together AI Token Counter & Cost Calculator

Massive 405-billion parameter open-weights flagship rivaling the largest proprietary models

Standard Input / 1M
$3.50
Prompt tokens
Cached Input / 1M
N/A
Not supported
Output / 1M
$3.50
Generation completion
Context Window
128k
Max out: 4.1k
API & Developer Integration Details
generally available
API Model Stringmeta-llama/Meta-Llama-3.1-405B-Instruct-Turbo
Provider & FamilyMeta (llama3)
Prompt Caching RulesNone
Long-Context TiersStandard flat rate throughout full window
First-party audit timestamp: 2026-09-15 15:00 UTCOfficial Documentation

Interactive Cost & Token Simulator for Llama 3.1 405B — Together AI

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Cost / Request$0.0116
Daily Spend (5,000 reqs)$57.75
Monthly Run-Rate (30d)$1,732.50

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00525N/A
Document Summarization10,0002,000$0.0420N/A
Codebase & Context Analysis100,00020,000$0.4200N/A
Batch Corpus Processing1,000,000100,000$3.85N/A
Technical Architecture & Pricing VerificationVerified: 2026-09-15 15:00 UTC
Model ArchitectureLlama 3.1 405B — Together AI
Provider OrganizationMeta
Tokenizer EncodingMeta Llama 3 Tiktoken BPE (~128k vocabulary)
API ConnectivityTogether AI serverless 405B cluster running 8xH100 tensor-parallel inference.
Batch API 50% DiscountNot Available
Tier 1 Rate Quota100 RPM · 500k TPM
Blended 3:1 Cost / 1M Tokens$3.50
Symmetric token pricing ($3.50/1M in and out) for the largest open model in production.Together AI Serverless Pricing

Verified Pricing History

September 2026: $3.5/M input · $3.5/M output ($0.875/M cached)
Verified Llama 3.1 405B Together AI rates.

When to Choose Llama 3.1 405B — Together AI

Ideal for production workloads demanding flagship capabilities, deep context depth (128k tokens), and reliability from Meta. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About Llama 3.1 405B — Together AI

How much does 1 million tokens cost with Llama 3.1 405B — Together AI?

For Llama 3.1 405B — Together AI, 1 million input tokens costs $3.50, while 1 million output tokens costs $3.50. If using prompt caching, repetitive input prefixes are discounted to $0.875 per million.

What is the context window for Llama 3.1 405B — Together AI?

Llama 3.1 405B — Together AI features a maximum context window of 131,072 tokens (~98,304 words), with a maximum output limit of 4,096 tokens per completion.

Which tokenizer does Llama 3.1 405B — Together AI use?

Llama 3.1 405B — Together AI utilizes the Meta Llama 3 Tiktoken BPE (~128k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does Llama 3.1 405B — Together AI support prompt caching discounts?

No, Llama 3.1 405B — Together AI does not currently advertise prompt caching discounts on its standard API tier.