{T}
TokenMath.net
Verified against Alibaba Cloud / Qwen Official Pricing
Alibaba / QwenbalancedGenerally AvailableActive Production

Qwen3 235B-A22B Token Counter & Cost Calculator

MoE flagship balancing frontier quality with framework-native serving costs

Standard Input / 1M
$0.455
Prompt tokens
Cached Input / 1M
$0.18
Prompt caching rate
Output / 1M
$1.82
Generation completion
Context Window
128k
Max out: 32k
API & Developer Integration Details
generally available
API Model Stringqwen3-235b-a22b
Provider & FamilyAlibaba / Qwen (qwen)
Prompt Caching Rules1,024 min tokens (80% off)
Long-Context TiersStandard flat rate throughout full window
First-party audit timestamp: 2026-09-23 04:43 UTCOfficial Documentation

Interactive Cost & Token Simulator for Qwen3 235B-A22B

Live Calculation
Input Tokens2,500
Output Tokens800
Requests / Day5,000
Prompt Caching Ratio (50%)$0.18/M cached rate
Cost / Request$0.00225
Daily Spend (5,000 reqs)$11.2488
Monthly Run-Rate (30d)$337.46

Standard Workload Cost Scenarios

ScenarioInput TokensOutput TokensUncached CostWith Prompt Caching
Short Chat Query1,000500$0.00136$0.00109
Document Summarization10,0002,000$0.00819$0.00544
Codebase & Context Analysis100,00020,000$0.0819$0.0544
Batch Corpus Processing1,000,000100,000$0.6370$0.3620
Technical Architecture & Pricing VerificationVerified: 2026-09-23 04:43 UTC
Model ArchitectureQwen3 235B-A22B
Provider OrganizationAlibaba / Qwen
Tokenizer EncodingQwen SentencePiece BPE (152k vocabulary)
API ConnectivityAlibaba Cloud Model Studio dashscope-compatible chat completions endpoint.
Batch API 50% DiscountSupported (50% off input & output)
Tier 1 Rate Quota (Free-Tier Reference)500 RPM · 1M TPM

Tier 1 (free tier) reference limits — your account's actual quota may be higher.

Blended 3:1 Cost / 1M Tokens$0.80
MoE flagship with framework-native serving. Automatic prefix caching gives an 80% cached-read discount ($0.18/1M); cache writes carry a 1.25x input fee ($1.125/1M).Alibaba Cloud / Qwen Official Pricing

Verified Pricing History

2026-09-23 04:43 UTC: $/M input · $/M output
September 2026: $0.9/M input · $1.6/M output ($0.18/M cached)
Qwen3 235B-A22B launch pricing.

When to Choose Qwen3 235B-A22B

Ideal for production workloads demanding balanced capabilities, deep context depth (128k tokens), and reliability from Alibaba / Qwen. Excellent when predictable tokenomics and prompt caching support are paramount.

When Another Model May Be Better

If your use-case requires sub-second streaming latency or ultra-high frequency classification at micro-cent pricing, consider lighter budget options such as Gemini Flash-Lite or Claude Haiku. For deep formal logic, consider dedicated reasoning models like o3.

Frequently Asked Questions About Qwen3 235B-A22B

How much does 1 million tokens cost with Qwen3 235B-A22B?

For Qwen3 235B-A22B, 1 million input tokens costs $0.455, while 1 million output tokens costs $1.82. If using prompt caching, repetitive input prefixes are discounted to $0.18 per million.

What is the context window for Qwen3 235B-A22B?

Qwen3 235B-A22B features a maximum context window of 131,072 tokens (~98,304 words), with a maximum output limit of 32,768 tokens per completion.

Which tokenizer does Qwen3 235B-A22B use?

Qwen3 235B-A22B utilizes the Qwen SentencePiece BPE (152k vocabulary). Token counting on TokenMath runs client-side to ensure maximum privacy.

Does Qwen3 235B-A22B support prompt caching discounts?

Yes. Qwen3 235B-A22B supports prompt caching with a cached input rate of $0.18/1M (saving up to 60% on repeated input context).