Back to Guides5 min read • Updated Sept 2026
Conversion Formula
Tokens vs Words: The Definitive Conversion Guide
Quick rule-of-thumb ratios often mislead engineers when estimating production LLM architectures. Here is how token counts vary across English prose, source code, structured JSON, and multilingual scripts.
1. The Baseline English Prose Ratio
For standard grammatical English text (articles, essays, emails):
Word to Token1 Word ≈ 1.33 TokensOr 750 words = 1,000 tokens
Character to Token1 Token ≈ 4.0 CharactersIncluding spaces and punctuation
2. Source Code, JSON & Technical Documents
Code does not follow natural language distribution. Identifiers like fetchCustomerSubscriptionHistory get fragmented into 4 or 5 subwords. In addition, syntax characters like braces ({ }), tabs, and semicolons often consume independent tokens.
- Python / TypeScript: 1 line of code typically averages 10 to 14 tokens.
- JSON Payloads: Repeating keys (
"status","created_at") and quotes cause high token inflation. A 10KB JSON payload is usually 2,800 to 3,500 tokens. - Single-Spaced PDF Page: A standard 8.5" x 11" single-spaced page (approx. 500 words) produces approximately 650 to 750 tokens.
3. Multilingual Multipliers
Because tokenizers dedicate a majority of vocabulary entries to English, non-Latin alphabets suffer from byte fallback penalties:
| Language Script | Tokens Per 1,000 Words | Cost Multiplier vs English |
|---|---|---|
| English | ~1,330 tokens | 1.0x (Baseline) |
| Spanish / French / German | ~1,500 - 1,700 tokens | 1.1x - 1.3x |
| Russian (Cyrillic) | ~2,200 - 2,600 tokens | 1.7x - 2.0x |
| Japanese / Chinese / Korean | ~2,400 - 3,200 tokens | 1.8x - 2.4x |