Back to Guides5 min read • Updated Sept 2026
Conversion Formula

Tokens vs Words: The Definitive Conversion Guide

Quick rule-of-thumb ratios often mislead engineers when estimating production LLM architectures. Here is how token counts vary across English prose, source code, structured JSON, and multilingual scripts.

1. The Baseline English Prose Ratio

For standard grammatical English text (articles, essays, emails):

Word to Token1 Word ≈ 1.33 TokensOr 750 words = 1,000 tokens
Character to Token1 Token ≈ 4.0 CharactersIncluding spaces and punctuation

2. Source Code, JSON & Technical Documents

Code does not follow natural language distribution. Identifiers like fetchCustomerSubscriptionHistory get fragmented into 4 or 5 subwords. In addition, syntax characters like braces ({ }), tabs, and semicolons often consume independent tokens.

  • Python / TypeScript: 1 line of code typically averages 10 to 14 tokens.
  • JSON Payloads: Repeating keys ("status", "created_at") and quotes cause high token inflation. A 10KB JSON payload is usually 2,800 to 3,500 tokens.
  • Single-Spaced PDF Page: A standard 8.5" x 11" single-spaced page (approx. 500 words) produces approximately 650 to 750 tokens.

3. Multilingual Multipliers

Because tokenizers dedicate a majority of vocabulary entries to English, non-Latin alphabets suffer from byte fallback penalties:

Language ScriptTokens Per 1,000 WordsCost Multiplier vs English
English~1,330 tokens1.0x (Baseline)
Spanish / French / German~1,500 - 1,700 tokens1.1x - 1.3x
Russian (Cyrillic)~2,200 - 2,600 tokens1.7x - 2.0x
Japanese / Chinese / Korean~2,400 - 3,200 tokens1.8x - 2.4x