Multimodal Vision Transformers512px Tiles & Native Crops
Multimodal Vision Token & Image Cost Calculator
Convert image dimensions and resolutions into exact input token counts and dollar charges. Fully models OpenAI's 512×512 patch grid, Anthropic's 1568px bounding box, and Google's native tile crops.
Vision & Multimodal Image Token Calculator
Calculate exact vision tokens and costs based on official 512×512 tiling and downscaling specifications across OpenAI, Claude, and Gemini.
Image Dimensions & Detail
100 images
Calculated Vision Tokens (1920×1080px)Batch: 100 items
OpenAI (high)
1,105
Tile-based 512×512
Claude Vision
1,844
(Pixels / 750) downscaled
Google Gemini
258
Standard image crop
| Model | Tokens / Img | Cost / 1 Image | Cost / 100 Imgs |
|---|---|---|---|
| Claude Sonnet 5 (Anthropic) | 1,844 | $0.00553 | $0.5532 |
| Claude Haiku 4.5 (Anthropic) | 1,844 | $0.00184 | $0.1844 |
| GPT-6 Astra (OpenAI) | 1,105 | $0.00221 | $0.2210 |
| GPT-5.6 (OpenAI) | 1,105 | $0.00138 | $0.1381 |
| Gemini 3.8 Flash (Google) | 258 | $0.00019 | $0.0194 |
| Gemini 3.1 Pro (Google) | 258 | $0.00052 | $0.0516 |