Why does Japanese text cost 2-3x more LLM tokens than its English translation?
answer
- same meaning, different token count
- vocabularies learned mostly from English
- tokens per unit of text
- rare characters fall back to bytes
- measure per locale, not per blog post
basics
~20 sTokenizer vocabularies are dominated by English text, so common English words become single tokens while Japanese fragments into short subword or byte pieces. The same meaning then consumes more tokens, more money, and more of the context window.
solid answer
~50 sThe measure is **fertility**: the average number of tokens a tokenizer emits per unit of text. Because training corpora and vocabularies skew English, English prose lands near four characters per token while Japanese, Thai, Hindi and other non-Latin scripts often approach one token per character, with rare characters falling back to several byte-level tokens each. Three consequences follow. Cost: a per-conversation price model calibrated on English is wrong by a multiple for Japanese traffic. Capacity: the *effective* context is smaller, so the same conversation history fits fewer turns. Latency: more tokens means more prefill and more decode. It is also an equity problem — users of some languages pay more and get less usable context for identical content. Recent large-vocabulary tokenizers have narrowed the gap but not closed it, so measure fertility per locale on your real corpus rather than assuming a published figure.
go deeper
Know that the same sentence in different languages does not cost the same number of tokens, and that non-English text usually costs more, so cost estimates cannot be copied across locales.
Explain fertility as tokens per unit of text, why English-dominated vocabularies cause it, and how byte-level fallback makes rare characters expensive. Be able to name typical ratios for prose, code and random identifiers.
Show the production consequence: cost models and history-trimming policies calibrated on English silently misbehave for other locales. Demonstrate that you measure fertility per locale on real traffic and act on the result.
Own it as a pricing and fairness question — per-locale unit economics, effective-context parity across markets, and tokenizer fertility as a first-class criterion in model selection alongside price and quality.
## Fertility, defined Fertility is the average number of tokens a tokenizer produces per unit of text — per word, per character, or per unit of meaning, depending on which comparison you want. It is the single number that explains why two texts saying the same thing bill differently. A tokenizer learns its vocabulary from a corpus; sequences that appeared often enough become single vocabulary entries, and everything else is assembled from smaller pieces. Whatever was rare in that corpus is expensive at inference time, forever. ## Why non-English text pays more Pretraining and tokenizer-training corpora are overwhelmingly English and, more broadly, Latin-script. So English function words, common stems and frequent suffixes each occupy a vocabulary slot, and a paragraph of English prose lands around 3.5–5 characters per token. Japanese gets much less of that treatment: fewer of its character sequences earned dedicated entries, so text splits into short pieces, often near one token per character. Worse, modern tokenizers are byte-level, meaning any character with no learned merge decomposes into its UTF-8 bytes — and a single Japanese character is three bytes, so an unlucky rare character can cost three tokens by itself. The comparison people actually feel is not per character but per *meaning*. Japanese writes densely: a sentence that takes eighty English characters may take thirty Japanese characters. Even so, the token count typically lands two to three times higher, because the per-character ratio moves against Japanese by more than the density moves for it. ## The concrete failure this produces Consider a customer-support product that priced itself per conversation using the ~4-characters-per-token rule and a sample of English transcripts. Ship the same product in Japan and each conversation carries two to three times the tokens for the same exchange. Gross margin on the Japanese tier collapses; nothing in the code changed and no alert fires, because the system is behaving exactly as built. The same arithmetic hits capacity: history-trimming tuned so that thirty English turns fit will hold roughly ten to fifteen Japanese ones, so the assistant starts forgetting earlier in the conversation for one set of users than another. ## Content type matters as much as language Fertility is not only a language property. Measure a mixed corpus and you get four different regimes. English prose sits near four characters per token. Ordinary source code sits near 2.5–3.5 — punctuation, operators and indentation runs are their own tokens, and `camelCase` or `snake_case` identifiers split at the case or underscore boundary. Minified JavaScript or TypeScript is worse than readable code, because mangled one- and two-character names share no learned merges. And a column of UUIDs, hex digests or base64 blobs is the floor: random character sequences have no statistical structure to compress, so they land near one to two characters per token. A retrieval system that stuffs raw identifiers into context is paying near the worst possible rate for text that carries little meaning. ## Why this is an equity issue, not just a billing quirk If price is per token, speakers of under-represented languages pay more for the same service, wait longer for the same answer, and get a smaller effective context. That asymmetry sits below the product, in the tokenizer, and no amount of prompt tuning removes it. The honest framing in an interview is that the gap is a known and measurable property of the vocabulary, that it has narrowed as vocabularies grew from tens of thousands to hundreds of thousands of entries and as multilingual data share increased, and that it has not disappeared. ## What to actually do Measure, do not assume. Take a representative sample per locale and per content type from your own traffic and compute characters per token for each; that table, not a blog-post constant, is what your cost model and your budgeting code should use. Price and forecast per locale. When choosing between models, include tokenizer fertility on your languages in the comparison — a model with a slightly higher per-token price can be cheaper end to end if it segments your text more efficiently, and the reverse trap is just as common. Where content is machine-generated, reduce fertility at the source: send identifiers by reference rather than pasting them, strip boilerplate, and avoid re-sending large blobs each turn. Finally, budget in tokens everywhere, so that the fertility difference is absorbed by the budget rather than discovered as a production incident.
- How does the same effect show up on code and machine-generated identifiers?Fertility is a content-type property too. Ordinary source code runs about 2.5–3.5 characters per token because punctuation and case boundaries split identifiers; minified code is worse, since mangled short names share no learned merges. UUIDs, hex digests and base64 are the floor at roughly one to two characters per token, because random sequences have no structure to compress. Pasting identifier columns into context pays the worst rate for the least meaning.
- Has the multilingual gap narrowed with newer models, and does that change your advice?It has narrowed. Vocabularies grew from tens of thousands to hundreds of thousands of entries and multilingual data share rose, so recent tokenizers segment non-Latin scripts more efficiently than earlier ones. The gap has not closed, and it varies by model family and by script. The advice is unchanged: measure fertility on your own corpus per locale for the specific model you deploy, and re-measure when you switch, rather than carrying a remembered ratio.
- How would you fold fertility into a model-selection decision?Compare end-to-end cost per unit of work, not headline price per token. Run a representative sample of your real traffic through each candidate's tokenizer, multiply by that model's input and output prices, and add the effective-context consequence — a fertile tokenizer shrinks how much history fits, which can force extra summarization calls. A cheaper per-token model that segments your languages badly frequently loses on the total.
saying these in an interview costs you the question
- Claims token cost depends only on character count
- Assumes every language hits about four characters per token
- Thinks translating to English before sending is a free fix
- Treats fertility as a language issue only, ignoring code and identifiers
- Says larger vocabularies have eliminated the multilingual gap