skip to content

Why do LLMs miscount letters and judge 9.11 as larger than 9.9?

level: middleimportance: must knowfreq 58%

answer

  1. the model reads IDs, not letters
  2. spelling is recalled, not inspected
  3. digit grouping breaks place value
  4. version-number priors also mislead
  5. move exact work into code

basics

~10 s

The model never sees characters. Text arrives as subword token IDs, so spelling is a memorized association rather than something readable, and long numbers split into chunks whose boundaries differ between similar-looking values.

solid answer

~50 s

Tokenization puts a layer between the text and the model: a word may arrive as one or two vocabulary IDs with no explicit list of its letters, so any character-level task — counting occurrences, reversing a string, taking the nth character — requires reconstructing spelling from statistical memory instead of reading it. Numbers are the same problem with higher stakes. Digit strings are segmented into chunks, and two visually similar numbers can decompose differently, so a comparison becomes a fuzzy match over pieces rather than arithmetic on magnitudes. For decimal comparisons like 9.11 versus 9.9, tokenization is one cause and training priors are another: 9.11 genuinely is later than 9.9 as a software version or a date, and that pattern is abundant in the corpus. The engineering answer is not to prompt harder but to move the work: emit structured values and compare or compute them in code.

go deeper

for a junior

Know that the model sees subword tokens rather than individual characters, which is why tasks like counting letters in a word or comparing long numbers go wrong even when the model writes them correctly.

for a middle

Explain the mechanism: spelling is a memorized association because token IDs hide characters, and digit grouping destroys place-value alignment. Be able to say why the 9.11 case also has a training-prior component.

for a senior

Show the engineering response — recognize the class of task that hits the representation mismatch and move it out of the model into structured output plus code, rather than escalating model size or prompt length.

for a principal

Own the boundary decision across the system: which operations must be deterministic, where the model is allowed to be the source of truth, and how failures surface as validation errors rather than plausible wrong answers.

## What the model actually receives A language model does not read text. The tokenizer converts a string into a sequence of integer IDs drawn from a fixed vocabulary, and every layer of the network operates on the embeddings of those IDs. Nothing in that pipeline hands the model a character array. Whatever the model knows about which letters compose a token, it learned indirectly — from spelling puzzles, hyphenations, typos and letter-by-letter text that happened to appear during training. That knowledge is real but statistical, and it degrades exactly where you would expect: rare words, unusual casing, long strings, and any task that requires enumerating characters in order rather than recalling a fact about the word. ## Why character tasks fail on words the model spells fluently The apparent paradox — a model writes a word perfectly yet miscounts its letters — dissolves once you see the two operations are different. Producing the word is emitting one or two token IDs it has seen a million times. Counting a letter inside it requires decomposing an ID into characters, iterating, and maintaining a tally, none of which the architecture does natively; it must be simulated in the forward pass from memorized associations. The same explains failures at reversing a string, extracting the nth character, judging string length, and counting words or syllables. Related tasks that look similar but are easier — does this word contain the letter x, does it start with a given prefix — succeed more often because they need a single recalled fact rather than an ordered traversal. ## Why numbers are a special case Digits get grouped. Depending on the tokenizer, a long number may split into multi-digit chunks or into individual digits, and the split depends on the exact character sequence. That has two consequences. First, place value is not preserved by the segmentation: the same digit can appear in differently-sized chunks in two numbers, so column alignment — the thing that makes hand arithmetic work — is not available for free. Second, similar-looking numbers can decompose incomparably, which is why long multiplication, exact division and precise comparison degrade as digit count grows while short arithmetic looks fine. Some tokenizer designs deliberately restrict number tokens to single digits or fixed three-digit groups precisely to make the segmentation regular and improve arithmetic, which is good evidence that the segmentation is causal. ## The 9.11 versus 9.9 case, honestly This example is popular because it is overdetermined, and a good answer says so. Tokenization contributes: the two strings do not decompose into pieces that can be compared as magnitudes, so the model is doing something closer to string matching than to arithmetic. But the training distribution contributes at least as much. In software versions, chapter numbers, dates and verse references, 9.11 genuinely comes after 9.9 — and those contexts are extremely common in text. The model has learned a strong prior that the token pattern *9 . 11* follows *9 . 9*, and that prior can override the decimal reading unless the context makes 'these are decimal numbers' unmistakable. Claiming tokenization alone explains it is overconfident; claiming tokenization is irrelevant is equally wrong. ## What to do about it in production Stop asking the model to do the thing it is bad at. Character-level manipulation and exact arithmetic should be delegated: have the model identify or extract the values, emit them as structured output, and let ordinary code do the counting, comparison, sorting or calculation. Where a tool or code-execution path exists, routing the computation there converts a probabilistic answer into a deterministic one. When delegation is unavailable, several mitigations help partially and none is a fix. Presenting a string with its characters separated makes each character its own token and turns a decomposition task into a reading task. Asking for explicit step-by-step working improves multi-digit arithmetic because it externalizes intermediate values into the context, but it does not make the underlying operation exact. Naming the domain explicitly — 'treat these as decimal numbers, not version strings' — pushes against the wrong prior for the 9.11 case. Sampling several answers and taking the majority raises accuracy at a cost, without guaranteeing correctness. ## How to talk about it The strongest framing in an interview is diagnostic. These failures are not evidence that the model is stupid about maths; they are evidence of a representation mismatch between what the user sees, characters and digits, and what the model consumes, subword tokens. Recognizing the class of task that hits the mismatch — anything requiring an exact operation over the surface form of text — tells you immediately to move that step outside the model rather than to escalate to a larger one, which is often the cheaper and more reliable engineering call.

  • If tokenization causes it, why do some models handle multi-digit arithmetic much better than others?
    Partly tokenizer design and partly training. Restricting number tokens to single digits or fixed three-digit groups makes segmentation regular so place value survives, and heavy training on arithmetic and step-by-step working teaches the model to externalize intermediate results into the context. Both raise accuracy substantially. Neither makes the operation exact, which is why exact arithmetic still belongs in code for anything that matters.
  • Does chain-of-thought reasoning fix character-counting tasks?
    It helps and does not fix. Writing the word out letter by letter converts a decomposition problem into a reading problem, because separated characters become separate tokens, and keeping a running tally in the visible context reduces the working-memory burden. Errors still occur on long or unusual strings, and the extra tokens cost money and latency. For anything with a correctness requirement, delegate to code instead.
  • How would you design around this in a product that must compare user-supplied numbers?
    Never let the comparison happen inside the model. Have it extract the values into a typed structured output, validate them against a schema, and perform the comparison, rounding and formatting in application code. Log any value the model failed to extract cleanly rather than letting a free-text answer through. This turns a probabilistic step into a deterministic one and makes the failure mode a visible parse error instead of a silently wrong answer.

saying these in an interview costs you the question

  • Says the model reads text character by character
  • Blames model size, so a bigger model would fix it
  • Claims prompting alone makes arithmetic exact
  • Attributes 9.11 versus 9.9 purely to tokenization
  • Thinks a spelling failure means the word is unknown

context