What does an LLM gain and lose by reasoning in words rather than latent state?
answer
- words are a narrow channel
- high-dimensional state, one discrete token
- what you can read you can check
- legible trace versus richer internal state
- latent thought keeps compute, loses the window
basics
~20 sVerbalizing forces each reasoning step through one discrete token, discarding most of the model's richer internal state - but it produces a trace that can be read, checked, cached, edited and monitored. Latent reasoning keeps the bandwidth and gives up the audit surface.
solid answer
~50 sEmitting reasoning as text is an information bottleneck: at each step a high-dimensional hidden state is collapsed into a single token drawn from a fixed vocabulary. In principle that throws away a lot, and it is why research explores keeping reasoning in continuous space - feeding hidden states back as inputs rather than decoding them, or using filler and pause tokens to buy computation without content. What verbalization buys is everything downstream of having a readable artifact: a human or an automated checker can inspect intermediate values, a verifier can score individual steps, a wrong step can be edited or retried, and the trace can be logged for incident review. There is also a governance argument that keeping reasoning in natural language preserves a monitoring surface that latent computation would remove. As of mid-2026 latent approaches remain largely research-stage; production systems reason in tokens, with the honest caveat that a readable trace is not automatically the true cause of the answer.
go deeper
Know that today's reasoning steps are ordinary text tokens, and that being text is what lets anyone read, check or log them.
Explain the bottleneck: each step compresses a large hidden state into one discrete token, losing information but producing an artifact that verifiers and humans can act on.
Weigh the tradeoff in operational terms - inspection, step-level verification, editing and replay, incident logs - and be honest that a legible trace is not a guaranteed faithful one.
Own this as an architectural and governance position: where reasoning lives determines what you can audit and monitor. State that latent approaches remain research-stage as of 2026 and that you would not build a compliance story on an inspectable trace you cannot prove is faithful.
## Two channels for the same computation A model's reasoning has to happen somewhere. In chain-of-thought it happens in the output channel: each step is decoded into a token, appended to the context, and read back by the next step. The alternative is to keep the reasoning inside the network - to let the model iterate on its own continuous representations without ever turning them into words. Both are ways of buying serial computation. They differ in what survives each step. ## The bottleneck argument for latent reasoning A hidden state is a vector with thousands of dimensions, carrying a graded, ambiguity-preserving representation of where the model is in the problem. Decoding turns that into one token from a vocabulary of perhaps a hundred thousand entries - a handful of bits. Most of the state is discarded, and what remains is forced into a single committed choice. That commitment is not neutral. It collapses a distribution over possible continuations into one branch, and everything after it is conditioned on that branch. Research on continuous or latent reasoning - feeding the model's own hidden representations back as the next input instead of decoding them, sometimes described as reasoning in a continuous thought space - is motivated by exactly this: keep the model's uncertainty alive across steps, and avoid paying for words when the words are not the point. Related work on filler or pause tokens tests the weaker version of the idea: can extra compute alone help, with no meaningful content emitted? Results there are real but modest and task-dependent, which itself tells you something - a good chunk of chain-of-thought's benefit comes from the content of the steps, not just from the passes they buy. ## What the words buy The case for verbalization is not that words are the best representation for computation. It is that a computation you can read is a computation you can operate. **Inspection.** A wrong answer in a dosing calculation is a mystery; a wrong answer with a visible conversion step is a bug you can point at. Intermediate values are only checkable if they exist as values. **Verification.** Because steps are discrete text, external verifiers can score them - a unit checker, a calculator, a schema validator, a judge model, a test suite. Process-level supervision of any kind needs a process you can see. **Intervention.** A token trace can be truncated, edited, branched from, or retried at a specific step. You can re-run just the tail. Latent state offers no comparable handle. **Reuse.** Tokens are cacheable and portable. A trace can be logged, replayed, attached to an incident report, shown to an auditor, or turned into training data for a smaller model. A continuous internal state is none of those things. **Monitoring.** There is a governance argument, made prominently by safety researchers across several labs, that reasoning conducted in natural language gives a rare window into a model's process - and that architectures which move reasoning out of tokens would close it. Whether or not you weight that heavily, it is a live consideration in how frontier systems are being built, and worth naming in a senior interview. ## The honest caveats Two, and a strong candidate volunteers both. First, readability is not faithfulness. A trace can be well formed, plausible and not the actual cause of the model's answer. Verbalization gives you an artifact to monitor; it does not guarantee the artifact reflects the computation. Treat the trace as evidence, not testimony. Second, the bottleneck may be doing useful work. Forcing each step through a discrete symbol imposes a kind of discipline: it commits the model to a definite intermediate claim that later steps must live with, and it keeps steps in the distribution the model was trained on. Unconstrained continuous iteration has fewer guardrails, and part of why latent approaches have not displaced token reasoning is that they are harder to train stably and harder to evaluate. ## Where this stands in 2026 Production systems reason in tokens. The practical variation is in where those tokens are shown, how many of them are spent, and what verification runs over them - not in whether reasoning is verbalized at all. Latent and continuous-thought approaches remain an active research direction with promising results on specific tasks, notably ones with heavy search or backtracking, and no broad displacement of chain-of-thought. ## Answering it well Name the bottleneck honestly - words are a narrow channel and the model has more state than it can emit. Then argue that the operational value of a legible trace (verification, intervention, logging, monitoring) currently outweighs the lost bandwidth, and that the field is exploring the alternative rather than having settled it. Close with the caveat that a legible trace is not a guaranteed faithful one, so you would still verify outcomes, not just read prose.
- If latent reasoning preserved more information, why has it not replaced token-based chains?Three reasons. It is harder to train stably, since there is no discrete supervision target for each step. It is far harder to evaluate, because there is no step to score. And it removes the inspection, verification and monitoring surface that production systems rely on. The gains so far are task-specific rather than general, so the tradeoff has not favoured it outside research.
- Does a legible reasoning trace prove the model reasoned that way?No. A trace can be fluent, well structured and not the actual cause of the final answer; models can produce post-hoc rationalisations. Use the trace as a debugging and monitoring artifact, and confirm behaviour with outcome checks and targeted perturbations rather than treating the prose as testimony about the model's internals.
- What is the practical argument for keeping reasoning in natural language at all, if it costs bandwidth and tokens?Operability. A token trace can be inspected by a human, scored by a verifier, edited and re-run from a specific step, cached, logged for incident review and monitored for concerning content. None of that is available for reasoning that stays in activations. For any system that has to be debugged or audited, that operability is currently worth more than the lost representational bandwidth.
It is the difference between an engineer working entirely in her head and one writing on a whiteboard. The whiteboard loses nuance and slows her down, but it is the only reason a colleague can spot the error on line three.
saying these in an interview costs you the question
- Assuming latent reasoning is already standard in production
- Treating a readable trace as proof of the true computation
- Believing the token bottleneck carries no cost at all
- Ignoring that legibility is what makes verification possible
- Presenting a contested research question as settled