Twenty cell formulas contain the identical pure subexpression — what may the engine do that it could not otherwise?
answer
- the same inputs, the same value
- compute once, use many times
- dropping evaluations must drop nothing
- watch for a draw or a write
- holding a value costs memory
basics
~20 sEvaluate that subexpression once for the pass and feed the single value to all twenty formulas. Purity makes the elimination sound: the same inputs give the same value, and skipping nineteen evaluations skips no effect.
solid answer
~40 sIt can eliminate the repetition — evaluate the shared expression once and substitute the resulting value into all twenty formulas. This is sound for two reasons, and both come from purity. The expression yields the same value for the same inputs, so nothing is lost by computing it once; and it produces no observable effect, so removing nineteen evaluations removes nothing anyone can see. Both halves matter. If the expression drew a fresh random value, the first condition fails and the twenty cells were never meant to agree. If it appended an audit line, the second fails and the sheet loses nineteen lines. The trade is memory for time: the single value must stay live while all twenty consumers need it, which is not automatically a win for a very cheap expression.
code
pseudocode · 8 lines# before: the same pure subexpression twice
A = (base * rate) + 1
B = (base * rate) * 2
# after: one evaluation, two uses
t = base * rate
A = t + 1
B = t * 2go deeper
Recognise that repeated identical work over unchanging inputs can be computed once, and that this is exactly what hoisting a repeated calculation out of a loop body does.
Give both halves of the justification — same inputs give the same value, and skipping an evaluation skips no effect — and name an expression that fails each half.
Argue the trade rather than the rule: a large value held live across a pass, or a failure that now surfaces at a different point, can outweigh the work saved.
Set the expectation for the codebase. Whether a repeated call can be hoisted safely is decided by whether teams keep computation and effect separate, which is a standard, not a per-review judgment.
## The transformation When the same expression appears in many places with the same inputs, an engine can compute it once, bind the result to a temporary, and let every occurrence read that temporary. Twenty evaluations become one. This is one of the oldest optimisations there is, and it is entirely a consequence of purity — which is why it is worth being able to justify rather than merely name. The question an engine has to answer before applying it is not *does this text look the same?* but *is this expression a function of inputs that are the same here, and does skipping an evaluation lose anything?* ## The two halves purity supplies 1. **Same inputs, same value.** If the expression is a function of its declared inputs alone, then all twenty occurrences — reading the same cells, during a pass in which those cells do not change — must produce the same value. Nineteen of the evaluations were guaranteed to be redundant before any of them ran. 2. **No effect is lost.** Evaluating a pure expression leaves no trace. Nobody can count evaluations, because there is nothing to count: no line written, no counter moved, no external call made. Removing nineteen of them is therefore invisible. Either half alone is insufficient, and the failure modes differ: | What the shared expression does | Safe to evaluate once? | Why | |---|---|---| | Arithmetic over cells the pass does not write | Yes | Both halves hold | | Draws a fresh random value each time | No | It was never a function of its inputs; the twenty results were meant to differ | | Appends a line to an audit log | No | Nineteen effects vanish, and the log is what somebody reads | | Reads a cell that is edited part-way through the pass | Only within a region where that input is stable | The inputs are not the same on both sides of the edit | ## Inside a pass, not across passes There is an important boundary here. Within one pass, the engine knows the inputs cannot change, so the reused value cannot go stale: it needs no key, no invalidation rule and no eviction policy. Holding the value so that a *later* pass can skip the work is a different proposition, which reintroduces all three of those questions — what identifies the entry, when it stops being valid, and how much of it you are willing to keep. Answer only what was asked: the elimination described here is a plan-level rewrite for one pass. ## The trade you are actually making Eliminating repeated work is not free, and describing it as pure win is a weak answer: - **The value must stay live** from the point it is computed until the last of the twenty consumers has read it. That is memory held for the span of the pass, and for a large intermediate value that can be the dominant cost. - **Recomputation is sometimes cheaper than retention.** For a trivial expression over values already at hand, computing it twenty times may cost less than keeping one result alive and reachable. - **The rewrite changes where a failure surfaces.** If evaluating the expression can fail, moving it to a single earlier point means the failure appears there, once, rather than at the first consumer that needed it. - **The consumers now share a value, not a computation.** That is exactly what was wanted here, but it is also why the random-value case is a bug rather than a speed-up. ## Recognising the pattern by hand Engineers apply the same transformation manually all the time — hoisting a repeated lookup out of a loop body, binding a repeated call to a local name before a branch. The reasoning is identical, and so is the check: *would running this once instead of many times change anything a reader of the system can observe?* When the expression is pure, the answer is no and the hoist is safe by construction. When it is not, the hoist is a behaviour change dressed as a refactoring, and it is one of the most common ways a well-intentioned clean-up introduces a defect. ## How this is asked Interviewers rarely ask for the name of the optimisation; they ask why it is allowed, and then supply an expression that breaks it. The strong answer states both halves of the justification up front, so that when the follow-up introduces a random draw or a log write, you already have the vocabulary to say which half failed and what the sheet now gets wrong.
- Some of the twenty formulas run before an input cell is edited and the rest after. What breaks?The same-inputs condition. A shared value is valid only across a region in which its inputs do not change, so the engine must either treat the edit as beginning a new pass or invalidate the shared value at that point. Carrying one value across an input change silently mixes two different answers into one sheet.
- How is reusing a value within one pass different from keeping it for later passes?Within a pass the inputs cannot change, so the reused value cannot go stale and needs no key, no invalidation and no size bound. Keeping it for later passes reintroduces all three questions at once, and each of them is a source of wrong answers rather than merely slow ones.
saying these in an interview costs you the question
- Says a value may be reused whenever the text of the expression matches.
- Believes a freshly drawn value is safe to compute once and share.
- Thinks eliminating repeated work is free, ignoring the value held live.
- Assumes an expression containing a write can be evaluated once unnoticed.