skip to content

A feature column was smoothed with a window centred on each row; why could its value at that row's stamp not have been computed then?

level: seniorimportance: must knowfreq 57%

answer

  1. length and alignment are separate
  2. where the answer is written down
  3. half the records come from later
  4. only a right-edge stamp is knowable then

basics

~20 s

A centred window is stamped in the middle of the records it covers, so roughly half of them come from after that stamp. The value therefore encodes observations that had not happened at the moment the row describes.

solid answer

~50 s

Alignment is a separate choice from length: it decides which position each answer is written to, relative to the records it was computed from. A **trailing** window writes its answer at the right edge, so the value at a position uses nothing later than that position. A **centred** window writes it in the middle, so with a window of seven records the value at a row is computed from the three records after it. A **leading** window is worse still. Only the trailing form has the property people mean when they call a smoothed column safe to consume — and that is a property of the alignment, not of the statistic or the length. The defaults are not uniform across designs; some window surfaces are trailing unless told otherwise, some smoothing surfaces are centred because their purpose is description rather than prediction. So state the alignment rather than infer it from the answer looking reasonable.

code

pseudocode · 7 lines
pseudocode
# trailing: the answer at position t is written at the window's right edge
for t in positions:
    answer[t] = mean(x[t-6 .. t])        # 7 records, none later than t

# centred: the answer at position t is written in the middle of the window
for t in positions:
    answer[t] = mean(x[t-3 .. t+3])      # 7 records, three of them later than t

go deeper

for a junior

Learn that a window has two separate settings: how many records it covers, and which position its answer is written to. The second one decides whether later records are involved.

for a middle

Explain the three alignments and what each implies. Be able to say why a trailing window lags the data by about half its length, and why centring removes that lag by using records from the other side.

for a senior

Show the diagnosis in production terms: pick a row in the middle, compare the latest covered stamp against the row's own, and be explicit that nothing raises, no absent values appear, and the relationship looks better than it should.

for a principal

The angle is a standing rule. Argue for a team convention that any column consumed as knowable-at-its-stamp declares its alignment in the same place it declares its length, and that descriptive smoothings live in a separate namespace from pipeline inputs.

## Where the answer is stamped A window computation produces one answer per position, from the records the window covers there. **Alignment** is a separate choice from length: it says *which position the answer is written to*, relative to the span of records it was computed from. | alignment | the answer is written to | records used, relative to that stamp | could the value have existed at that stamp? | |---|---|---|---| | trailing | the window's right edge | all at or before it | yes | | centred | the middle of the window | roughly half of them after it | no | | leading | the window's left edge | all at or after it | no | The property in the last column is what people mean when they say a smoothed column is safe to consume. It belongs to the alignment, not to the statistic and not to the length. A centred mean and a centred median are both unsafe in this sense; a trailing mean over a window of two and one over a window of two hundred are both safe. One clarification, because the names mislead: a window anchored at the start is safe under this test **while its right edge is the current position**, because it too uses only records at or before the stamp. What is not safe is a statistic computed once over the whole sequence and then written back onto every row — that is the extreme case, where every row but the last carries information from after itself. ## Why centred smoothing exists, and why it leaks here A centred window is not a mistake. It is the better description of a local level *when you are describing a sequence after the fact*, because it introduces no phase shift. A trailing window's output lags the data by roughly half the window length, so a peak appears later in the smoothed column than it did in the raw one. Centring removes that lag by borrowing records from both sides. The borrowing is the whole problem when the column will be consumed as if it were a fact about its own stamp. With a window of seven records centred on each position, the value at a given row is computed in part from the three records after it. If that column is then lined up against an outcome and used to explain it, the explanation has access to observations that had not happened yet. Nothing in the data marks this. The column has the right length, the right type, no absent values in its interior, and a smoother, better-looking relationship with whatever you compare it to than the honest version would have had. **Leakage** is the general name for information reaching a row that did not exist at the moment that row describes, and a centred or leading window is one of the tidiest ways to manufacture it: one argument, no warning. ## The defaults are not the same everywhere Alignment is defaulted, and the defaults are not uniform across the designs in this family. Some window surfaces are trailing unless you say otherwise. Some smoothing surfaces are centred unless you say otherwise, precisely because their purpose is to describe a sequence rather than to build an input for something else. Two routines in the same program, both described as "a moving average", can differ on this. The rule that survives porting is to **state the alignment explicitly** wherever the output will be consumed as knowable-at-its-stamp. ## The edges hide it where you are most likely to look At the very first and very last positions a centred window has nothing on one side, so it quietly behaves like a leading or a trailing window there. Inspecting the head or the tail of the output is therefore exactly the inspection that will fail to find the problem. Look in the middle. ## The check 1. Choose a row in the middle of the data, not at either end. 2. Write down that row's own stamp. 3. Work out which records the window covered for that row, and write down the stamp of the **latest** one. 4. If that latest stamp is after the row's own, the value at that row could not have been computed at that moment. 5. Apply the check once per derived column rather than once per pipeline. One centred column among twenty trailing ones is enough to poison a result, and it will not announce itself. ## What to do instead If you want the lag-free picture for a chart or a report, use the centred form and say so in the caption — that is the right tool for describing what happened. If the column will be consumed as an input to anything that also consumes later outcomes, use the trailing form, accept the lag, and if the lag genuinely matters, shorten the window rather than recentring it. Where a tool exposes only a centred surface, the honest construction is to compute it and then write each answer at the position of the latest record it used, which converts it into a trailing window of the same length — with the lag back, because the lag is what the honesty costs.

  • Is a trailing window automatically safe, whatever its length?
    Under this particular test, yes: every record it covers is at or before the position the answer is written to, so the value could have been computed at that moment. Length changes how much history the value reflects and how sluggishly it responds, but not whether it looks forward. The failure mode being described here is about alignment alone.
  • What is the cost of switching a centred smoothing to trailing?
    Lag. A trailing window's output follows the data by roughly half the window length, so peaks and level shifts appear later in the smoothed column than in the raw one. That is the honest price of using only what was known. If the lag is unacceptable, shorten the window — do not recentre it, which trades a visible lag for an invisible error.
  • Why is inspecting the first and last rows a poor way to find this?
    Because a centred window is truncated at both ends: at the first positions it has nothing to its left and at the last nothing to its right, so it behaves there like a trailing or a leading window respectively. The rows that look most normal are the ends, and the rows that carry the problem are in the middle.

saying these in an interview costs you the question

  • Says a moving average is safe as an input because it only looks backward.
  • Assumes the alignment default is trailing in every design.
  • Thinks a short centred window is harmless because it only peeks a little.
  • Checks the first and last rows, where a centred window is truncated anyway.
  • Confuses where the answer is stamped with how long the window is.