skip to content

Why can systematic sampling of every 7th day of logs give a biased estimate?

level: middleimportance: should knowfreq 33%

answer

  1. every k-th record after a random start
  2. the list has its own rhythm
  3. k lines up with the cycle length
  4. seven Mondays, no Saturdays

basics

~20 s

Taking every 7th day always lands on the same weekday. Systematic sampling picks every k-th unit after a random start, so when the ordering carries a cycle of period k, the sample sees one phase of that cycle and misses the rest.

solid answer

~50 s

Systematic sampling means choosing a random start between 1 and `k` and then taking every `k`-th unit thereafter, with `k` roughly `N/n`. It is cheap, spreads the sample evenly across the list, and when the list is ordered on something related to the outcome it behaves like implicit stratification and can beat a simple random sample. Its failure mode is **periodicity**: if the ordering has a cycle whose period equals `k` or divides it, every selected unit sits at the same point in that cycle. Daily log files with `k = 7` hand you seven Mondays, or seven Sundays, and whatever weekday effect exists in traffic or conversion is baked into the estimate as if it were the average. The fixes are to pick a `k` coprime to the known cycle, to randomise the start within each cycle, or to abandon the interval and stratify explicitly on day of week.

go deeper

for a junior

Know the recipe: a random start, then every k-th unit, with k set to the population size divided by the target sample size.

for a middle

Explain periodicity concretely — k = 7 on daily records means one weekday forever — and name a fix, such as choosing an interval coprime to the cycle.

for a senior

Show that you inspect how the frame is ordered before trusting a systematic draw, and that you would rather stratify on a known cycle than tune the interval around it.

for a principal

Weigh the operational simplicity of an every-k-th rule any engineer can implement against a design whose validity silently depends on nobody re-sorting the source table.

## The mechanics Systematic sampling selects a sampling interval `k`, draws a random start `r` uniformly from `1` to `k`, and then takes units `r`, `r + k`, `r + 2k`, and so on. With `N` units on the frame and a target of `n`, you set `k` to about `N/n`. It is popular because it is trivial to execute — a query, a cursor, a person counting down a shelf — and needs no random number per unit, only one at the start. The random start is what makes it a probability design: each unit's inclusion probability is `n/N`, so the sample mean is unbiased for the population mean across repeated samples. Start deterministically at record 1 and you have no randomisation at all; whatever pattern lives in the ordering is locked in permanently. ## Why the ordering of the frame is the whole story Unlike a simple random sample, systematic sampling never draws two neighbouring units and always draws units exactly `k` apart. That makes the design's behaviour a function of how the list is sorted. **Ordering unrelated to the outcome** — say, account IDs assigned at signup with no meaning — makes systematic sampling behave much like a simple random sample. **Ordering related to the outcome, without cycles** — say, customers sorted by tenure — makes it behave better than a simple random sample. Taking every 20th customer from a tenure-sorted list guarantees the sample spans the full tenure range in proportion, which is exactly what stratifying on tenure would achieve. This is called *implicit stratification*, and it is the reason experienced survey designers sort the frame deliberately before a systematic draw. **Ordering with a cycle** is where the design breaks. ## Periodicity, concretely Daily log files sit in date order, and human activity is weekly. Choose `k = 7` and every selected date falls on the same weekday: seven Mondays, or seven Saturdays, depending only on the random start. If Saturday traffic is half of Tuesday traffic, an estimate built on seven Saturdays is not a noisy estimate of the weekly average — it is a precise estimate of the wrong quantity. More samples do not help; they converge on the Saturday mean. The damage is not limited to `k` exactly equal to the period. Any `k` that shares a large common factor with the cycle concentrates the sample on few phases: with a weekly cycle, `k = 14` or `k = 21` is just as bad as `k = 7`. What you want is `k` **coprime** to the period. With a 7-day cycle and `k = 10`, successive selections shift by 3 weekdays each time and walk through all seven positions before repeating. The same trap appears wherever the frame has a rhythm: hourly server metrics sampled every 24th record, shift-ordered production records sampled at the shift length, a customer list ordered household-by-household where every fourth row is the head of household. ## The random start is not a fix A common misconception is that randomising the start neutralises periodicity. It does not. The random start decides *which* phase of the cycle you get; it does not give you more than one phase. Averaged over all possible starts the estimator is unbiased, which is a statement about the long run over hypothetical repetitions — but you only run the sample once, and that single realisation carries the full weekday effect. A genuine fix randomises *within* each cycle rather than once at the beginning: pick one random day inside each week rather than one random start followed by fixed steps. At that point you have stopped doing systematic sampling and started doing stratified sampling with the week as the stratum, which is exactly the right move. ## Variance is awkward too A subtler cost: a single systematic sample gives no unbiased estimate of its own sampling variability, because the design draws one cluster of units out of `k` possible clusters — there is only one "draw" in the design sense. The standard workaround is *repeated systematic sampling*: run several independent systematic samples with smaller intervals and different random starts, which restores replication. Many practitioners instead compute variability as if the sample were simple random, which is an approximation that is optimistic when implicit stratification helped and dangerously wrong when periodicity struck. ## What to do in practice Before accepting a systematic draw, look at how the source is sorted, ask whether the domain has a natural rhythm (weekly, daily, shift-based, billing-cycle), and check whether `k` shares factors with it. Prefer sorting the frame on something useful and choosing a `k` coprime to any known cycle. If the cycle is important enough to worry about, do not try to dodge it with arithmetic — make it a stratum and sample inside it.

  • When is systematic sampling actually better than a simple random sample?
    When the frame is sorted on something related to the outcome and carries no cycle. Sorting customers by tenure and taking every 20th spreads the sample evenly across the whole tenure range, which acts like implicit stratification and beats a simple random draw that could, by chance, over-represent new customers.
  • How would you choose k to avoid a weekly cycle in daily data?
    Pick a k coprime to 7 — 5 or 10, for example — so successive selections walk through every weekday instead of repeating one. Note that 14 and 21 are as bad as 7, since they share the cycle's factor. Better still, treat day of week as a stratum and sample inside each, which guarantees balance rather than relying on arithmetic.
  • What does the random start actually buy you?
    It makes each unit's inclusion probability n/N, so the estimator is unbiased over repeated samples, and it stops the design from being fully deterministic. What it does not do is break periodicity: it only selects which phase of the cycle your one realised sample lands on.

It is like judging a restaurant by visiting every seventh day: you only ever see it on a Tuesday, and never on the Saturday night rush.

saying these in an interview costs you the question

  • Treats systematic sampling as identical to simple random sampling
  • Thinks a random start removes periodicity bias
  • Never checks how the source frame is sorted
  • Assumes evenly spaced automatically means representative
  • Believes only k equal to the cycle length is dangerous

context