skip to content

A trailing check field fails on a received configuration blob: what can the receiver do about it, and what can it not?

level: seniorimportance: should knowfreq 45%

answer

  1. a verdict, not a map
  2. the whole covered unit is suspect
  3. no redundancy, no repair
  4. refetch needs a back channel
  5. coverage spans computation to verification

basics

~20 s

A failed check proves only that the covered bytes differ from what the sender summed. It names no position, quantifies no damage, repairs nothing, and leaves exactly one recovery path: discard the whole covered unit and obtain it again.

solid answer

~40 s

The check value is a function of the entire covered span, so a mismatch is a verdict about that whole span and nothing finer. It does not say which bytes changed, how many bits flipped, or whether the fault was on the link, in the sender's buffer or in the receiver's. Because a detection-only check carries no redundancy beyond the verdict, there is nothing to reconstruct from, so the receiver discards the unit. Recovery then depends on conditions outside the check: a back channel, a sender that still holds the data, and a retry budget. When any of those is missing the data is simply lost, and the right engineering response is a documented fallback — refuse the blob and keep the last known-good configuration — plus an alarm on the failure rate.

go deeper

for a junior

Remember the basic division of labour: a check field says whether the data is intact, and the response to a failure is to throw the unit away and ask again rather than to patch it.

for a middle

Explain why no repair is possible — the value is one verdict over the whole covered span and carries no redundancy — and why it names neither a position nor a count.

for a senior

This is your tier: reason about rates rather than single events, about the span the check actually covers, and about whether a refetch is even available on this path before promising reliability.

for a principal

Own the placement decision. Where the check is computed and verified determines what it can possibly cover, and whether the system needs detection plus retransmission or redundancy that repairs without a back channel.

## What a mismatch proves The trailing field is computed over a defined span of bytes. The receiver recomputes over the span it received and compares. A mismatch supports exactly one conclusion: **at least one covered byte is not what the sender summed**. That is a genuine and valuable fact — it is the whole reason the field is there — but it is a single bit of information about a span that may be thousands of bytes wide. ## What it does not prove - **Where.** The value is a function of the whole span; it carries no map. A 4 KB block that fails tells you 4 KB are suspect. - **How much.** One flipped bit and a wholly replaced payload produce the same verdict. - **Which side.** The corruption could have happened on the link, in the sender's memory after the data was assembled, or in the receiver's buffer before verification. The check cannot distinguish them. - **That a pass is safe.** A passing check says the covered bytes agree with the value computed at the sender — and only within the span between computation and verification, and only for the corruption patterns that check is sensitive to. ## The only recovery a detection-only check offers There is no repair. Repairing requires enough redundancy to reconstruct the original bytes, which is a different and heavier family of codes; a check field deliberately does not carry it. So the receiver's move is fixed: **discard the whole covered unit and fetch it again.** That move has preconditions the check itself cannot supply: 1. a **back channel** on which to ask; 2. a **sender that still holds the data** — a source that has already discarded it cannot answer; 3. a **retry budget**, because a permanently degraded path will fail every attempt and turn retries into a load amplifier. When those hold, detection plus retransmission is a complete and very cheap reliability story. When they do not — a one-way push, a device with no upstream link, a source that streams and forgets — detection tells you the data is gone without giving you any way to get it back, and the honest design answer is to carry repair redundancy instead, or to accept the loss explicitly. ## The span a check actually covers A check protects only the interval between where it is computed and where it is verified. Corruption before the computation is summed in as though it were intended, and the receiver sees a perfect match over corrupt data. Corruption after verification is equally invisible. A blob that is verified at an intermediary, rewritten and re-checked is covered in two separate intervals with a gap between them — and a fault inside that gap produces a value that agrees at every step. The practical consequence is a design rule rather than an algorithm choice: compute the check as close to the true producer as you can, and verify as close to the true consumer as you can. Widening the covered span is worth more than widening the check field. ## Operating it | Signal | What it usually means | Reasonable response | |---|---|---| | one isolated mismatch | a transient on the path | discard, refetch, count it | | a rising mismatch rate | a systematic fault: degrading link, marginal connector, a producer writing past a buffer | alarm and investigate the path, not the payload | | mismatches only on large units | length-correlated corruption, or a span the check does not fully cover | check the covered range before blaming the medium | | a check that has never once failed | plausible, but also the signature of verification being skipped | prove the verification path fires by corrupting a unit deliberately | The last row is the one experienced candidates raise unprompted. A check nobody verifies is indistinguishable from a check that always passes, and it is a surprisingly common defect: the field is written, carried and never read. ## In the interview The strong answer separates the verdict from the diagnosis. A failed check means discard and refetch; it does not mean the link is at fault, does not point at a byte, and does not license a repair attempt. The strongest answers then add the two systemic points: recovery needs a back channel and a source that still has the data, and a check only covers what lies between its computation and its verification.

  • What does a rising rate of check failures tell an operator that a single failure does not?
    A single mismatch says one unit was bad; a rate says the fault is systematic — a degrading link, a marginal connector, a producer writing past a buffer. Track the ratio and alarm on it. Watch for the opposite signal too: a check that never fails can mean it is never being verified.
  • What happens when corruption occurs in the sender's buffer before the check is computed?
    The value is computed over the already-corrupt bytes, so it matches and the receiver accepts corrupt data as good. A check covers only the interval between computation and verification, so the fix is to widen that interval — compute at the true producer, verify at the true consumer — rather than to widen the field.
  • When is discard-and-refetch not available, and what follows?
    On a one-way push, when the source has already discarded the data, or when the path is degraded enough that every retry fails. Detection then only tells you the data is lost. The design answers are to carry enough redundancy to repair in place, to persist the data at the source until acknowledged, or to accept the loss explicitly and fall back.

A smoke alarm tells you something is wrong somewhere in the house. It does not tell you which room, and it does not put the fire out.

saying these in an interview costs you the question

  • Expects the check value to locate the corrupted bytes
  • Thinks a detection-only check can repair one flipped bit
  • Believes a passing check covers corruption that happened before it was computed
  • Assumes a mismatch always means the link corrupted the data
  • Treats retransmission as free and always available
  • Reads a check that never fails as proof the path is clean