skip to content

questions

5

What happens to unknown fields when an older reader decodes a newer producer's message and re-emits it?

level: middleimportance: must knowfreq 66%

answer

  1. decode then re-encode, not edit in place
  2. the reader's model is the filter
  3. drop, preserve, or reject
  4. a side buffer of unrecognised bytes
  5. both hops succeed, so nothing is logged

basics

~20 s

They are usually dropped. Decoding keeps only the fields the reader's own schema version names, and the output is rebuilt from that smaller in-memory value, so the newer producer's data disappears with no error anywhere. Preserving it requires a deliberate side buffer.

solid answer

~40 s

The danger is not decoding, it is **re-emitting**. A reader decodes bytes into a value shaped by its own schema version, edits one field, and then encodes a fresh message *from that value* — the input bytes are not edited in place. Anything the reader did not model never entered the value, so it cannot leave in the output. A reader has three possible policies for fields it does not know: drop them, preserve them in a side buffer and write them back on encode, or reject the message outright. Dropping is the common behaviour unless preservation was deliberately chosen. The loss is silent because both the decode and the re-encode succeed, so no hop reports anything abnormal; the damage surfaces later, at a sink that expected a field the producer really did send.

code

pseudocode · 11 lines
pseudocode
function handle(input_bytes):
    value, unknown = decode_keeping_unknown(input_bytes)
    # unknown: list of (field_id, raw_bytes) the schema version does not define

    value.region = lookup_region(value.customerId)   # this hop's one job

    output = encode(value)                           # built from the model only
    for each (field_id, raw_bytes) in unknown:       # a dropping reader skips this loop
        output = append_raw_field(output, field_id, raw_bytes)

    emit(output)

go deeper

for a junior

Recall that a reader only sees the fields its own schema version names, and that re-encoding builds a fresh message from what it saw rather than patching the bytes that arrived.

for a middle

Explain the decode-edit-encode sequence and name all three reader policies — drop, preserve, reject — saying which one a reader follows when nobody has chosen deliberately.

for a senior

Show how you would detect the loss in a running pipeline: per-hop counts of unrecognised fields, a round-trip fixture in each service's tests, and end-to-end field-set comparison on sampled traffic.

for a principal

Weigh preservation's cost — every transit hop carrying and storing data it cannot validate — against constraining the topology so fields only cross hops that understand them.

## What counts as an unknown field A **wire contract** is the set of fields a message may carry plus the rules for reading them. Two parties rarely hold the same version of that contract at the same instant: a producer ships a new field on Monday, an enrichment service in the middle of the pipeline upgrades on Thursday, and the sink at the end upgrades whenever its team gets to it. In the gap, the middle reader meets bytes for a field its own schema version does not define. That field is **unknown** to it — not malformed and not corrupt, simply outside its model. What the reader does with those bytes has exactly three outcomes: **drop**, **preserve**, or **reject**. Most readers drop, unless someone deliberately chose otherwise. ## Why the loss happens on encode, not on decode The dangerous shape is a reader that does not merely consume a message but **re-emits** one: 1. **Decode.** The bytes are walked and turned into an in-memory value shaped by the reader's schema version. Recognised fields become members of that value; unrecognised ones are skipped — skipping is exactly what a decoder must do to find the next field boundary. 2. **Edit.** The service does its one job: it stamps a region, attaches a score, marks the message as seen. 3. **Encode.** The output is built **from the in-memory value**, not from the input bytes. Nothing was edited in place. The output was rebuilt from a smaller model, and a field that never entered the model cannot leave in the output. This is why the failure is described as *round-trip data loss* rather than a parsing bug: each hop worked exactly as written. ## The three reader policies | Policy | What the reader does with the bytes | What the pipeline gets | Where it fits | |---|---|---|---| | **Drop** | Skips them; they never enter the value | Simplicity; silent loss on any re-emit | Terminal consumers that never re-emit | | **Preserve** | Keeps them in a side buffer, writes them back on encode | Data survives hops that implement preservation | Transit hops, enrichers, routers | | **Reject** | Fails the decode and refuses the message | Loud, immediate feedback to the sender | Edges where a human authored the document | Preservation is a property of the **reader implementation plus the contract that demands it**, not something an encoder does by accident. An encoder writes whatever the value holds; if the value holds nothing for a field, nothing is written. ## Why it is silent Three things conspire: - **No error exists to raise.** A skipped field is a normal decoding path, not an exception. - **No feedback channel runs backwards.** The producer gets an acknowledgement for a message that was accepted, because it *was* accepted. - **The symptom appears at a different service.** The sink sees an absent field and, if the encoding substitutes defaults, cannot even tell the field was ever present. The team that debugs it is rarely the team that owns the hop that dropped it. ## What preservation does and does not give you Preservation means the unrecognised bytes are held alongside the decoded value and written back out. It does **not** mean the hop understands them, can validate them, or can make decisions based on them. A preserving hop is a faithful courier for data it is deliberately blind to — which is a feature on an internal transit hop and a liability at a boundary whose whole job is to vet what crosses it. Preservation also has a real cost: every hop now carries, and possibly stores, fields nobody at that hop can interpret, and the in-memory value grows a second, untyped compartment that the service's own code must not accidentally discard when it constructs a fresh message rather than editing the decoded one. ## Making the behaviour observable Because the failure is quiet, it has to be made loud on purpose: - **Count unrecognised fields per hop** and alert when a hop that is supposed to preserve starts seeing them without forwarding them. - **Keep a round-trip conformance fixture** in every service's test suite: a message carrying a field the service's schema does not define, asserted to come out the far side intact. - **Compare field sets end to end** — what the producer emitted against what the sink persisted — on a sampled fraction of traffic. - **Constrain the topology** so that a field only crosses hops known to preserve it, rather than assuming every hop does. The interview answer an experienced engineer gives is the mechanism plus the mitigation: decode-edit-encode rebuilds the message from a smaller model, so the fix is a side buffer plus a test that proves it works, not a stricter parser.

  • When is dropping unknown fields the correct behaviour rather than a bug?
    At a terminal consumer that persists its own projection and re-emits nothing, dropping costs nothing. It is also the right call at a boundary whose job is to vet what crosses it: a preserving hop forwards data nobody at that hop reviewed, which is precisely what such a boundary exists to prevent.
  • Why does unknown-field loss usually get reported by a team that did not cause it?
    The hop that drops the field succeeds and stays quiet. The symptom appears wherever the field was expected — often several services later — as an absent value or a substituted default. Without a per-hop count of unrecognised fields, the investigation starts at the sink and walks upstream by elimination.
  • Does preserving unknown fields require understanding them?
    No, and that is the point. The reader keeps the raw bytes with enough structure to write them back — the field's identifier and its encoded value — without interpreting them. It cannot validate, filter or act on them, so preservation is safe on a transit hop and questionable at a trust boundary.

A clerk copies every arriving form onto the office's own blank, which has fewer boxes. Nothing is refused and nobody complains; the answers with no box simply never reach the copy.

saying these in an interview costs you the question

  • Says a decoder keeps every field it saw, whether modelled or not
  • Thinks editing one field leaves the rest of the input bytes untouched
  • Assumes the producer gets an error when its fields are dropped
  • Believes only binary encodings lose unknown fields, never text ones
  • Confuses preserving unknown fields with validating or understanding them
  • Reaches for a stricter parser when the actual fix is a side buffer
open as a page

How do you choose the default for a new field so older readers that ignore it stay correct?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Pick the value that reproduces the behaviour the system had before the field existed, and check that the same value is safe when the field is lost in transit. If no single value satisfies both, the state does not belong in a defaulted field.

open as a page

Why can a reader that substitutes implicit defaults not tell a writer's explicit zero from an unset field?

level: middleimportance: should knowfreq 45%

basics

~20 s

Because the default is materialised during decoding: an absent field and a field carrying the type's default produce the identical in-memory value. Presence information is destroyed before any application code runs, so no later check can recover it.

open as a page

Which reader behaviour for unknown fields — strict rejection or lenient ignoring — would you write into a wire contract, and when?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Strict rejection where a human or a tool authored the document and a misspelled field must never be silently discarded; lenient ignoring on machine-to-machine streams with many independently deployed producers. Either way, the contract states which, rather than leaving it to each decoder's configuration.

open as a page

Across a pipeline of services on different schema versions, how would you set one policy for unknown-field handling?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Classify edges rather than services: strict at boundaries where a human or tool authored the input, preserve-and-forward on internal transit hops, drop at terminal sinks. Then make the policy real with a conformance fixture, a per-hop counter, and a named owner.

open as a page