skip to content

Why might one system deliberately use a different encoding on its internal hop, its public webhook, and its nightly archive?

level: seniorimportance: should knowfreq 50%

answer

  1. one system, several boundaries
  2. criteria belong to the hop
  3. internal optimises bytes and codegen
  4. public optimises unknown, unupgradable consumers
  5. archive optimises readers years later

basics

~20 s

Those three hops face different constraints: an internal hop optimises bytes and generated code between services that deploy together, a public hop optimises for consumers nobody controls, and an archive optimises for a reader who arrives years later with different tools.

solid answer

~50 s

Because the criteria that decide an encoding are properties of the hop, and these three hops differ on almost every one. The internal hop is high volume with both ends under one deployment, so compactness and generated accessors win and evolution is cheap to coordinate. The public hop is low volume with unknown consumers on their own upgrade schedules, so breadth of decoders and debuggability win, and every change must be survivable by a reader that never upgrades. The archive is written once and read by code that does not exist yet, so self-description and durable, widely available tooling beat per-record size. A single house encoding is still a real option — one toolchain and one set of failure modes is worth a lot — but it should be chosen as a deliberate trade, not assumed because uniformity feels tidy.

go deeper

for a junior

Know that the same data can be encoded differently at different boundaries, and that the reason is who reads the bytes and when, not fashion or inconsistency.

for a middle

Explain how each criterion shifts between hops: why coordinated deployment makes evolution cheap internally, why an unknown consumer makes it expensive externally, and why an archive's reader constrains the choice hardest.

for a senior

Argue both directions honestly. Name what uniformity buys — one toolchain, one set of failure modes — and then show the constraint difference that makes a second encoding worth its translation points.

for a principal

Set defaults per class of hop rather than blessing one encoding, keep a deliberate count of live toolchains, and pick a canonical representation so that fidelity across boundaries has one answer instead of one per translation.

## One system, several boundaries People find per-hop variation suspicious because it looks like inconsistency. It is not: it is the same selection criteria applied to genuinely different workloads. An encoding choice is only ever about one boundary, and a product of any size has several boundaries whose constraints barely overlap. The question is not "should we be consistent" but "what is different about these hops, and is the difference worth a second toolchain". ## Three hops, three constraint sets | Hop | Who reads it | When | Dominant criteria | |---|---|---|---| | Internal service-to-service | Code your team owns and deploys | Now, continuously, at volume | Per-message size, decode cost, generated accessors | | Public webhook or partner feed | Consumers you neither own nor can upgrade | Now, and on their schedule forever | Decoder availability everywhere, debuggability, survivable change | | Nightly archive | Code that may not be written yet | In months or years | Self-description, durable tooling, readability without the original toolchain | Read the table by column rather than by row. **Who reads it** decides how much you may assume about their tools. **When** decides how long a mistake persists: an internal choice is revisited with one coordinated release, a public one requires everybody else to move, and an archive one is discovered when the person who could have fixed it has left. **Dominant criteria** is simply what falls out of the first two. The consequences are concrete: - On the internal hop, evolution pressure is low because both ends ship together, so an encoding that demands a coordinated change is affordable, and a compact representation pays off every message. - On the public hop, evolution pressure is at its maximum: an old reader will meet new bytes, and you cannot schedule its upgrade. That, plus incident debuggability, is why external surfaces so often end up readable even in systems that are compact everywhere else. - An archive usually outlives the code that wrote it, so the criterion is not size but whether a future reader can make sense of the bytes with tooling that still exists. Storage is cheap; an unreadable decade of records is not. ## What uniformity actually buys One house encoding is not a naive position, and a candidate who dismisses it has missed half the trade: - One toolchain to install, version, patch and harden, rather than three. - One set of failure modes engineers learn once, so an on-call engineer recognises a decode failure anywhere in the system. - One place to enforce limits, and one decoder surface to review when a vulnerability lands. - No translation points, and therefore no places where a value can quietly lose fidelity crossing between two different data models. ## The price of heterogeneity Against that, three encodings cost: - Translation code at every boundary between them, which is where numeric precision, absent-versus-null and time representations get quietly flattened. - More libraries in the dependency graph, each with its own limits, defaults and advisories. - A larger surface a new engineer must learn before they can debug across the whole system. The honest answer is that per-hop variation is justified when the constraint sets genuinely diverge — an external surface, an archive, an analytic dump — and is not justified because a team preferred a different encoding for two adjacent internal services. ## Keeping per-hop choice from becoming a free-for-all 1. Name a **default** for each class of hop rather than for the product: an internal default, an external default, an archive default. Most choices are then nobody's decision. 2. Require a hop that departs from its class default to say which criterion drove it, so departures are evidence-backed rather than preference-backed. 3. Pick one representation as **canonical** for each value that crosses classes, and translate at exactly one place, so fidelity questions have one answer rather than one per boundary. 4. Count the toolchains deliberately. Three is a decision; seven is an accident, and it usually means the defaults were never written down. ## Where this goes wrong - Extending the internal encoding to external consumers because it was already there, then discovering it cannot be changed once partners depend on it. - Archiving in whatever the live hop happened to use, so the archive inherits an evolution model built for code that redeploys daily. - Translating between encodings at several boundaries and losing precision or the distinction between absent and null at each one. - Treating consistency as a value in itself and paying a large per-message cost on the highest-volume hop in the system to preserve it.

  • What is the real price of running three encodings in one system?
    Translation points and toolchains. Every boundary between two encodings is code where numeric precision, the absent-versus-null distinction and time representations can be flattened, and every extra library brings its own limits, defaults and advisories. Heterogeneity is worth paying for when constraint sets genuinely diverge, and is pure overhead when it reflects team preference between two adjacent internal hops.
  • Which hop most often ends up with the wrong encoding through inertia?
    The archive. It usually inherits whatever the live hop was already emitting, which was chosen for services that redeploy together and can coordinate a breaking change. The archive's reader cannot coordinate anything: it arrives years later, possibly with none of the original tooling, and needs bytes that describe themselves well enough to be understood on their own.
  • How do you stop per-hop freedom from becoming a different encoding per team?
    Set defaults per class of hop rather than per product — an internal default, an external default, an archive default — so the usual case is nobody's decision. Require any departure to name the criterion that drove it, and keep a deliberate count of live toolchains. Freedom without a default reliably drifts into one encoding per team.

saying these in an interview costs you the question

  • Insists one encoding everywhere is always simpler and therefore better.
  • Treats a public partner hop and an internal hop as the same problem.
  • Forgets that an archive is read long after its writing code is gone.
  • Counts wire cost only, ignoring the cost of extra toolchains.
  • Converts between encodings at many boundaries without noticing fidelity loss.
  • Extends the internal encoding outward because it was already there.