An API gateway switching to encoding/json/v2 starts rejecting payloads it used to forward. How do you find them before the switch?
answer
- Measure before you flip
- Decode twice, serve one result
- Two classes shout, one is silent
- Label by sender, never log the body
- A clean sample is not a proof
basics
~20 sShadow-decode live traffic: run every payload through both encoding/json and encoding/json/v2, keep the v1 result, and record only the disagreements, classified as duplicate name, invalid UTF-8 or an unmatched member. Then attribute each class to a sender before flipping any route.
solid answer
~50 sDo not flip and watch the error rate; measure first. Add a shadow path that decodes each inbound payload with both packages, serves the `encoding/json` result exactly as today, and records whether `encoding/json/v2` disagreed -- with a counter labelled by class (duplicate object name, invalid UTF-8, member matched only case-insensitively) and by route and sender, never the body itself. Two of the three classes surface as an error from v2; the third does not, so the shadow path must also compare the decoded values, otherwise a member silently unmatched by case looks clean. Run it long enough to cover a real cycle of partner behaviour, attribute every hit to a named sender, and only then switch route by route, keeping the loosening options -- `jsontext.AllowDuplicateNames(true)` and friends -- as a per-route lever for the one partner still being fixed. A clean week proves only that the sampled traffic was clean, not that nothing will ever be rejected.
code
go · 14 lines// jsonv1 "encoding/json"; json is "encoding/json/v2"
var got Payload
if err := jsonv1.Unmarshal(body, &got); err != nil {
return err // serve exactly as before
}
var shadow Payload
switch err := json.Unmarshal(body, &shadow); {
case err != nil:
countRejected(route, sender, classify(err))
case !samePayload(shadow, got):
countDiffered(route, sender) // the silent class: no error, different value
}go deeper
Know that the two packages can coexist in one binary, and that you can decode the same bytes with both to compare them. The instinct to measure before changing behaviour is the point here.
Explain the three disagreement classes you would count and why one of them cannot be detected from the error alone. Be able to name the per-call option that restores each behaviour.
An interviewer expects the full plan: shadow decode with per-sender labels, no bodies in logs, sampling that respects systematic per-producer defects, a route-by-route switch, and an explicit statement of what a clean sample does not prove.
Own the negotiation around the data. The measurement converts a strictness upgrade into a list of named partners and dates, which is the artefact that lets you decide who absorbs the cost of the change and when the exceptions expire.
## The problem with the obvious approach Switching the gateway's import and watching the error rate is a live experiment on other people's traffic. The failure is asymmetric: the gateway is the component that sits between senders it does not control and services that trust it, so a rejection it introduces is a partner outage attributed to you, and it arrives all at once. ## Shadow decode The measurement is cheap because both packages are in the same binary. On the inbound path, decode with `encoding/json` as today and serve that result; alongside it, decode the same bytes with `encoding/json/v2` and throw the result away except for a comparison. Emit one counter per outcome: - both succeed and the decoded values agree -- the normal case; - v2 returned an error -- record its class; - both succeeded but the decoded values differ -- the quiet case. That third bucket is the one an inexperienced version of this plan omits. Duplicate names and invalid UTF-8 announce themselves as errors from v2. A member that used to be matched case-insensitively simply goes unmatched under v2's default, leaving a zero-valued field and no error at all, so it is only visible by comparing decoded values rather than error status. What you label the counters with matters as much as having them. Route, sender or credential id, and failure class -- enough to phone the right team. Not the payload: this is inbound traffic from outside, it may carry personal data, and a log line with the whole body is a new incident. Record the offending member name and its location in the document, which is what v2's errors identify, and nothing else. ## Cost, and how to keep it honest Double-decoding doubles the parse cost on the hot path. If that is not affordable at full volume, sample -- but sample by *sender* rather than uniformly by request, because the defects here are systematic per producer. A partner whose encoder emits duplicate names emits them on every request, so a 1% uniform sample finds them anyway; a partner who sends invalid UTF-8 only in a rarely used free-text field may hide from any sample, which is an argument for running the shadow path over a full cycle of the business -- a month-end batch, a seasonal upload -- rather than an afternoon. ## Rolling out Adoption is an import change per call site, so it stages naturally. Move one route at a time, starting with the ones whose senders are internal and whose shadow counters have been zero. For a route with a known offender, switch the import anyway and pass the matching loosening option on that call: `jsontext.AllowDuplicateNames(true)` restores tolerance of repeats, `jsontext.AllowInvalidUTF8(true)` of bad bytes, `json.MatchCaseInsensitiveNames(true)` of the old name matching. That is better than leaving the route on v1, because the exception is now a visible argument at one call site with a comment naming the partner and the date it should be removed, rather than an invisible difference between two imports. Keep those options as the rollback lever. Rebuilding the binary with `GOEXPERIMENT=nojsonv2` is not a rollback for this: it is a build-level revert of the whole implementation that tells you nothing about which route broke, and it does not undo an import you changed. ## What a clean result does and does not prove This is the same epistemics as a clean race-detector run. A week of zero disagreements says the traffic you sampled contained nothing the stricter package rejects. It does not say a partner will not deploy a new encoder next month, and it does not cover routes that were idle. So the switch ships with an alert on the new rejection class, a documented runbook -- which option to pass on which route, and who to call -- and a note that the exception expires. The measurement buys you a staged rollout, not a guarantee.
- Why is comparing only the error status of the two decoders insufficient?Because one of the three changed defaults never produces an error. A member that the original package matched case-insensitively is simply not matched by encoding/json/v2, the field keeps its zero value, and both decoders return nil. Only comparing the decoded values exposes it. That is also why the eventual migration test asserts field values rather than just that Unmarshal succeeded.
- One partner sends duplicate object names on a single route. What is the narrowest fix?Switch that route's import to encoding/json/v2 like the others and pass `jsontext.AllowDuplicateNames(true)` on that one call, with a comment naming the partner and an expiry date. The rest of the gateway stays strict, the exception is greppable and removable, and nothing about the build changes. Reverting the import or rebuilding with GOEXPERIMENT=nojsonv2 both trade a one-line exception for an invisible one.
- How do you decide how long to run the shadow path?Long enough to cover a full cycle of sender behaviour, not a fixed number of days. Systematic defects show up in minutes because a bad encoder is bad on every request; the ones that hide are tied to rare paths -- a month-end batch, a seasonal upload, one partner's quarterly file. Sample by sender rather than uniformly, since per-sender coverage is what actually matters.
- What do you log when the strict decoder rejects an inbound payload?Route, sender identity, failure class, and the member name and document location that the error identifies. Not the body, and not the offending value: this is untrusted inbound traffic that may carry personal data, and a log line containing it turns a migration into a data-handling incident. The counter tells you the size of the problem; the member name tells the partner what to fix.
saying these in an interview costs you the question
- Flips the import in production and watches the error rate
- Compares only error status, missing silently unmatched members
- Logs whole rejected bodies from untrusted senders
- Treats a clean sample week as proof no payload will be rejected
- Reaches for GOEXPERIMENT=nojsonv2 as the rollback lever
- Rolls the whole gateway at once instead of route by route