skip to content

Your service decodes every payload into map[string]any and type-asserts at use sites. Why does that panic in production?

level: seniorimportance: should knowfreq 44%

answer

  1. a successful decode proves only that it was JSON
  2. all the checking got deferred to the readers
  3. the crash names your code, not the input
  4. type what you touch, at the edge
  5. comma-ok, or a struct that validates for you

basics

~10 s

Decoding into map[string]any succeeds for any valid JSON, so nothing checks shape, and each single-result type assertion is an unchecked bet. Drift panics deep in business logic instead of failing at the decode.

solid answer

~50 s

Decoding into `map[string]any` succeeds for any syntactically valid JSON, so the decode step tells you nothing about shape; every field access afterwards is an unvalidated assertion. Single-result assertions such as `m["count"].(int)` panic on mismatch, and that one always panics because decoded numbers are `float64`. Others panic only for the payloads that drift: a field that arrives as `"12"` instead of `12`, an explicit `null` leaving a nil interface, a single object where the sample had an array. The panic then lands at the use site, layers into business logic and possibly on a goroutine no handler recovers, so the stack trace names your code rather than the payload. The fix is to validate at the boundary: decode into a typed struct so `encoding/json` returns an `*json.UnmarshalTypeError` naming the field, keep `map[string]any` for genuinely schemaless passthrough, and use comma-ok assertions everywhere it survives.

code

go · 18 lines
go
// always panics: a decoded JSON number is float64, never int
_ = m["count"].(int)

// panics only when the producer starts sending "12" instead of 12
_ = int(m["count"].(float64))

// the boundary does the checking and names the offending field
type Batch struct {
	Count int64 `json:"count"`
}
var b Batch
if err := json.Unmarshal(body, &b); err != nil {
	var te *json.UnmarshalTypeError
	if errors.As(err, &te) {
		return fmt.Errorf("field %s: got JSON %s: %w", te.Field, te.Value, err)
	}
	return err
}

go deeper

for a junior

Remember that a type assertion written with one result crashes when the type differs, and that the comma-ok form gives you a boolean to branch on instead.

for a middle

Explain why decoding into an untyped map cannot fail on shape, and show how a typed struct converts the same mismatch into an error that names the field.

for a senior

Demonstrate the diagnosis path from panic message to offending payload, and argue for validating at the boundary, including what you would keep untyped and why.

for a principal

Own it as an integration standard: where schemas are declared, what a producer must do before changing a field's type, and how contract fixtures keep drift from becoming an outage.

## Why this shape is so common It usually starts well. Somebody onboarding to an unfamiliar integration writes a small terminal program that reads a payload, decodes it into `any`, and prints `%T` and the value for each field to learn the schema. That is an excellent way to *explore* a document. The mistake is shipping the exploration: the map stays, the printed types get baked into assertions, and the service now carries a schema that exists only in the heads of whoever read one sample. ## What actually fails Decoding into `map[string]any` succeeds for every syntactically valid JSON document. There is no schema, so there is no mismatch to report — the decode's nil error means only "this was JSON". All the type checking has been deferred to the use sites, where it takes the form of assertions: - `m["count"].(int)` panics on every payload, because a decoded number is a `float64`; this one at least fails immediately in testing - `m["user"].(map[string]any)["id"].(string)` panics when `user` is `null`, because the intermediate is a nil interface - `m["items"].([]any)` panics when a producer that used to send a list starts sending a single object - `m["id"].(float64)` succeeds but rounds, when the id is large — that one does not even panic, it corrupts The first is a bug you find on day one. The rest are latent, and they fire on the day an upstream team makes a change they consider backwards-compatible. ## Why the panic is expensive A failed assertion panics at the point of use. That is typically several call frames into business logic, so the stack trace names your code and not the input, and the log line does not contain the payload. If the work happens in a goroutine spawned by the request rather than on the handler's own goroutine, no server-level recovery applies and the panic takes the whole process down — one malformed payload becomes an outage rather than a rejected request. Blanket recovery is not a fix either: it converts a crash into a silently dropped message. ## Moving the failure to the boundary The structural fix is to make the shape explicit once, where the data enters: ```go type Order struct { ID int64 `json:"id"` Items []Item `json:"items"` User *User `json:"user"` } ``` Now `encoding/json` performs the checking. A string where a number was expected returns an `*json.UnmarshalTypeError` that names the field, the JSON type it found and the Go type it wanted; the decoder fills what it can and reports that at the end, so you can log a precise complaint and reject the message with a 400 instead of panicking with a 500. The int64 field also removes the float64 rounding. And the struct is documentation: a new engineer reads the expected shape instead of reverse-engineering it from assertions. ## When a map is still right There are honest uses for `map[string]any`: a genuinely schemaless field such as user-supplied metadata, a component that forwards documents it deliberately does not interpret, or a tool that inspects arbitrary payloads. The rule that keeps those safe is to type what you touch. If your code reads a field, that field belongs in a struct; the parts you never look at can stay untyped. For a forwarder that must not corrupt numbers it never reads, decode with `UseNumber` so numeric literals survive as `json.Number` rather than passing through `float64`. Where assertions do remain, always use the comma-ok form, and prefer a `switch v := x.(type)` over a chain of them so the unexpected case has a place to go. The difference between `v := x.(string)` and `v, ok := x.(string)` is the difference between a crash and a branch. ## How you would find it in a running system Start from the panic: the message names the two types involved ("interface conversion: interface {} is float64, not int"), which identifies the field's real wire type immediately. Then capture a failing payload — logging the raw body on the error path, not on the happy path — and compare it with what the assertions expect. If it is drift rather than a longstanding bug, the durable fix is the typed boundary plus a test that decodes a stored sample of each producer's payload, so the next change fails in CI rather than at 3am.

  • Why is recovering from the panic in the handler not an adequate fix?
    It converts a crash into a silently dropped or half-processed message, and it leaves the real defect — an unvalidated boundary — in place. It also does not help when the assertion runs on a goroutine the handler spawned, because a panic there is not recovered by the handler and takes the process down. Recovery is a blast-radius control, not a substitute for validating input where it arrives.
  • The team says the payloads are genuinely schemaless, so a struct is impossible. What do you propose?
    Type what you touch. Any field the code actually reads goes into a struct, even if the rest of the document stays untyped alongside it, and the untyped remainder is only forwarded, never interpreted. Where assertions survive, require the comma-ok form or a type switch with a default branch, so an unexpected shape becomes a rejection with a clear message instead of a panic.
  • What does an UnmarshalTypeError give you that a panic does not?
    It names the struct field, the JSON type that was found and the Go type that was wanted, and it arrives while you still hold the raw body, so the log line can identify both the producer and the offending value. Decoding also continues past it, so you learn about the document as a whole rather than stopping at the first surprise, and the request can be rejected as a client error rather than crashing as a server error.
  • How would you stop the same class of drift from reaching production again?
    Store a real sample payload from each producer as a test fixture and decode it in CI against the typed boundary, so a shape change fails the build. Pair that with logging the raw body on the decode-error path only, so the first production occurrence carries its own evidence. Neither is expensive, and both fail loudly at the edge rather than quietly in the middle.

It is like accepting parcels without opening them and then having each department reach blindly into the box for the item it expects. The first mismatch is discovered by whoever put their hand in, far from the loading bay where it could have been rejected.

saying these in an interview costs you the question

  • Says a successful decode proves the payload was valid
  • Wraps every handler in a recover and calls it fixed
  • Uses single-result type assertions on decoded values
  • Asserts a decoded number to int without going through float64
  • Reverse-engineers the schema from one sample payload