A Go type switch with 20 cases sits on a hot path, defended by a 2 ns/op benchmark — what do you challenge?
answer
- ask what the benchmark fed it
- one dynamic type is the best case
- not all arms cost the same
- the copy is outside the timed region
- profile the service before tuning the switch
basics
~10 sChallenge what the benchmark fed it. One dynamic type hits one arm with a predicted branch; real traffic mixes kinds, arms naming an interface cost a runtime check, and the interface copies go uncounted.
solid answer
~50 sMy first question is what the benchmark passed in. If it boxes one event type and loops, it measures the best case: the same arm every iteration, a perfectly predicted branch, one type descriptor hot in cache. Production traffic mixes twenty kinds, so I want the benchmark re-run over a representative mix, with `-benchmem`, and ideally against a CPU profile of the real service rather than a microbenchmark at all. Then I look at the arms: cases naming concrete types are dynamic-type comparisons and the compiler is free to dispatch a long run of them better than a linear chain, but any case naming an interface type is a runtime implements check, evaluated in source order, so an expensive one placed early is paid by everything that falls past it. Finally, 2 ns/op is nowhere near the real per-message cost, because the copy into `any` and back out is not in that number.
code
go · 10 linesswitch e := ev.(type) {
case OrderCreated: // dynamic-type comparison
handleCreated(e)
case Validatable: // runtime implements check, evaluated in order
handleValidatable(e)
case OrderCancelled:
handleCancelled(e)
default:
deadLetter(ev) // counted, not silently dropped
}go deeper
Know what a type switch does and that its arms are tried against the dynamic type of the interface value; the benchmark critique is not expected of you yet.
Be able to say that arms naming concrete types and arms naming interface types are not the same check, and that a benchmark feeding one type exercises only one arm.
Demonstrate the review instinct: interrogate the benchmark's input and timed region, ask for a CPU profile of the real service, and separate the dispatch cost from the per-message copy into and out of the interface.
Own the call between a central switch and behaviour on the types. Weigh who adds event kinds, whether a missing arm should be a compile error, and whether a nanosecond-scale saving justifies changing a contract other teams depend on.
## The claim and what it actually measures "The type switch is 2 ns/op, so it is free" is a benchmark reading, not a cost model. Before arguing about the switch, read the benchmark. **A single dynamic type.** The overwhelmingly common shape is a benchmark that constructs one event, boxes it into `any` once, and loops. That measures a monomorphic call site: the same arm is taken every iteration, the branch predictor is right every time, the type descriptor and the event's memory stay in L1, and nothing else competes for cache. A dispatcher fed twenty kinds in unpredictable order behaves differently for reasons that have nothing to do with Go — the arms and the data both stop being predictable. **The boxing is outside the timed region.** If the value is boxed before the loop, the per-message cost of getting a value into `any` and of copying it back out on assertion is excluded. On a real path that copy is per message, and for a large event it dominates the check itself. `-benchmem` is the minimum ask here; it will show whether the boxing the benchmark does perform allocates. **The loop may not be doing what it looks like.** If the result is unused, the compiler is entitled to remove work. Benchmarks in this area should consume the result — assign it to a package-level sink — and, on modern Go, use `testing.B.Loop`, which is designed to keep the benchmarked value alive across iterations rather than relying on a hand-written sink around `b.N`. ## Then read the arms A type switch is not one uniform mechanism. - **Concrete-type arms** are dynamic-type comparisons. The compiler is free to dispatch a long run of them better than testing each in turn, so twenty concrete cases is not twenty tests per message. - **Interface-type arms** are implements checks the runtime has to answer, and they are evaluated in source order. An interface arm placed near the top is paid by every message that does not match it. If one of the twenty cases is `case Validatable:`, that is the arm worth measuring, not the nineteen concrete ones. - **The default arm** is a design question as much as a cost one. What does an unrecognised event kind do — get counted and dead-lettered, or fall through silently? A dispatcher whose default arm does nothing turns a new producer's event type into silent data loss. ## What I would ask for instead 1. A CPU profile of the service under real traffic, showing that this dispatch is actually hot. Most of the time it is not, and the whole discussion is premature. 2. If it is hot: a benchmark over a representative mix of event kinds, with `-benchmem`, plus the boxing inside the timed region, because that is what production pays. 3. If the numbers hold up, restructure rather than micro-tune: decide the type once at the boundary where the event is decoded and dispatch through a method or a small kind field afterwards, so the per-message path does not re-ask a question that was already answered upstream. ## The reviewer's other objection There is a maintenance argument that usually outweighs the nanoseconds. A twenty-arm type switch is a central place every new event type must be edited into, and forgetting is a silent default-arm fall-through rather than a compile error. Moving the behaviour onto the event types themselves gives you a compile-time obligation instead. That is a design tradeoff — a closed set of types defined in one package is a perfectly good reason to keep the switch — but it should be the argument being had, rather than a 2 ns/op number that measures a case the service never runs. ## What not to say Do not claim the switch "compiles to a jump table" or "is O(1) regardless of cases" — the compiler has latitude here and the honest statement is that concrete arms are cheap comparisons and interface arms are not. And do not accept the inverse claim either, that twenty cases means twenty checks per message; that is the pessimistic story and it is equally unsupported.
- What would you ask the author to measure instead?A CPU profile of the service under production traffic first, to show the dispatch is hot at all. If it is, a benchmark whose input is a representative mix of event kinds, with the boxing inside the timed region and `-benchmem` reported, so the copy into and out of the interface is counted rather than hoisted out of the loop.
- The hot arm turns out to be an interface case. What do you do?Either move it below the concrete arms so most messages never evaluate it, or remove the question from the per-message path entirely: resolve the interface once where the event is decoded, store the resolved handler or a small kind field, and dispatch on that afterwards. Re-profile under the same traffic mix rather than trusting the reasoning.
- When is a twenty-arm type switch the right design anyway?When the set of types is closed and defined in the same package as the switch, so a new type is a deliberate edit in one place and reviewers see all the routing at once. It stops being right when other teams add event types, because a missing arm is then a silent default-arm fall-through rather than a compile error.
saying these in an interview costs you the question
- Accepts a microbenchmark fed one dynamic type as representative
- Says a type switch always costs one check per case
- Claims a type switch compiles to a constant-time jump table
- Ignores that boxing and unboxing sit outside the timed loop
- Optimises the dispatch without a CPU profile of the service
- Leaves the default arm silently dropping unknown event kinds