How do you cache a reflect-built field plan keyed by reflect.Type, and what does that stop repeating?
answer
- two phases, not one
- the plan depends only on the type
- the type itself is the key
- store index paths, never Values
- removes discovery, not boxing
basics
~20 sWalk each destination type once, store the resulting field-index plan in a map keyed by its reflect.Type, and look the plan up per row. reflect.Type values are canonical and comparable, so they work directly as map keys.
solid answer
~50 sThe expensive part of a naive reflective mapper is discovery: on every row it asks the type how many fields it has, matches each destination field, and builds the same descriptors again. That work depends only on the type, so do it once. Build a plan — typically a slice of field-index paths plus a per-field conversion function — and store it in a `sync.Map` keyed by the `reflect.Type`, using `LoadOrStore` so a duplicate build under concurrency is merely wasted, never wrong. `reflect.Type` values are comparable and canonical, so equal Go types give one entry, and the cache is bounded by the number of distinct types the program maps, not by rows. What must never go in the plan is a `reflect.Value`: values are bound to one particular instance. And the cache does not remove the interface boxing or the dynamic field access — only the rediscovery.
code
go · 14 linestype fieldPlan struct {
index [][]int // one field-index path per destination position
}
var plans sync.Map // reflect.Type -> *fieldPlan
func planFor(t reflect.Type) *fieldPlan {
if p, ok := plans.Load(t); ok {
return p.(*fieldPlan)
}
p := buildPlan(t) // walks t exactly once per type
actual, _ := plans.LoadOrStore(t, p)
return actual.(*fieldPlan)
}go deeper
Know the shape of the idea: work that depends only on the type can be done once and reused, instead of on every row. Being able to describe that split in plain terms is what is expected here.
Be able to write it: a plan of field-index paths, a sync.Map keyed by reflect.Type, LoadOrStore for the race, and no reflect.Value in the cache. Explain why the type value itself is a valid key.
Show that you know what the cache leaves behind — boxing and dynamic access — and that you would measure the residue before proposing a bigger change. Mention that the entry count is bounded by distinct types, so the cache cannot grow with traffic.
Own the framing that this is the cheap, contained fix with no build-system consequences, and that its measured result is the evidence for or against paying for code generation.
## The idea Split a reflective mapper into two phases: a **plan** that depends only on the type, and an **execution** that depends on the row. The plan is derived once and reused; execution runs ten million times. A mapper written without this split does both phases per row, which is why it re-derives descriptors it already computed for the previous row. ## Why `reflect.Type` is a legitimate map key `reflect.Type` is an interface, and the concrete values behind it are canonical: for any given Go type there is exactly one `reflect.Type` value, so `==` on two `reflect.Type` values means "the same Go type". That makes it directly usable as a map key with no stringification. Using `t.String()` or `t.Name()` as the key instead is a bug waiting to happen — two types in different packages can share a name, and `Name()` is empty for unnamed types such as `[]byte` or an anonymous struct. ## What belongs in the plan The plan holds *descriptions*, never live values: - the field-index path for each destination position, as `[]int`, usable with `Value.FieldByIndex` (an embedded field needs a path rather than a single index); - per-field conversion or validation decided once from the field's `reflect.Type`; - anything else derived purely from the type, such as which destination positions have no source and must be skipped. What must **not** go in it is a `reflect.Value`. A `Value` is bound to one particular instance; caching one per type and reusing it across rows is a correctness bug, not an optimisation, and it keeps whatever it points at alive for the life of the process. ## Concurrency and the shape of the cache A package-level `map[reflect.Type]*fieldPlan` filled lazily is not safe if worker goroutines read it while another fills it — Go's built-in map detects concurrent map read and write and throws, killing the process, and the race detector flags it immediately. Two safe shapes: - `sync.Map` with `Load` then `LoadOrStore`. Plan construction is idempotent, so two goroutines racing to build the plan for the same type is harmless: one plan wins and the loser's copy is garbage. This is the classic read-mostly, write-once-per-key workload `sync.Map` is documented for. - A plain map behind a `sync.RWMutex`, which is fine and often faster than people expect when the key set is small. There is also the zero-concurrency option: build every plan at package initialisation from a known list of types, and let a missing type be a programming error rather than a lazy fill. ## Memory behaviour, and why this cache is well-behaved The entry count is bounded by the number of distinct Go types the program maps — a fixed, small number known at compile time. This is exactly what makes it a safe cache: it cannot grow with traffic. The one way to break that property is to key it on something unbounded, for example a type generated per request through reflection, which would turn a bounded cache into a leak. ## What the plan cache does and does not buy It buys the removal of rediscovery: the field count, the index paths, the per-field decisions. On a wide struct that is often the majority of the reflective overhead, and it costs you nothing structural — no build step, no change to the API other teams call, no generated files. It is the first thing to do and frequently the last thing you need. It does **not** buy: - **The interface boxing.** Values still cross the `any` boundary going in and coming out, and that is where the per-row allocations live. - **Static field access.** `FieldByIndex` with a cached path is much cheaper than searching, but it is still a call resolving an offset at run time, and it does not inline. - **Compile-time checking.** A struct field renamed tomorrow still fails at run time, on row one of the nightly batch, not at build time. That residue is precisely the argument for the alternatives: generated per-type mappers or type parameters remove the dynamism itself. But the plan cache is where you start, because it is a contained change you can make and measure inside one package, and it is the version whose numbers tell you whether the bigger change is worth its price. ## The interview tell Candidates who have actually done this volunteer three things without prompting: that the plan is keyed by `reflect.Type` rather than by a name, that `reflect.Value` must stay out of the cache, and that the cache removes discovery but not boxing. Candidates who have only read about it usually claim the cache makes reflection "about as fast as static code", which the numbers do not support.
- Why is `LoadOrStore` preferable to checking the cache and then storing unconditionally?Because two goroutines can both miss and both build a plan for the same type. `LoadOrStore` publishes one and returns whichever won, so every caller for that type shares a single plan afterwards. An unconditional `Store` would let the second writer replace a plan the first caller is already using, which is harmless for an idempotent plan but wasteful and makes identity comparisons unreliable.
- Could you key the cache on the type's name instead of on reflect.Type?You should not. Two types in different packages can share a name, and unnamed types such as `[]byte` or an anonymous struct have no name at all, so a name key collides or comes up empty. `reflect.Type` values are canonical for each Go type, so `==` on them is exactly the identity you want, and using them directly also skips the string work.
- After adding the plan cache the batch is still slower than a hand-written mapper. Where has the remaining cost gone?Into the parts the cache cannot touch: values still crossing the `any` boundary on the way in and out, which allocates per field, and field access that still resolves an offset through a non-inlinable call. That residue is what generated code or type parameters remove, and it is the number to weigh against the cost of adding a code-generation step.
saying these in an interview costs you the question
- Caches reflect.Value objects rather than field-index paths
- Keys the cache on the type's name string
- Fills a plain map from several goroutines with no synchronisation
- Claims the cache makes reflection as fast as static code
- Rebuilds the plan per row but calls it caching because it is a local variable