An in-memory LRU cache budgets its entries with unsafe.Sizeof and the process is OOM-killed in production; what went wrong and how do you fix the accounting?
answer
- the number counted only the words in the struct
- nothing behind the pointers was ever charged
- charge capacity, not length
- make the build fail when the struct changes
basics
~20 sunsafe.Sizeof charges each entry only for its fixed words, not for the bytes behind its string and slice pointers, so the budget undercounts by orders of magnitude. Charge len of the key and cap of the payload as well, then check the total against measured heap usage.
solid answer
~50 sThe cache is charging every entry the same constant — the struct's shallow size — while the memory that actually fills the heap sits behind the pointers inside it. A 48-byte entry holding a one-megabyte payload is billed as 48 bytes, so a budget of a few hundred megabytes admits many gigabytes. The fix is a real size function: `unsafe.Sizeof(e)` for the struct, plus `len(e.key)` for the string's bytes, plus `cap(e.data)` — capacity, not length, because the whole backing array stays alive — plus a measured allowance for any map. Then validate it: fill the cache under a representative workload and compare the accounted total against a heap profile, treating a persistent gap as unaccounted structure rather than a number to fudge. Finally, pin the struct with a compile-time size assertion so that adding a field fails the build and forces the size function to be revisited.
code
go · 10 linestype entry struct {
key string
data []byte
meta map[string]string
}
// If entry stops being 48 bytes, one of these lengths is negative
// and the build fails. Update size() when it does.
var _ [unsafe.Sizeof(entry{}) - 48]byte
var _ [48 - unsafe.Sizeof(entry{})]bytego deeper
The lesson to carry away is that a size taken from unsafe.Sizeof is the struct's own size, so a cache that adds it up is counting handles rather than data.
Be able to write the deep size term by term — struct size, string length, slice capacity, an estimate for the map — and to say why capacity rather than length is the right charge.
Show the diagnosis and the guard: identify the undercount, calibrate the model against a heap profile under real load, and pin the struct with a compile-time assertion so a later field addition fails the build rather than the service.
Own whether the service models memory at all. Byte accounting is a permanent maintenance cost that drifts with every struct change; weigh it against bounding by count, a soft memory limit and an alert on the gap between the accounted and measured totals.
## What the budget actually measured `unsafe.Sizeof` is a compile-time constant describing the value's own representation. For an entry like ```go type entry struct { key string // 16: pointer + length data []byte // 24: pointer + length + capacity meta map[string]string // 8: one pointer } ``` it is 48 bytes on a 64-bit platform and stays 48 forever. Every field is a handle; none of the memory they refer to is inside the struct. So a cache that adds `unsafe.Sizeof(e)` to a running total on insert is counting the *bookkeeping* and ignoring the *payload*. With one-megabyte payloads the model is wrong by a factor of about twenty thousand, which is why the limit never triggered and the kernel's OOM killer had the last word. A related trap makes it worse: the number is a compile-time constant, so it is also the number a reviewer sees and believes. Nothing at run time can contradict it. ## Building an honest size function A deep size is written by hand, term by term: ```go func size(e entry) int { n := int(unsafe.Sizeof(e)) n += len(e.key) // the string's backing bytes n += cap(e.data) // the whole backing array, not just the used part for k, v := range e.meta { n += len(k) + len(v) + int(unsafe.Sizeof(""))*2 } return n } ``` Three judgment calls are hiding in there. **Capacity, not length.** A slice keeps its entire backing array alive, and a slice grown by `append` usually carries spare capacity beyond its length. Charging `len` understates what the entry pins. **Maps cannot be derived.** A map's runtime hash table allocates room for more slots than it currently holds, and the per-entry overhead is an implementation detail that changes between releases. Estimate it from measurement and mark the estimate as such. **Allocator rounding is real.** The runtime rounds each allocation up to a size class, so a 33-byte key does not cost 33 bytes. For a cache holding millions of small entries this alone can be a double-digit percentage. The conclusion most teams reach is that the arithmetic should be calibrated, not trusted: run a representative workload, take a heap profile at steady state, and compare the accounted total with what the profile says. A stable ratio means the model is usable with a safety factor; a growing gap means something is unaccounted. ## Stopping the model from drifting The failure that brought the service down is not really the wrong formula — it is that nothing connects the formula to the struct. Someone adds a field, the size function is not updated, and the budget quietly becomes fiction again. Because `unsafe.Sizeof` is a compile-time constant, you can wire the two together and let the compiler enforce it: ```go // The build fails if entry stops being 48 bytes. var _ [unsafe.Sizeof(entry{}) - 48]byte var _ [48 - unsafe.Sizeof(entry{})]byte ``` Both declarations are needed. If the struct grows, the first is a legal positive length but the second is a negative constant that cannot be represented as `uintptr`, and the build stops. If it shrinks, the roles swap. A comment next to the pair pointing at the size function tells the next person exactly what to update. This is the honest use of the whole family: not measuring memory, but pinning a layout so that a change announces itself at build time instead of at 3am. ## What else to change - **Bound by something you can measure.** Counting entries and enforcing an entry-count limit, tuned by observed heap usage, is often more robust than a byte model, particularly when values vary in size by orders of magnitude. - **Give the runtime a backstop.** A soft memory limit lets the collector push back before the kernel intervenes, turning a hard kill into degraded performance you can alert on. - **Alert on the gap, not just the total.** Export both the accounted total and the measured heap size; the ratio between them is the health of your model, and it is the metric that will warn you the next time a struct changes. ## The summary line for an interview Sizeof told the truth about the struct and nothing about the data. A memory budget needs a runtime measurement of what the pointers reach, calibrated against a heap profile, and a compile-time assertion so the struct cannot change underneath it.
- Why charge cap(e.data) rather than len(e.data)?The slice keeps its entire backing array alive for as long as the entry lives, and that array is `cap` elements long. A slice built by `append` typically has spare capacity beyond its length, so charging `len` understates what the entry actually pins — sometimes by close to a factor of two.
- How do you check that the corrected accounting is close to reality?Fill the cache with representative data and compare the accounted total against measured heap use, from `runtime.ReadMemStats` or a heap profile taken at steady state. Treat a persistent gap as unaccounted structure — map overhead, allocator size-class rounding, per-entry metadata — rather than tuning the constant until the numbers agree.
- What does the compile-time size assertion actually buy you?It converts silent drift into a build failure. Adding a field changes `unsafe.Sizeof(entry{})`, one of the two array declarations becomes a negative length, and the build stops until someone updates both the expected constant and the size function. Without it, the accounting rots the first time the struct is edited.
- When would you abandon byte accounting altogether?When entry sizes are roughly uniform, or when the model's error bars are wider than the safety margin you would apply anyway. Bounding by entry count, tuned from observed heap usage and backed by a soft memory limit, is less code and does not silently rot when a struct changes.
saying these in an interview costs you the question
- Trusts unsafe.Sizeof as an entry's total memory cost
- Charges len of a payload instead of cap
- Assumes a map field's contents are counted by Sizeof
- Never compares the accounted total with measured heap usage
- Leaves the size function unconnected to the struct it models
- Treats the allocator's size-class rounding as negligible for tiny entries