skip to content

What must a recurring per-prefix memory report contain for a team to act on its share before the tier reaches its ceiling?

level: principalimportance: nice to knowfreq 35%

answer

  1. numbers alone change nothing
  2. an owner per line
  3. a date, not just a share
  4. print the unattributed remainder
  5. cadence set by growth rate

basics

~20 s

Attributed bytes and entry count per prefix as a share of the ceiling, an explicit unattributed remainder, the change since the last run with a projected exhaustion date, a named owner, and the vantage the numbers were taken from.

solid answer

~50 s

A breakdown only changes behaviour if it arrives early, names someone, and carries a deadline. So the report carries five things: **bytes and entry count per prefix** with each prefix's share of the memory ceiling; an **explicit unattributed remainder** rather than silently dropped keys; the **change since the last run** and the date headroom runs out at that rate; a **named owner** per prefix; and the **vantage** — which node, at what time, full walk or scaled sample, sized by the server or by measuring fetched values. Cadence is set by growth, not by the calendar: short enough that the fastest-growing prefix cannot consume the remaining headroom between two runs. What the report deliberately does *not* decide is what happens to an owner who overruns — enforcement is a separate conversation about how the tier is shared. The report's job is to make the fact undeniable while there is still room to act on it.

go deeper

for a junior

Notice what turns numbers into action: every line needs an owner and a date, not just a percentage of the total.

for a middle

Explain why entry count belongs beside bytes, and why keys matching no prefix are reported as a remainder rather than discarded.

for a senior

Derive the cadence from the measured growth rate, keep the vantage and sizing method identical between runs, and state the report's error bar explicitly.

for a principal

Own the framing: attribution accounts, it does not enforce. Keep the enforcement decision separate, or the numbers will be argued away along with the proposal attached to them.

## What the report is for An attribution walk produces numbers. A report is what makes somebody change their code. The difference between the two is almost entirely in the framing, and a principal is expected to own the framing rather than the walk. The failure mode is familiar: a beautifully produced breakdown, generated during the incident, showing that a team you have never met holds sixty percent of a tier that is already refusing writes. Every fact in it is correct and none of it is useful, because the decision it should have informed was taken by the store an hour ago. ## What it has to carry 1. **Bytes and entry count per prefix, with share of the memory ceiling.** Bytes are the ask; the count is the diagnosis. Rising bytes with a flat count means entries got bigger; a rising count with a flat average means a population grew. The remedies have nothing in common, and a report with only one column forces the reader to guess. 2. **An explicit unattributed remainder.** Keys matching no convention, plus the gap between your summed estimates and the server's own total. Dropping these makes the report look tidy and quietly wrong. 3. **Change since the last run, and a projected exhaustion date.** A share is an opinion; a date is a deadline. "Nineteen percent, growing four points a month, headroom gone in March" is a sentence a team acts on. 4. **A named owner per prefix.** A prefix with no owner is a line item nobody reads twice. This is also the moment the naming convention is audited: an unowned prefix is a finding. 5. **The vantage.** Which node, at what wall-clock window, full walk or scaled sample at what rate, sized from the server's own estimates or from measuring fetched values. Without this the reader cannot discount the numbers correctly, and a reader who discovers the omission later discounts them to zero. | Column | Why it is there | |---|---| | Bytes, and share of the ceiling | the size of the ask | | Entry count, and derived average size | separates a growing population from growing entries | | Change since last run, projected exhaustion | turns a number into a deadline | | Named owner | a report with no owner changes nothing | | Vantage: node, window, full or sampled | lets the reader discount it correctly | | Unattributed remainder | measures the convention's decay and your own error | ## The unattributed remainder is the honest part It has two independent causes and they need separating before anyone acts: - **Keys written outside the convention** — a new service, ad-hoc keys from a migration, a prefix nobody retired. This has an owner and is fixable, and a remainder that grows every month is the convention decaying in plain sight. - **Measurement error** — the difference between summed per-entry estimates and the total the server reports. This is not anyone's fault, but it bounds how much precision the rest of the report deserves. A report that folds both into one number invites an argument about whether the whole thing can be trusted. A report that names both invites an argument about the data, which is the argument you wanted. ## Timing and cadence Set cadence against growth, not the calendar. The test: could the fastest-growing prefix consume the remaining headroom between two runs? If headroom is a fortnight and something doubles weekly, monthly reporting is decorative. Two further scheduling points are worth owning explicitly: - Run the walk the **same way every time** — same vantage, same sizing method, same sampling rate — because the trend is what people act on, and a trend built from two different measurements is noise. - Deliberately run it **when nothing is wrong**. A procedure first attempted under pressure is the one that gets the batch size wrong. ## What the report does not decide Attribution is an accounting tool. It says who holds what; it bounds nobody's damage and grants nobody a guarantee. Whether an overrunning owner gets a quota, a conversation or an instance of their own is a separate decision about how the tier is shared, and one that varies with what the other occupants are holding — replaceable copies, or state that exists nowhere else. Keep the two apart. A report that arrives with an enforcement proposal attached gets argued with on the proposal, and the numbers never get read at all.

  • The unattributed remainder is a third of the stored-data size. What does that tell you?
    That one of two things is badly out, and you separate them before acting. Either large populations of keys are being written outside the convention, which has an owner and a fix, or your per-entry estimates diverge from the server's total, which is measurement error and caps the precision of every other line.
  • How often should the report run?
    Often enough that the fastest-growing prefix cannot eat the remaining headroom between two runs. Derive the interval from the growth you already measured rather than defaulting to monthly, and keep the method identical run to run so the trend means something.
  • Why not simply alert on the per-prefix numbers instead of reporting them?
    Because there is no action a machine can take on "this team holds more than last month", and an alert nobody can act on trains people to ignore alerts. The ceiling and the process footprint are what deserve a page; the attribution is what deserves a meeting, early.

saying these in an interview costs you the question

  • Publishes prefix totals with no owner attached
  • Drops keys matching no convention instead of reporting them
  • Reports bytes only, so a growing population looks like growing entries
  • Produces the breakdown during the incident and calls it capacity planning
  • Sets the cadence by the calendar rather than by the growth rate
  • Presents attribution as though it bounded anyone's damage