When a tier cannot afford even a paced traversal, how do you produce a per-prefix memory breakdown off the serving path?
answer
- move the cost off the serving path
- a node that serves no traffic first
- taking the copy is not free
- on-disk size is not resident memory
- the copy is production data
basics
~20 sAnalyse a copy rather than the live tier: a point-in-time whole copy, a following copy that serves no traffic, or an operator-interface export, walked on another machine. It costs staleness, taking the copy, and moving production data.
solid answer
~50 sMove the walk to something that is not serving. Three routes exist: run the paced traversal against a **following copy of the tier** that answers no reads; walk a **point-in-time whole copy** on another machine; or take whatever **export an operator interface** offers where you cannot reach the host. Each lands the cost of the measurement where nobody is waiting, and each charges for it. Taking a whole copy is not free — it can add a copy-on-write pause to serving and hold extra memory while it runs — so take it from a node that serves no traffic where one exists. The breakdown is then as of the copy's timestamp, understating a fast-growing prefix, and sizes read from a copy reflect a serialized form rather than memory residency, so treat them as relative shape. Some stores in this class keep nothing on disk, and many deployments switch it off deliberately.
go deeper
The idea is simple: analyse something that is not answering traffic. A copy of the data can be walked anywhere without the serving tier paying for it.
Name the three routes — a following copy that serves no reads, a point-in-time whole copy, an operator interface export — and say what each one costs.
Price taking the copy, including the copy-on-write pause and the extra memory it holds while running, and say why on-disk sizes give relative shape rather than resident bytes.
Set the standing rule: where these copies may be analysed, who may hold one, how long it lives, and whether the tier's data class permits this route at all.
## Why you would leave the serving path at all A paced traversal is usually enough. You leave it behind for one of three reasons: the tier has no spare capacity even for a slow walk; you cannot reach the host and the operator interface offers something better than a client-side walk; or the walk needs to be repeated regularly and you would rather it never touched production. ## Three routes 1. **A following copy of the tier.** Point the paced walk at a node that follows the primary and answers no reads. The measurement's cost lands on a machine nobody is waiting on. Two caveats: the node runs slightly behind, and if it is quietly serving reads it is not off the serving path at all. 2. **A point-in-time whole copy.** If the deployment already writes one on a schedule, walking it on another machine costs the serving tier nothing extra. If it does not, you are asking for one to be taken, which is a real charge against the tier. 3. **An export from an operator interface.** On a managed instance you may have no host access and no file. What the operator interface offers — an export, a key inventory, a downloadable copy — may be the only route that does not run through the serving process. ## What taking a copy actually costs - **The copy-on-write pause.** Taking a whole copy of a large in-memory dataset can stall serving briefly while the copy is established, and the process can hold noticeably more memory for as long as the copy is being written, because pages changed during the write exist twice. On a tier already near its ceiling, that is precisely the wrong moment. - **Where you take it from.** Take it from a node that serves no traffic where the shape allows. Taking it from the primary at peak is the version of this answer that gets someone paged. - **Whether it exists at all.** Some stores in this class produce no on-disk copy whatsoever, and many deployments disable it deliberately because the tier holds only replaceable data. Assuming a copy is lying around is assuming one product's configuration. ## What the breakdown loses | Loss | Why it happens | What to do about it | |---|---|---| | Staleness | The copy is as of its timestamp | Compare the age against the growth rate you already measured | | Bytes are not memory | On-disk form is serialized, not resident | Report relative shape — which prefixes dominate — not exact bytes | | One node only | Each node holds its own entries | Take and merge one copy per node, and keep the node in the report | | Data leaves the box | The copy contains production values | Treat it as production data wherever it is analysed | The second row is the one most often got wrong. A copy on disk tells you reliably that one prefix is ten times another; it does not tell you how many bytes of the memory ceiling either occupies, because the representation in memory and the representation in the copy are different things. If the question is "who should shrink", relative shape is enough. If the question is "how much headroom do we get back", it is not, and you need a figure from the running store. ## Handling the data A volatile tier frequently holds sessions, claims, tokens and personal data. A copy of it is a copy of all of that. Walking it on a workstation moves production data onto a workstation, and an analysis run "just to find the big prefixes" is not an exemption. Do the walk somewhere the data is already allowed to be, keep only the rollup, and destroy the copy afterwards. If the only thing you actually need is key names and sizes, prefer a route that never materialises values. ## Choosing between routes 1. Is there a node that serves no traffic? Run the paced walk there — it is live, it is exact, and it costs nothing you care about. 2. Is a whole copy already being written on a schedule? Walk that, and accept relative shape. 3. No host access? See what the operator interface will export, and otherwise fall back to the paced walk from a client, which needs only a connection. 4. None of the above, or the tier holds nothing on disk by design? The paced walk is the answer, slowed down until it is affordable, or the callers instrument themselves.
- The instance is managed and you have no host access. What is left?Whatever the operator interface exports, and the paced walk itself, which needs nothing but a connection and is usually the route that survives on managed instances. Failing both, the callers that write the keys instrument themselves. Anything requiring the process, its files or the host is simply unavailable.
- How stale is too stale?Judge it against the growth rate the last breakdown gave you. If headroom is a fortnight and nothing doubles in a week, last night's copy is fine. If a prefix can double in an hour, a breakdown from a copy taken hours ago names yesterday's holder and you need a live paced walk.
saying these in an interview costs you the question
- Takes a whole copy from the primary at peak without pricing it
- Quotes sizes read from a copy as memory occupied
- Assumes every deployment has an on-disk copy to analyse
- Reports one node's copy as the whole tier's breakdown
- Moves production data to a workstation to analyse it
- Uses a stale breakdown during a fast-growing incident