skip to content

How should a platform team set the default storage request it offers — size, writer count, room to grow — when those choices resist reversal?

level: principalimportance: nice to knowfreq 30%

answer

  1. pick the reversible direction
  2. grow yes, shrink no
  3. writer count means migration
  4. small default needs a watcher
  5. design the exception route

basics

~20 s

Default to a single writer, a deliberately modest size, and expansion enabled, because growth is usually possible and shrinking and changing writer count are not. Then make used-versus-requested capacity a watched signal, since a full volume hurts more than unused capacity costs.

solid answer

~40 s

Both numbers are asymmetric, and the default should lean towards the reversible direction. Size can usually be grown on a bound store and essentially never shrunk, so a modest default with **expansion enabled** beats a generous default that nobody dares revisit — provided the platform watches used capacity against the request, because a full volume takes workloads down. Writer count is the harder one: moving from single-writer to shared storage is a data migration, not a spec edit, so defaulting to **one writer per copy** keeps the cheap, widely available option and forces a conversation for the exception. State the trade-off you are choosing: you are buying occasional unused capacity and an occasional migration in exchange for avoiding the two failures that hurt most — a full volume and a scaling dead end.

go deeper

for a junior

The takeaway to carry: storage requests are easy to grow and effectively impossible to shrink, so a request is not a number to inflate for safety.

for a middle

Be able to explain why expansion is the reversible direction and why changing the writer count is not, since one is a property of a bound store and the other decides which store you are on.

for a senior

Argue a default from the failure modes: an unbound request fails safely before start, while a full volume fails a running workload mid-write and reads as an application problem for the first twenty minutes.

for a principal

Own the trade-off explicitly. Name what you are buying — occasional expansions and occasional migrations — name the monitoring obligation the default creates, and design the exception route so teams do not route around the standard.

## Why this is a standard, not a per-workload choice Every team on the platform writes a storage request, and almost none of them will think about it again until something breaks. The default you hand them therefore decides, at estate scale, how often you meet a full volume, how often a team discovers at scale-out time that it cannot add a writer, and how much capacity you are paying for that nobody uses. That makes it a platform decision with a defensible rationale, not a field. ## The asymmetries that drive it - **Size grows, it does not shrink.** Most platforms can expand a bound store in place where the backing technology supports it, and the filesystem on the device then has to be extended too — sometimes while running, sometimes needing a restart. What essentially no platform offers is shrinking. Getting the size too low costs you an expansion; too high costs you money forever. - **Writer count is close to irreversible.** Going from a single-writer store to one that honours many writers means a different backing store, which means copying the data and cutting over. It is a project, not an edit. - **A full volume is worse than an unbound request.** An unbound request fails before anything starts, loudly and safely. A volume that fills at 03:00 fails a running workload mid-write, and the failure surfaces as application errors that nobody attributes to storage for a while. - **Defaults are sticky.** Whatever you ship as the default becomes what most teams run, including teams for whom it is wrong. Choose it as if it were the only option, then design the exception path. ## What a defensible default offering looks like 1. **One writer, per copy.** Cheap, fast, offered by every backing technology, and it forces the scale-out conversation to happen at design time rather than at incident time. 2. **A modest size with expansion enabled.** Small enough that unused capacity is not the estate's largest line item, with growth available without a migration. 3. **Used-versus-requested capacity as a watched signal.** This is the control that makes a modest default safe. Without it, "start small and grow" is a promise you have no way to keep. 4. **One or two named kinds of store, not a catalogue.** Every additional kind is another way for a request to be unsatisfiable in one environment and fine in another. | Choice | Reversible? | What a wrong default costs | |---|---|---| | Requested size, too small | yes — expand | one expansion, or an outage if unwatched | | Requested size, too large | no | unused capacity paid for indefinitely | | Single writer, wrong | no — migration | a scale-out blocked until data moves | | Many writers, wrong | partially | latency on every operation, plus a coordination story nobody wrote | ## The exception path matters as much as the default A standard with no exception route is one teams route around. Make the many-writer request available, and make it require a short conversation in which the team says how concurrent writers will be coordinated — because the storage layer will not do it. That conversation is the actual value of the standard: it catches, before the data exists, the workloads that were about to point several copies at one shared tree and hope. The same applies to size. Allow a large request where the team can state the growth rate that justifies it, and treat "we might need it" as a reason to rely on expansion instead. ## What you are explicitly buying Say the trade-off out loud, because that is what distinguishes a lead's answer from a preference: - You are accepting **occasional expansions** to avoid paying for capacity nobody uses. - You are accepting **occasional migrations** for the minority of workloads that genuinely need many writers, to keep the majority on the cheap, simple, universally available option. - You are accepting a **monitoring obligation**, because a modest default is only safe if someone sees the volume filling. - You are **not** accepting the failure that is hardest to fix — a workload whose data is already large and whose writer count is wrong. ## How to answer this in an interview There is no single right default, and claiming one is the weak answer. The strong answer names the asymmetries, picks a direction from them, states the control that makes the pick safe, and describes the exception route. A candidate who can also say which failure they are deliberately choosing to accept is demonstrating exactly the judgment the question is for.

  • What makes 'start small and expand' unsafe without any other change?
    It quietly hands every team an operational obligation nobody took. Expansion is possible, but somebody has to notice the volume filling in time to do it — and expansion may require the filesystem on the device to be extended and, on some platforms, the workload to restart. Without a watched used-versus-requested signal, a modest default simply converts a cost problem into an outage.
  • Why not default to a store that honours many writers so nobody is ever blocked?
    Because it is a worse default for most workloads. Every operation crosses the network, the store is not offered everywhere, and it grants concurrent access without any coordination — so a workload designed to own its files corrupts them instead of failing to scale. You would trade a rare, visible migration for a common, silent hazard.
  • How do you handle a team that already requested far more capacity than it uses?
    Accept it for now and fix the default going forward, because shrinking a bound store is generally not available; recovering that capacity means provisioning a smaller store, copying the data and cutting over. That cost is exactly the argument for a modest default with expansion, and it is worth showing to the next team that asks for a generous request 'to be safe'.

saying these in an interview costs you the question

  • Believes a bound volume can be shrunk as easily as grown
  • Defaults to many writers so teams are never blocked
  • Ships a small default with no capacity monitoring
  • Treats changing writer count as a spec edit
  • Offers a large catalogue of store kinds as flexibility