A workload will not start because its 200-gigabyte storage request is still unbound, so how do you work out which cause it is?
answer
- no process, so no logs
- read the request, not the workload
- was anything even attempted
- size is a floor
- bound plus stuck means attach
basics
~20 sRead the request's own status rather than the workload's, because no process ever ran. Three causes account for nearly all of it: no store large enough, nothing configured to create one on demand, or a declared writer count no candidate can honour. Each leaves a different trace.
solid answer
~40 sThe workload has no logs because it was never started, so triage begins at the request. Ask three questions in order. **Was a store attempted at all?** If nothing was created and no candidate was considered, the on-demand provisioning path is missing or the requested kind is one nobody offers. **Was a candidate found and rejected?** Then compare the request against it: an unsatisfiable size and a writer count no candidate can honour are the two rejections that matter, and they are distinguishable because the size is a floor while the writer count is a capability. **Is the request actually bound after all?** Then this is not a binding problem — a bound request with a stuck copy is an attach conflict, which looks similar and is fixed completely differently.
go deeper
The key fact to hold: a workload whose storage request has not bound was never started, so there are no logs to read. The evidence lives on the request.
Explain the three families of cause — no store big enough, nothing to create one on demand, a writer count nobody can honour — and say what distinguishes them in the request's own status.
Show the ordering: is anything running, is the request bound, was a store ever attempted. Then separate the look-alikes, especially a bound request whose copy simply cannot attach.
This is mostly an environment-parity problem. Decide how a spec proves it can bind before it reaches production, so that the first person to meet a missing provisioning path is not the on-call engineer.
## Start where the evidence is An unbound request is a **pre-start** symptom. No process ran, so the workload has no logs, no exit code and no application-level story. Everything you need is in the platform's status record for the request itself: whether it is bound, whether a store was ever created or considered, and what it was asked for. That is the first thing to say in an interview, because the wrong instinct — going to the workload's logs — burns ten minutes and finds an empty file. ## The three causes, and what each one leaves behind | Cause | What the status shows | Why it happens | |---|---|---| | No store big enough | candidates considered, none accepted; nothing created | a pre-provisioned pool whose largest free store is below the request | | Nothing to create storage on demand | no candidate, no creation attempt at all | the on-demand path is not configured, or the requested kind is one nobody offers | | Writer count cannot be honoured | a candidate exists and is rejected on a property, not on size | many writers requested against storage that attaches to one node | The distinguishing evidence is coarse and reliable: 1. **Nothing attempted** points at the provisioning path. This is the classic "a spec that worked in one environment hangs in another" failure: one environment creates stores on demand, the other expects an operator to have made them. 2. **Attempted and refused on size** points at capacity. Remember the request is a floor — a 200-gigabyte request is satisfied by a 256-gigabyte store and never by a 128-gigabyte one, so a pool full of small stores can look busy and still satisfy nothing. 3. **Attempted and refused on a property** points at the writer count or the requested kind. A many-writer request will never bind to a store held on one node at a time. ## The near-miss that is not a binding problem Before going further, confirm the request is genuinely unbound. Two look-alikes are fixed completely differently: - **Bound, but one copy cannot attach.** The request found its store and another copy is holding it. This is the single-writer conflict, and the giveaway is that some copy of the workload is running normally. - **Bound and attached, but the process cannot write.** Then something ran, logs exist, and the problem is ownership and permission bits on the mounted tree — a different subject entirely, and one you can rule out in seconds because a process produced output. A quick ordering that settles all of this: *is anything running?* → *is the request bound?* → *was a store ever attempted?* ## What to do about each - **No capacity.** Either an operator provisions a store that satisfies the floor, or the request is re-issued against a kind the platform can create on demand. Lowering the requested size is a legitimate short-term move only if the workload genuinely fits in less; it is not a fix if it merely delays a full volume. - **No on-demand path.** Configure it, or change the request to name a kind the environment actually offers. This is the failure worth catching before it reaches production, because it is environment-shaped rather than workload-shaped. - **Writer count.** Decide honestly whether the workload needs many writers. Usually the request was written for many writers by habit, and one writer per copy is both available and better. ## Why this is a senior question Because the reasoning chain, not the remedy, is what is being assessed. The candidate who says "check the logs" has not internalised that nothing started. The candidate who says "an unbound request means the platform had nothing to give it, so I read the request's status and ask whether a store was even attempted" has, and everything after that follows from three comparisons against what the request declared. It is also the failure most likely to appear the first time a spec moves between environments, which is exactly when the person debugging it has the least context about how that environment provisions storage.
- The same spec binds immediately in one environment and hangs in another. What is the most likely difference?How storage is satisfied. One environment has a component that creates stores on demand for whatever the request declares; the other expects an operator to have pre-provisioned them, so the request waits for a candidate that nobody made. The second most likely difference is that the requested kind of store is offered in one place and not the other.
- Why is lowering the requested size a risky way to get past a capacity refusal?Because it buys a start, not a fit. If the workload genuinely needed 200 gigabytes, binding it to something smaller only moves the failure to the day the volume fills — and a full volume takes the workload down in a far more confusing way than a request that refused to bind. Shrink the request only when you have a reason to believe the smaller number is honest.
- How do you rule out an ownership or permission problem in the first minute?Ask whether anything ran. Ownership and permission failures happen after the volume is attached and the process has started, so they come with logs and usually an explicit write error. An unbound request produces no process and no output at all, so the presence of any application log already rules it out.
saying these in an interview costs you the question
- Goes to the workload's logs when no process ever ran
- Treats every unbound request as a capacity shortage
- Cannot separate an unbound request from a refused attach
- Lowers the requested size to force a start
- Assumes every environment creates storage on demand