skip to content

A search indexer with one copy and a bound 200-gigabyte single-writer volume is scaled to three copies, so why do two never start?

level: seniorimportance: must knowfreq 55%

answer

  1. count the requests, not the copies
  2. bound once, held once
  3. refused at attach, before start
  4. same host can attach anyway
  5. own store, shared store, or one writer

basics

~20 s

One request bound one store, and that store honours a single writer at a time. The first copy attaches it; the other two are refused at attach and wait, never started. Scaling the workload did not multiply the request.

solid answer

~40 s

Scaling changed the replica count, not the storage. There is still exactly one request bound to exactly one backing store, and that store was requested as single-writer — it attaches to one node at a time. Copy one takes it and runs. Copies two and three ask for the same store, the attach is refused, and they sit unstarted with no logs. Three real fixes exist: give **each copy its own request**, so each writes its own volume; move to a backing store that honours **many writers**, and then solve concurrent writes in the application; or keep one writer and let the extra copies serve reads without that volume. Note the failure can look intermittent, because a copy landing on the same host as the holder may attach successfully.

code

pseudocode · 17 lines
pseudocode
request = { requestedSize: 200GB, writers: "single" }
volume = bind(request)          # one request binds one store
volume.attachedHost = none

for each copy in workload.copies:          # three copies now
    if volume.writers == "single"
       and volume.attachedHost != none
       and volume.attachedHost != copy.host:
        refuse attach(volume, copy)
        copy.state = "waiting, never started"
    else:
        attach(volume, copy)
        volume.attachedHost = copy.host
        copy.state = "running"

# copy 1 attaches and sets attachedHost; copies 2 and 3 take the refusal branch
# unless they landed on that same host

go deeper

for a junior

Hold on to the counting: raising the replica count does not create more storage. One request means one backing store, and a single-writer store can be held in one place at a time.

for a middle

Explain the refusal at attach and why it differs from a crash loop — no process ever ran, so there are no logs. Then name the fixes rather than just the cause.

for a senior

Show the triage instinct: one copy running while the rest sit before start is the signature. Raise the same-host variant that makes it look flaky, and say which fix you would pick for a workload that owns an index.

for a principal

The real decision is whether this workload should fan out writes at all. Per-copy stores, a shared writable store and a single distinguished writer are three different architectures, and only one of them is a spec change.

## What the numbers are Start by stating them, because the arithmetic is the answer. Before the change: one copy, one storage request, one bound 200-gigabyte store, one writer. After: **three** copies, still **one** request, still **one** store, still **one** writer allowed. One copy attaches and runs; **two** are refused at attach and never start. Nothing about raising the replica count creates a second store. ## Why the other two are refused A single-writer store is one that can be held for writing in one place at a time. When the second copy is placed, the platform tries to attach the already-bound store to that copy's node, and the store is currently held by the first copy's node. The attach is refused, and the copy stays in a pre-start state: no process, no logs, nothing to grep. That is a meaningfully different symptom from the ones next to it: - **A crash loop** means a process ran, wrote logs and exited. Here nothing ran. - **An unbound request** means the store was never found or created. Here the request is bound — one copy is happily using it. - **A permissions problem** means the process ran and its writes were rejected. Again, nothing ran. So "bound, one copy running, others stuck before start" is almost a signature for the single-writer limit. ## The intermittent-looking variant Single-writer usually means **single-node**, not strictly single-container. If the scheduler happens to put the second copy on the same host as the first, the device is already attached there and that copy may start fine. Teams hit this and conclude the problem is flaky, when what varies is placement. Say this out loud in an interview — it is the difference between having read about the limit and having debugged it. ## Three ways out, and what each costs 1. **One request per copy.** Each copy gets its own storage request bound to its own store, and each writes only its own data. This is the standard shape for a workload that owns data, and it scales cleanly. The cost is that the copies no longer share a dataset: an index built by one copy is not visible to the others, so something upstream must give each copy the work or the data it needs. 2. **A store that honours many writers.** Re-request storage backed by something reachable by many clients at once. All three copies can then write one tree. The cost is twofold: every operation crosses the network, and — the part that bites — concurrent access is not coordination. Three copies appending to one index will corrupt it unless the application locks or partitions the paths. Moving to this store is also a data migration, not a spec edit. 3. **Keep one writer.** Let exactly one copy hold the volume and build the index, and let the extra copies serve queries from their own copy of the data or from a read-only mount where the platform supports one. The cost is that the writer is now a distinguished copy, and something must decide which one it is. ## Choosing between them | Option | Fan-out achieved | Migration needed | New failure mode bought | |---|---|---|---| | One request per copy | full write fan-out | no, but data is not shared | divergent per-copy datasets | | Many-writer store | full write fan-out | yes, copy the data | concurrent writers corrupting shared files | | Single writer, read fan-out | reads only | no | one copy is special and must be chosen | For a search indexer specifically, option 1 or option 3 is almost always the honest answer, because an index is exactly the kind of structure that assumes one process owns its files. Option 2 looks like the smallest change on paper and is the one most likely to end in corruption. ## What this question is not about It is not about where the platform chose to put the copies, and it is not about whether the copies have stable names or start in a fixed order — those are separate subjects with their own machinery. Keep the answer on the request-and-attach model: one request, one store, one writer, two copies with nowhere to attach.

  • Would requesting a larger volume help the two stuck copies?
    No. Size and writer count are independent properties of the request. A 2-terabyte single-writer store is still held in one place at a time, so the second and third copies are refused for exactly the same reason. Growing the request changes how much the one writer can store, not how many copies may write.
  • What symptom distinguishes this from the same workload having an unbound request?
    Whether anything is running. With a single-writer conflict, the request is bound and one copy is serving normally while the others sit before start. With an unbound request, no copy starts at all and the request itself reports no store. One copy up and the rest stuck is the single-writer signature.
  • The team reports the problem is intermittent — sometimes a second copy starts. What explains that?
    Placement. A single-writer store is typically held per node rather than per container, so a second copy that lands on the host already holding the device can often attach and run. Nothing in the spec changed between the good and bad runs; where the copy landed did. It also means the workload is quietly running two writers when it does start, which is the worse outcome.

It is a locker with one key: the first copy took the key, and the other two are standing at the locker room door rather than having failed at anything of their own.

saying these in an interview costs you the question

  • Thinks scaling the workload creates a volume per copy
  • Says the stuck copies are crash-looping and reads their logs
  • Proposes enlarging the volume to admit more writers
  • Assumes a many-writer store fixes concurrent index corruption
  • Believes the platform detaches the first copy for the second