skip to content

Why must each copy of a data-owning workload keep a stable identity and its own storage across replacement?

level: middleimportance: must knowfreq 66%

answer

  1. identity, not interchangeability
  2. the replacement must be the same copy
  3. storage follows the name, not the instance
  4. healthy but empty is the symptom
  5. requested per identity, reattached on replacement

basics

~20 s

A data-owning copy holds data no other copy has, so the replacement is useful only if it comes back as that same copy. A durable per-copy name plus storage bound to that name, not to the instance, is what makes that possible.

solid answer

~40 s

A container platform is built on copies being interchangeable: any instance serves any request, so it can kill one and start another anywhere. A copy that owns data is not interchangeable - it holds a slice of the accepted writes that only it has. Two things have to hold for replacement to work. First, the copy has a durable per-copy identity that is part of the workload's declared shape: `ledger-0` stays `ledger-0`, addressable as itself, whoever the running instance is. Second, storage is requested per identity and reattached to whatever instance currently holds it. Without the second you get a healthy-looking empty copy; without the first nothing can decide which store to reattach. Both together are what turns a replaceable slot into a member that can be replaced without losing what it held.

code

yaml · 10 lines
yaml
workload: ledger
copies: 3
identityPerCopy: true          # ledger-0, ledger-1, ledger-2 outlive their instances
startOrder: sequential
storagePerCopy:
  size: 200Gi
  accessMode: singleWriter
  followsIdentity: true        # reattached to whoever holds the name
scratchPerInstance:
  size: 5Gi                    # discarded with the instance, by design

go deeper

for a junior

Recall that a workload holding data needs storage that outlives the instance, and that a replaced copy starts with nothing unless the platform reattaches something to it.

for a middle

Explain the two mechanics: a durable per-copy name that the replacement takes over, and a storage request made per name so the same backing store comes back. Say what each one fails to do alone.

for a senior

Show that you have seen the failure that looks like success - a replacement that is addressable, reports healthy and holds an empty store - and name the health signal that let it through.

for a principal

Frame it as a decision: which workloads should be allowed to own data on the platform at all, and what standard the organisation sets for the ones that do.

## The bargain the platform is built on A container platform assumes the copies of a workload are **interchangeable**. A control loop compares the declared number of copies against what is actually running, and when one is missing it starts a replacement wherever there is room. Nothing carries over from the dead instance because nothing needed to: the copy held no data of its own, so a fresh instance answers the same requests the same way. Almost everything the platform does well - replacing a failed copy, draining a host, growing and shrinking with load - rests on that assumption. A workload that **owns data** breaks it. A payments ledger copy holds accepted writes that only it has. A message broker copy holds messages only it acknowledged. The replacement instance is worth nothing unless it comes back *as that copy*, with that data. So the platform has to hand back the two properties it normally throws away. ## The two properties a data-owning copy needs 1. **A durable per-copy identity.** Each copy gets a name of its own - `ledger-0`, `ledger-1`, `ledger-2` - that belongs to the declared workload rather than to the process currently running. When an instance dies, the replacement takes that same name and is addressable under it, so peers and operators can talk to *that member* instead of to whichever copy a routing layer happens to pick. This is not the same thing as a stable name in front of the whole set, and it is not the credential a workload presents when it calls the platform's API; it is the copy's own durable label. 2. **Storage bound to the identity, not to the instance.** The storage is requested *per identity*, and the platform reattaches the same backing store to whichever instance currently holds that identity. The instance-scoped writable area still dies with the instance - that never changes. What changes is that the data the workload cares about was never in it. A third property usually travels with these two, and this leaf owns it as well: a **defined start and stop order**, so that a group of copies that must agree about who holds what does not come up all at once with no first member. ## Replaceable against data-owning | | Replaceable copy | Data-owning copy | |---|---|---| | Identity across replacement | none needed; any instance is any copy | durable per-copy name the replacement takes over | | Storage | instance-scoped scratch, discarded with it | requested per identity and reattached | | Placement | any host with capacity | only where that copy's storage can be reached | | "Healthy" means | the process is serving | the process is serving *its own* data | | Losing one | a replacement covers it in seconds | recovery is bounded by the data, not the process | | Scaling in | remove any copy | removing a copy removes the only holder of a slice | ## Where this goes wrong in practice - **One store declared for the whole workload.** Every copy points at the same backing store. If that store permits a single writer, all but one copy sits waiting; if it permits several, the copies quietly overwrite each other's files. - **Identity without storage.** The replacement is addressable as `ledger-1`, reports healthy, and has an empty store. This is the failure that looks like success, and it is usually discovered by a read that returns nothing rather than by an alert. - **Storage without identity.** The data is on a device nobody can map back to a member, so a restore becomes guesswork about which copy owned which store. - **Trusting the health signal.** A process that reports ready when it has bound a port, rather than when it has opened and validated its data, makes an empty copy look identical to a correct one. - **Scaling in as if the workload were stateless.** Removing a copy from a data-owning set removes the only holder of that copy's slice; the platform will do it cheerfully because nothing in the declared shape says otherwise. - **Expecting the image to help.** Redeploying restores the process and none of the data. Code and data recover by different routes and on different clocks. ## What identity and per-copy storage still do not buy They make replacement survivable; they do not make the workload cheap to operate. Placement is now pinned to wherever that storage can be attached, which narrows where the copy can ever run. Backups remain a separate obligation, because a reattached store is the *same* store - it carries a corruption or a bad write straight through the replacement. And the members' own agreement about which of them holds what data is the data engine's problem, not the platform's: the platform can guarantee that `ledger-1` gets `ledger-1`'s bytes back, and nothing beyond that. The practical test for any workload is one question: if this copy is destroyed right now and a fresh instance starts elsewhere, is anything lost? If the answer is no, treat it as replaceable and keep it simple. If the answer is yes, it needs both properties, and it needs the operational work that comes with them.

  • What breaks if a data-owning workload keeps per-copy identity but declares one shared store for all copies?
    The identity stops meaning anything, because every replacement reattaches the same bytes. On a store that permits one writer, only one copy can start and the rest wait; on a store that permits several, the copies write over each other because nothing partitions their data. Per-copy identity is only useful when the storage request is per copy too.
  • During a restore, why does the mapping between a stored copy of the data and a member name matter?
    Because each member's store holds a different slice. Restoring member 2's data into member 1's store gives you a workload that starts, reports healthy and answers from the wrong data - the worst failure shape, since nothing flags it. The mapping is part of what a backup has to record, not something to reconstruct from timestamps.
  • How do you tell whether a workload genuinely owns data or only looks like it does?
    Ask what is lost if this copy is destroyed and a fresh instance starts elsewhere. A cache, a derived index or a rebuildable working set loses nothing that its upstream cannot regenerate, so treat it as replaceable. If the answer names writes that exist nowhere else, it owns data and needs identity, per-copy storage and a backup.

saying these in an interview costs you the question

  • Thinks any replacement instance can serve any copy's data
  • Believes the data lives in the image and a redeploy restores it
  • Declares one shared store and expects each copy to get its own
  • Treats a healthy signal as proof the copy has its own data
  • Confuses a stable name for the whole set with per-copy identity
  • Scales a data-owning workload in as if copies were interchangeable