A helper container in a running co-located group needs a new image — what does the platform actually replace, and at what cost?
answer
- the unit is the group, not the member
- a changed member means a new instance
- everything instance-specific is renewed
- setup members run again
- release cadences merge across the group
basics
~20 sThe whole group. Changing any member's declared image produces a new group instance, so every member restarts, the group becomes reachable at a new address and its scratch area starts empty — the main process pays for the helper's release.
solid answer
~40 sReplacement is group-granular. A running group instance is fixed: to change one member's image you declare a new spec, and the platform creates a replacement group and removes the old one rather than swapping a member inside it. Everything that was specific to the instance goes with it — the address, the contents of the shared scratch area — and any setup members run again from scratch. So the untouched indexer restarts because its helper shipped. The cost is coupling: the group's release cadence becomes the highest cadence of any member in it, and its restart blast radius is the whole group. That is the real constraint on group membership — the members that must share the address or the scratch area, and nothing else.
go deeper
Remember that you cannot update one container inside a running group; a changed spec replaces the whole group.
Explain what a replacement instance does not inherit — the address, the scratch contents, and the finished work of setup members, which all run again.
Cost it out the way an incident review would: a helper's release schedule becomes the main process's restart schedule, including its cold start and cache warm-up.
Set the rule for the estate. Membership of a group is justified by the shared address or the shared scratch area; every other co-location is an unpriced coupling of release cadence and blast radius.
## Replacement happens to the group The group is the unit the platform manages, and that holds for updates as firmly as it holds for placement and deletion. There is no operation that swaps one member's image inside a running group. What exists is: declare a new spec, and the platform brings up a **new group instance** matching it and removes the old one. Every member in that new instance is started fresh, including the ones whose declaration did not change at all. This surprises people because the mental model underneath is often 'a group is a little machine and the members are services on it' — and on a machine you would restart one service. A group is not that. It is an immutable bundle that gets reissued. ## What a replacement instance does not inherit - **The address.** The new instance is reachable at a new address. Anything that cached the old one is holding a stale value. - **The contents of the shared scratch area.** The new instance gets a fresh, empty area. Whatever the previous instance left there is gone. - **The completed work of setup members.** They run again, from the beginning, for the new instance — including any fetch or preparation step they perform. - **Anything held in a member's writable layer.** Each member starts from its image, as it did the first time. | | In-place member restart | Group replacement | |---|---|---| | Trigger | a member's process exits | a changed spec | | Scope | that member only | every member | | Address | unchanged | new | | Shared scratch area | kept | fresh and empty | | Setup members | not re-run | re-run | Getting this pair the right way round is the whole point. A crash is handled in place and cheaply; a spec change is handled by discarding the instance. Describing a crash as a replacement, or a spec change as a restart, gets both the cost and the surviving state wrong. ## What the coupling costs The interesting consequence is not mechanical, it is operational. Once two components are members of one group: - **Release cadences merge.** If the helper ships weekly and the indexer ships quarterly, the indexer now restarts weekly. For a process with a large in-memory index, a warm cache or a slow start-up, that is the expensive part of the whole system being paid on someone else's schedule. - **Blast radius merges.** A bad helper image does not degrade a helper — it takes the group's instances down as they roll, and the main workload with them. - **Rollback merges.** Reverting the helper means reverting the group's spec, so the main member is restarted a second time to undo a change that was never about it. - **Start-up cost is paid again.** Every replacement re-runs setup members and re-warms anything the members build at start-up. None of that is an argument against co-location. It is the price of the shared address and the shared scratch area, and when a member genuinely needs those, the price is worth paying. ## How to decide what goes in the group The question to ask about any candidate member is narrow: **does it need the group's address, or the group's scratch area?** A component that must be reachable as the same address as the main workload, or that must read and write the same files, has a real claim. A component that only needs to run somewhere near the workload does not — it can be its own workload, released and restarted on its own schedule, without dragging the main process along. When the answer is yes but the cadence mismatch is painful, the usual moves are to stabilise the frequently-changing member — take its behaviour from configuration rather than from a new image where that is possible — or to accept the restarts explicitly and make the main process cheap to restart, which is worth doing anyway. Platforms differ in how a replacement is sequenced: some start the new instance before stopping the old one, others stop first. That choice belongs to the workload's update policy, not to the group model. What does not differ is the unit being started and stopped: it is the entire group, every time.
- Is this the same thing that happens when one member crashes?No. A crash is handled in place: the node agent restarts that member while the others keep running, and the group keeps its address and its scratch area. Replacement is what a changed spec triggers, and it discards the entire instance, address and scratch contents included.
- Does the replacement group start before the old one stops?That depends on the workload's update policy rather than on the group model — some replacements start the new instance before stopping the old, others stop first and accept a gap. Either way, the thing being started and stopped is the whole group, not one member of it.
- How should this shape what you put in a group?Admit only what needs the shared address or the shared scratch area. A helper that releases weekly next to a process with an expensive cold start means weekly cold starts, and a group is the wrong home for two components whose lifetimes have nothing to do with each other.
saying these in an interview costs you the question
- Expecting the platform to swap one member in place
- Assuming the untouched members keep running
- Thinking the group keeps its address across a replacement
- Believing the shared scratch area carries over
- Packing unrelated components into one group for convenience