skip to content

How does a setup container in a co-located group hand a data snapshot to the indexer that starts after it?

level: middleimportance: should knowfreq 48%

answer

  1. runs first, exits, then the others start
  2. the handoff is through the shared scratch area
  3. it must exit successfully, not just finish
  4. fresh scratch area means it reruns
  5. failure means the group never starts

basics

~20 s

It is declared as a member that must run to completion first: it writes the snapshot into the group's shared scratch area and exits successfully, and only then does the platform start the long-running members, which read the same area.

solid answer

~40 s

The group declares two kinds of member. The **setup member** runs first and on its own, and has to exit successfully; the long-running members are started only once it has. Both mount the group's **shared scratch area**, so the setup member writes the snapshot there and the indexer reads it from the same place — the handoff is a file on shared storage, not a network call and not an ordering hint. It also means the setup member runs again for every new group instance, because a replacement group gets a fresh scratch area. If the setup member fails, nothing downstream starts at all: the platform applies the group's restart rule to it, and the group simply never reaches a serving state.

code

yaml · 11 lines
yaml
unit: search-indexer
sharedScratch: /work
setupMembers:                        # run in order; each must exit 0
  - name: fetch-snapshot
    image: snapshot-fetcher@<digest>
    writes: /work/index-snapshot
members:
  - name: indexer
    image: indexer@<digest>
    reads: /work/index-snapshot      # started only after setup members exit 0
    listenPort: 8080

go deeper

for a junior

Know the sequence: the setup member runs and exits first, then the long-running members start, and both of them see the same scratch area.

for a middle

Explain the two guarantees you are buying — ordering and successful exit — and that the handoff itself is nothing more exotic than a file in shared storage.

for a senior

Reason about the retry path and the diagnosis: a failing setup member holds the whole group out of service, and the symptom is a group that never becomes ready with no output at all from the main member.

for a principal

Decide what belongs in a setup step at all: preparation that must be atomic and attributable, versus data that should be baked in at build time so no start-up path depends on a remote fetch.

## Two kinds of member in one group A co-located group can declare members that are expected to **run to completion before anything else starts**, alongside the members that run for the life of the group. The first kind is the setup member: it exists to put the environment into a state the long-running members can assume. Fetching a data snapshot, preparing a directory layout, waiting for a dependency to answer, rendering a file from configuration — all of these are work that has to be finished, exactly once, before the real process begins. The platform gives you two guarantees about such a member, and those two guarantees are the whole feature: 1. **Ordering.** The long-running members are not started until the setup member has finished. Where several setup members are declared, the model is sequential: each runs after the previous one, in declaration order. 2. **Successful exit.** Finishing is not enough — the setup member must exit successfully. A non-zero exit does not release the rest of the group. ## The handoff is a file, not a message The setup member and the indexer never talk to each other. They cannot: when the setup member is running, the indexer has not been started, so there is nothing listening. The channel between them is the group's **shared scratch area** — a storage area the group declares once and each member mounts, possibly at different paths. The setup member writes the snapshot there and exits; the indexer starts, finds the file at the path it mounts, and reads it. That is why this mechanism belongs to the group rather than to orchestration in general. Two separate workloads cannot hand files to each other this way at all; two members of one group can, because they are on one host with one declared area between them. | Setup member | Long-running member | |---|---| | runs before the others, alone | runs for the life of the group | | must exit, and exit successfully | is not expected to exit at all | | holds the group out of service while it runs | serves once it is ready | | runs again for every new group instance | restarted in place if it crashes | ## What failure looks like A failing setup member does not produce an application error, because the application was never started. It produces a group that never comes up. The platform applies the group's restart rule to the setup member, typically retrying it with a growing backoff, and the long-running members stay unstarted for as long as that continues. Platforms differ in whether they retry forever or eventually mark the group as failed, so the shape of the end state is not something to assert in the abstract — but the important half is identical everywhere: nothing downstream runs. The practical consequence is where to look. If a group never becomes ready and its main member has no output at all, read the setup member's output, not the main member's. 'No logs from the main process' is itself the clue: it never ran. ## Why it runs again every time The shared scratch area belongs to the group instance. When the group is replaced — a changed image, a new spec, a move after host loss — the new instance gets a fresh area, empty. So the setup member runs again and fetches the snapshot again. This is a real cost to price in: every replacement pays the fetch, and the group's start-up time includes it. If the snapshot is large and rarely changes, that argues for baking it into an image at build time, or for a durable volume rather than scratch storage, rather than for repeating the download on every replacement. ## Why not just do the work in the main member's start-up? You can, and sometimes should. The reasons to split it out are concrete: - **A different image.** The tooling that fetches or prepares the data does not have to ship inside the runtime image, which keeps the long-running image small and its dependency surface narrow. - **A narrower privilege window.** The setup step can be given what it needs to do its job and then exit, while the long-running member runs with less for the whole life of the group. - **Attributable failure.** A setup member that exits non-zero is a distinct, named failure. The same work buried in a start-up script shows up as a main process that died for unclear reasons. - **No half-started process.** The main process is never observed in a state where it is running but its data is not there yet, which is the state that produces confusing readiness behaviour.

  • What happens to the group if the setup member exits non-zero?
    The long-running members never start. The platform applies the group's restart rule to the setup member, so it is usually retried with growing backoff, and the group stays in a not-yet-started state instead of reporting an application error. Platforms differ in whether they retry indefinitely or eventually mark the group failed.
  • If two setup members are declared, do they run at the same time?
    No. The model is sequential: they run one after another in declaration order, and each has to exit successfully before the next begins. The long-running members start only after the last one has finished, which is what lets you express a genuine chain of preparation steps.
  • Can a setup member need privileges the main member does not?
    Yes, and that is one good reason to split the work out. The setup step can be granted what it needs to fetch or prepare data and then exit, while the long-running member runs with a narrower privilege set for the entire life of the group.

The setup member is the stagehand who has to finish dressing the stage before the curtain can go up. If the set never arrives, the performance simply does not begin, and nobody on stage improvises around the gap.

saying these in an interview costs you the question

  • Thinking setup and long-running members run in parallel
  • Expecting the main member to start after a failed setup
  • Assuming the snapshot survives into a replacement group
  • Believing the members coordinate the handoff over the network
  • Treating a setup member as a background helper