skip to content

Why must a persisted-operation manifest reach the server before the client build that uses it ships?

level: middleimportance: should knowfreq 46%

answer

  1. Order of operations, not code
  2. The key is useless before the map
  3. No text left to fall back on
  4. Publish, read back, then release
  5. Cumulative registry, never a snapshot

basics

~20 s

Because the identifier means nothing until the server holds the mapping, and the shipped client has no document text to fall back on. Ship the client first and every one of its requests fails until the manifest lands.

solid answer

~50 s

The identifier is only useful if the server can resolve it, and a build-time registry deliberately gives the client no way to supply the text the server is missing — that text is not in the artefact. So publishing late fails **totally**, not partially: from the moment the build is reachable until the manifest lands, every operation it sends comes back unresolvable, and a retry sends the same unknown key. The pipeline order is fixed: publish the manifest, read it back through the path a request would use, then release the client. Two details decide whether that holds. A manifest baked into the server image is only partly published during a rolling deploy, so failures become intermittent and look like flakiness — a store every instance reads avoids that. And the registry must be cumulative, because a rollback that restores an older manifest un-publishes identifiers live clients are still sending.

code

pseudocode · 14 lines
pseudocode
# release pipeline: publication is a gate, not a step
manifest = buildClientArtefact().manifest

registry.publish(manifest)                 # additive: never replaces existing entries

deadline = now() + 120.seconds
repeat:
    sample = pickRandom(manifest.ids, 5)
    ok = all(serverEndpoint.canResolve(id) for id in sample)
    if ok: break
    if now() > deadline: failDeploy("manifest not readable by the server")
    wait(2.seconds)

cdn.promote(clientArtefact)                # only now is the build reachable

go deeper

for a junior

Remember the sequence: the manifest reaches the server first, and only then does the client that uses it go live. A client that ships first sends keys nothing can resolve.

for a middle

Explain why the failure is total rather than degraded — the shipped artefact holds no document text — and describe where in a release pipeline the publish step has to sit and how it is verified.

for a senior

Show that you have thought about the fleet and the rollback: a manifest embedded in a server image is only partly published during a rolling deploy, and a rollback that restores an older manifest withdraws identifiers live clients still send.

for a principal

Own the release contract between the client and server teams — which artefact gates which, who owns the registry store and its availability, and how long a build already in the field stays supported.

## The identifier is a promise the server must already be able to keep Give a client an identifier and it will send it. Whether that works depends entirely on whether the server can resolve it, and a build-time registry offers no negotiation: the client cannot supply the document text it is referring to, because the build removed that text from the artefact. That makes late publication a **total** failure, not a degraded one. It is not that some fields go missing or some operations run slowly; every operation from that client version, for every user on it, comes back as an unresolvable identifier from the moment the version is reachable until the manifest lands. There is no partial success to mask it, and usually no retry that helps, because a retry sends the same unknown key. A concrete shape for the incident: the freight console's bundle reached its CDN at 14:07 and the pipeline's manifest-publish step ran at 14:11. Every scan, every consignment lookup and every dispatch confirmation attempted in those four minutes failed — from a change nobody had made to the schema, the resolvers, or the client's logic. The postmortem is about job ordering. ## The ordering that follows Publish the manifest. Confirm the server can read it. Then release the client. In pipeline terms the publish is a **gate**, not a step: the release job writes the manifest, reads it back through the same path a request would use, and only flips the client artefact live once that read-back succeeds. Failing a deploy because the manifest is unreadable is cheap; the alternative is a failure window whose length is whatever the timing between two unrelated jobs happens to be. Two details decide whether that ordering actually holds in production. **Where the registry lives.** If the manifest is baked into the server image and loaded at boot, then "published" is not a moment — it is the duration of a rolling deploy. During the roll some instances hold the new entries and some do not, so a client released in that window fails on a fraction of its requests and succeeds on the rest. That is worse than a clean outage, because intermittent failure reads as flakiness and sends people looking at the network. Publishing to a store every instance reads — with a refreshable cache rather than a boot-time load — turns the roll back into a single event. **Whether the registry is cumulative.** A registry is a union of everything still reachable, not a snapshot of the current build. If the publish step replaces the contents, a server rollback restores an older manifest and un-publishes identifiers that live clients are still sending — the same total failure, arriving during the change that was meant to end an incident. Entries are added, and much later retired deliberately; they are never dropped as a side effect of a deploy. ## "The client ships" is not a moment either Even on the web, releasing a new bundle does not retire the old one. Users hold cached bundles, long-lived tabs keep running the code they loaded this morning, and a CDN may serve the previous artefact until its entry expires. So the union has to span the versions that are actually reachable, not the version you most recently built. Rolling a web client *back* has the same trap in reverse: the older bundle's identifiers must still be in the registry — which they will be, if nothing ever removes entries on deploy. An installed application makes this dramatically worse, because you cannot withdraw it at all. The publish-first rule is unchanged, but the tail stretches from minutes to months, which converts the ordering problem into a retention problem. ## What the client sees, and what to do with it The response to an unknown identifier is not specified. Servers vary between a request error inside the GraphQL envelope — an `errors` entry and no `data` key, since nothing executed — and a plain HTTP 4xx with no envelope at all. Either way it is worth pinning as a distinct case in the client, because the correct handling is not the handling for a failed field. A retry will not help; a background refresh will not help. The situation is "this build cannot talk to this server", and the useful responses are a forced reload for a web client or an upgrade prompt for an installed one. Interviewers like this question because the answer is barely about GraphQL. It is about a coupling between two artefacts that are normally deployed independently, and about noticing that a scheme which removes text from requests has also removed the client's ability to recover on its own.

  • What does the server actually return for an identifier it does not hold?
    It is unspecified, and saying so is part of the answer. Servers vary between a request error inside the GraphQL envelope — an `errors` entry with no `data` key, since nothing executed — and an HTTP 4xx with no envelope at all. Pin one and handle it as its own case in the client: the right response is not a retry, it is 'this build cannot talk to this server'.
  • The bundle went live before the manifest and the incident is over. What do you change so it cannot recur?
    Make publication a gate rather than a step. The release job publishes the manifest, polls until a read-back through the serving path succeeds, and only then promotes the client artefact. Failing the deploy on an unreadable manifest is cheap; the alternative leaves a failure window whose length is whatever the timing between two unrelated jobs happens to be.
  • Does the ordering constraint change for an installed mobile client?
    Publish-first is unchanged, but the tail is far longer. A web bundle stops being served shortly after you roll it back; an installed app does not, so on mobile the ordering problem becomes a retention problem — the registry must keep serving the identifiers of every build still in the field, for as long as those builds exist.

The paperwork has to reach the terminal before the truck does. Arrive first and the driver is standing at the gate with a reference number nobody can look up.

saying these in an interview costs you the question

  • Assumes the client can resend the document text on a miss
  • Thinks a missing entry degrades only some fields
  • Uploads the manifest after the client bundle goes live
  • Replaces the whole registry contents on every deploy
  • Forgets cached bundles keep old identifiers in use
  • Treats a partially rolled fleet as fully published

context