skip to content

In a service's main, how do you decide which collaborators must be reachable at startup and which may start degraded?

level: principalimportance: nice to knowfreq 31%

answer

  1. can the service still keep its promise?
  2. read acceleration versus the write path
  3. a crash-loop is a signal someone already handles
  4. degraded is fine only when it is visible
  5. write the policy per collaborator, not per incident

basics

~20 s

Refuse to start only for collaborators the service cannot keep its contract without: its store, its listener, its configuration. Optional ones such as a cache may start degraded, provided readiness and logs report that truthfully.

solid answer

~50 s

I classify each collaborator by whether the service is still correct without it. A thumbnailing service without its blob store can only return errors, so failing startup is the honest signal: the rollout stops, the platform's restart and rollback machinery does its job, and no traffic reaches a process that cannot work. A thumbnail cache or a metrics exporter is different — slower or blinder, still correct — so it may come up with a stand-in that no-ops or reconnects lazily. Two conditions make that acceptable: the degradation is visible (startup warning, a running-degraded metric, honest readiness), and it lives behind the same type the wiring passes in, so no call site grows a nil check. It is a decision rather than a preference because starting degraded converts a loud crash-loop the deploy catches into a quiet error rate on-call finds later.

go deeper

for a junior

Know that a service can refuse to start when something it needs is unreachable, and that a clear startup error naming the missing collaborator is more useful than a process that comes up broken.

for a middle

Be able to classify a dependency as required or optional by asking whether the service is still correct without it, and to implement the optional case as a stand-in rather than a nil field checked everywhere.

for a senior

Show the visibility obligations you attach to a degraded start — startup log, a running-degraded metric, honest readiness — and describe bounding any startup retry so a deploy fails rather than stalls.

for a principal

Own the policy across services and name who it costs. Explain that degraded starts move failure from a contained crash-loop to an error budget on-call pays, and that the decision is recorded per collaborator, with an agreed route for on-call to overrule it.

## The question behind the question Every service's entry point ends up asking, for each thing it builds: *if this is unreachable right now, should the process exist at all?* Answer "always fatal" and a cache blip crash-loops the fleet. Answer "always continue" and you ship processes that pass health checks and serve errors. Neither default survives contact with production, so the call is made per collaborator, deliberately, and written down. ## The classification **Hard dependencies** are the ones the service's core contract is defined in terms of. For an image-thumbnailing service that is the blob store the thumbnails live in, the listening socket, and the configuration that says where those are. Without them the process can do nothing correct. Failing startup here is not pessimism, it is honesty: a process that exits is a signal the deployment system already knows how to act on — the rollout halts, the previous version keeps serving, and a human sees a clear message naming the collaborator, the address and the underlying cause. **Soft dependencies** are the ones that change how well the service runs, not whether it is right. A thumbnail cache, a metrics sink, a feature-flag source with sane defaults, a warm-up prefetch. The service is slower, blinder or colder without them and still correct. These may start absent — provided you take the two obligations below seriously. The boundary is not always obvious, and the useful test is the read/write asymmetry: a collaborator on the *write* path is usually hard, because degrading means silently losing work; a collaborator on the *read-acceleration* path is usually soft. Ask what a user experiences: an error, or a slower correct answer? ## Obligation one: the degradation must be visible The genuine danger of degraded startup is not being degraded — it is not knowing. Three things make it honest: - A startup log line at warning level naming exactly what is missing and what behaviour changed. - A metric or gauge that says "running without X", so a dashboard can show a fleet coming up degraded rather than one instance. - Readiness that reflects the truth. If "ready" means "will serve correct responses", a degraded-but-correct service is ready. If the missing piece means some requests will fail, it is not, and the process must not claim otherwise merely to make the deploy go green. The anti-pattern is a health endpoint that reports only that the HTTP server answers. That is what turns a half-built process into a successful deployment. ## Obligation two: the degradation must not deform the graph If "cache unavailable" becomes a nil field, every call site grows a check, and one of them will be forgotten. Keep the wiring's shape constant: substitute a stand-in of the same type that stores nothing and always misses, or one that reconnects lazily in the background. The wiring function makes one decision at one line; nothing downstream changes. That also keeps the degraded path warm, because the same stand-in is what a local run or a test uses — a fallback exercised only during an incident is a fallback that has already rotted. ## Obligation three, the one people skip: bound the waiting "Retry until it comes up" is a third answer that looks generous and behaves badly: the process neither starts nor fails, the deployment sits at a partial rollout, and nothing pages. If you retry at startup, bound it — a small number of attempts over a few seconds — and then take the hard-or-soft decision anyway. ## Who owns this and who can overrule it The service owner writes the policy, but it is not theirs alone, because the choice moves cost between teams. Starting degraded converts a crash-loop — loud, caught by the deploy, contained to a rollout — into an elevated error rate or a cold-cache latency spike that shows up on the on-call's pager and in the error budget. On-call can and should force a change after an incident where a degraded start served hours of bad responses, just as product can force one after an outage where a non-essential dependency took the whole fleet down. The engineering answer is to record the decision per collaborator, next to the wiring code, with the reason — "cache: soft, misses only, readiness stays green, alert on running-degraded > 5 minutes" — so the next incident argues with the note instead of re-deriving it at 3am. ## What a good answer sounds like Name the classification test, give one hard and one soft example from a real service, state the visibility obligation, keep the stand-in behind the same type so the graph is unchanged, bound any retry, and finish by saying who is entitled to overrule you and on what evidence. What a weak answer does is pick a side globally — everything fatal, or everything degraded — and defend it as a principle.

  • How do you stop the degraded path from rotting between incidents?
    Make it the path used every day. The stand-in that no-ops or misses is what local runs and tests wire in, so the code is exercised constantly rather than first thing during an outage. Add one alert on "running degraded" that fires after a few minutes, so an instance that comes up without an optional collaborator and stays that way is noticed rather than normalised.
  • When is starting degraded actively the wrong call?
    When the missing collaborator is on the write path, or when the degradation cannot be made visible. A service that comes up without the store it must write to can only return errors, and every minute it stays up is a minute the deploy system thinks the rollout succeeded. Crash-looping is the more honest signal, and the one the platform already knows how to act on.
  • Who can overrule this decision, and on what evidence?
    On-call, after an incident where a degraded start served bad responses for hours; or the service's stakeholders, after an outage where a non-essential dependency took the fleet down. Both are legitimate. The durable answer is a written per-collaborator policy next to the wiring, with the reason and the alert threshold, so the next incident amends a note rather than re-litigating the principle.

saying these in an interview costs you the question

  • Treats every dependency as fatal, so a cache blip crash-loops the fleet
  • Starts degraded silently while readiness still claims full health
  • Adds a nil check at every call site instead of a stand-in of the same type
  • Retries forever at startup so the deploy stalls instead of failing
  • Decides per incident rather than recording a policy per collaborator
  • Treats a health check that proves the server answers as proof it works