The platform's name-resolution and instance metadata services are degraded, yet the databases behind them are healthy — why does your service still fail?
answer
- the lookup layer, not the service
- healthy target, unreachable by name
- identity and placement come from the same hop
- bootstrap has nothing cached
- shared by every tenant, isolated for none
basics
~20 sHealth and reachability are separate facts. Name resolution turns a dependency's name into an address and the metadata service hands a workload its identity and placement, so both sit in front of everything — a healthy database nobody can find or authenticate to is unreachable, and neither hop is usually drawn.
solid answer
~50 sBoth are **lookup services**: thin, provider-operated, shared by every tenant, and sitting in front of the things you actually care about. Name resolution converts a dependency's name into an address, and the instance metadata service is where a workload gets its own identity and placement facts. Neither stores your data and neither appears on most architecture diagrams, which is precisely why their failure is disorienting — the database is up, every check against it passes, and the caller still cannot get there. The blast radius depends on where in the lifecycle the lookup happens. A long-running process that resolved and fetched once at startup is holding what it needs and keeps serving. A process that has just started, or that performs the lookup on every request, has nothing to fall back on and fails immediately.
go deeper
Know that before a service can talk to a database it has to turn a name into an address, and before it can prove who it is it asks the platform. Those two steps can fail on their own while the database is perfectly fine.
Explain who consults which lookup and when: name resolution on connect, identity and placement at bootstrap and renewal. That timing is what decides which processes fail first and which keep serving.
Demonstrate the diagnosis — separate the target's health from the caller's reachability out loud — and the mitigations: hold looked-up answers for a bounded window, keep the lookup off the request path, and avoid replacing processes during the window.
Own the inventory. Require that every design review names the shared platform lookups it depends on and states what happens when each is unavailable, and decide which few dependencies justify a static fallback and the staleness risk it carries.
## The hop that is rarely drawn An architecture diagram shows the service, the queue, the cache and the database. It almost never shows the two hops that happen before any of those can be used: turning a name into an address, and a workload discovering who it is. Both are provided by the platform, both are shared by every tenant on it, and neither holds any of your data — which is why they are easy to forget and why their failure produces the most confusing kind of incident report. The target is healthy by every measure. The caller still cannot reach it. This is the direction that matters and the one teams get wrong under pressure: **a healthy target is not a reachable target**. Reporting 'the database is down' when the lookup layer is degraded sends responders to examine a system that will look perfect to everyone who opens it, and costs the first half hour of the incident. ## Two lookups, two different blast radii They fail differently because they are consulted at different moments. | | Name resolution | Instance metadata service | |---|---|---| | What it answers | Where is the thing called X | Who am I, and where am I running | | Typically consulted | On connect, and again when a connection is re-established | During bootstrap, and whenever a workload identity is renewed | | Held afterwards | For a bounded period, then looked up again | For the life of what was fetched | | Who fails first | Anything opening a new connection | Anything starting up | | Who survives longest | A process with established connections | A process holding an unexpired workload identity | The common thread is that each is a **shared, correlated dependency**: the provider runs one of them for everybody, so a degradation is simultaneous across every tenant, and the per-tenant isolation that applies to your compute and your storage does not apply here. There is no version of this that only affects you, and there is no version that only affects one of your services. ## What actually breaks, in order 1. **Processes starting up.** Bootstrap is where the dependency is unavoidable. A process that has not yet resolved its dependencies or fetched its identity has nothing cached, so it never reaches a serving state at all. This is why replacing processes during such an incident converts a degradation into an outage. 2. **Processes opening new connections.** An established connection already has an address and does not consult the lookup again; the next connection does. A service under load that is scaling its pools feels this quickly. 3. **Processes renewing a workload identity.** The renewal path goes back through the metadata service, and failing renewal means an eventual expiry. 4. **Processes serving from memory or from open connections.** These are the survivors, and naming them in advance is what lets you offer any product at all during the window. ## Reducing the exposure - **Get the lookup off the hot path.** Resolve and fetch once, hold the result for a bounded period, and re-attempt in the background rather than on the request that needs it. - **Keep the held value usable past its normal refresh point.** A value that is slightly stale is worth much more than a failed lookup, provided you bound how long you will trust it and say so explicitly. - **Have a static fallback for the few dependencies that justify one.** A configured address for a critical dependency is ugly and it goes wrong silently when the address changes, so this is a deliberate exception for one or two things, not a policy. - **Fail the right way.** A lookup failure should surface as 'cannot reach dependency', not as 'dependency returned an error', so the incident starts in the right place. ## Finding these before an incident does The honest inventory is not built from the diagram, because the diagram is what is missing them. Build it from behaviour: list everything a process contacts **before it serves its first request**, and everything the platform client library does that no application code mentions. Ask, for each dependency, whether it is consulted once at bootstrap or on every call, and what the process does when the answer never comes. The set of shared platform lookups is small, which is the good news — and once written down it stays written down, whereas relearning it live costs the opening of every incident.
- Which of these lookups happen once at startup and which happen repeatedly?A workload's identity and placement facts are typically fetched during bootstrap and then renewed on a schedule, so they are held for a long time. Name resolution is consulted whenever a connection is established, which for a service with pooled connections is far less often than once per request but far more often than once per process. That difference is exactly why a restarting process fails first and a busy process with warm pools fails last.
- How would you build an inventory of dependencies like these before an incident forces one?Take it from behaviour rather than from the diagram. List everything a process contacts before it serves its first request, and everything the platform client library does that no application code names. For each, record whether it is consulted at bootstrap or per call, how long the held answer is trusted, and what the process does when the answer never arrives. That last column is the one nobody has filled in.
- Why is caching the answer forever the wrong fix?Because the lookup exists to reflect change: an address that moved, an identity that was rotated, a placement fact that is now wrong. An unbounded cache converts a rare correlated outage into a permanent source of silent, hard-to-diagnose staleness. Bound it instead — hold the answer past its normal refresh point for a defined window, log loudly while you are doing it, and fail once the window closes.
saying these in an interview costs you the question
- Reports the database as down when only the lookup layer is degraded
- Assumes a healthy target is automatically a reachable target
- Thinks per-tenant isolation applies to the platform's own shared lookup services
- Treats the metadata hop as free because it never appears in a latency budget
- Caches the looked-up answer forever to avoid ever depending on the service again