A service can re-read its configuration on a reload signal or by watching the mounted file — how does each fail?
answer
- who decides when the read happens
- external trigger versus self-trigger
- a swap the watch never sees
- a dead watcher looks like no change
- candidate, validate, swap one reference
basics
~20 sReload on a signal fails when nothing sends the signal to every instance, or the handler rebuilds only part of the state. Watching the file fails when the watcher misses a whole-directory swap, storms on repeated events, or dies silently and never updates again.
solid answer
~50 sBoth put a second read into the process; they differ in who decides when. A reload signal is externally triggered, so it is explicit, auditable and easy to reason about — and it fails when nobody arranges to send it to all forty replicas, when the instance that was briefly unreachable never gets it, or when the handler reloads the table but not the derived state the request path actually reads. A file watch is self-triggered, so no external step is needed — and it fails when the platform replaces the whole mounted set rather than editing the file in place, leaving a watch registered on something that no longer exists; when several files change and it re-reads a half-consistent set; or when the watching thread dies and the process keeps serving forever with no error and no updates. The safe shape for either is the same: build a complete candidate, validate it, swap one reference, and report which version is live.
code
pseudocode · 16 lines# start-up: the only read a process gets for free
limits = parse(readAll("/etc/gateway/limits"))
on reloadSignal:
candidate = parse(readAll("/etc/gateway/limits"))
if candidate.isValid:
limits = candidate # one reference, swapped whole
report(configVersion = checksum(candidate.rawBytes))
log("reload adopted")
else:
log("reload REJECTED, still serving previous table")
watch directory "/etc/gateway/limits":
on change:
debounce(shortInterval) # one edit can fire many events
raise reloadSignal # same path, so both triggers agreego deeper
Know that a running process re-reads configuration only if it was built to, and that the two common triggers are an external reload signal and the process watching the file it was given.
Explain the trade-off by who decides when the read happens, and name a concrete failure for each: a signal nobody sends to every instance, and a watch that misses a whole-set swap or dies silently.
Show the safe reload shape — candidate, validate, swap one reference, keep the old set on failure, report the adopted version — and explain why an unreported reload is unauditable across forty replicas.
Weigh whether reload support is worth owning at all. It is code your team maintains and tests forever, against the cost of replacing instances; for workloads with long-lived connections the reload path usually earns its keep.
## Two triggers for the same second read A process that never re-reads its configuration is stuck with the copy it took at start-up. Both mechanisms here add a second read; they differ only in **who decides when that read happens**. - **Re-read on a signal** — something outside the process tells it to reload. The decision is external, explicit and observable: someone or something did an act you can point at. - **Watching the mounted file** — the process asks the operating system to tell it when the content under a path changes, and re-reads when told. The decision is internal and automatic; nothing outside the process needs to know a reload happened. A third option exists and is often the right one: replace the instance so a new process reads at start-up. That is not a reload at all, and it costs the warm state and the long-lived connections a chat gateway is full of. ## How the signal path fails - **Nobody sends it.** The mechanism works perfectly and is never used, because delivering the signal to all forty replicas is itself an operational exercise. The value changes, no signal goes out, and the fleet is unchanged. - **Some instances miss it.** An instance that was restarting, unreachable or newly created at the wrong moment never receives it, and the fleet ends up on two versions. - **The handler is partial.** It reloads the parsed table but not the buckets, matchers or clients derived from it, so the visible behaviour is unchanged or, worse, half-changed. - **The handler is slow or unsafe.** A reload that rebuilds state on the thread that answers requests adds a latency spike to the moment of every change, and a reload that mutates the live table in place can let one request see a half-updated set. - **The signal does not reach the process you think.** If the container starts a wrapper that does not pass signals through to the real process, the reload never arrives — the same class of failure that bites shutdown handling. ## How the watch path fails - **The whole set is swapped, not edited.** Platforms commonly build the new value set in a fresh directory and then swap the pointer in one step. A watch registered on the individual file is then watching an object that has been replaced, and it reports nothing forever. Watching the containing directory and re-opening by path survives that; watching one inode does not. - **The watcher dies quietly.** The registration is dropped, the thread throws once and exits, or the notification limit is exhausted. The process keeps serving happily with a configuration that can no longer change, and there is nothing to alert on unless the process itself reports a heartbeat. - **Event storms.** One logical edit can produce several notifications. Without a short debounce, the process re-reads and rebuilds repeatedly for one change. - **Torn multi-file sets.** If the configuration spans several files that are not swapped together, an eager watcher can read a new file beside an old one and build a set that never existed as a whole. - **Nobody knows it happened.** A self-triggered reload leaves no external record. Without a log line and a reported version, the operator cannot tell a replica that reloaded from one that did not. ## The shape that makes either one safe 1. **Read everything into a candidate.** Never mutate the live structure while reading. 2. **Validate the candidate as a whole** — required keys present, values in range, derived structures buildable. 3. **Swap a single reference** that the request path reads, so any given request sees entirely the old set or entirely the new one. 4. **On failure, keep the previous set and say so loudly.** A rejected reload that fails silently is indistinguishable from no reload. 5. **Report the version that is live** — a checksum of the bytes adopted plus the time — in a log line and on a status surface, so a fleet can be audited. | | Reload on a signal | Watching the mounted file | |---|---|---| | Who decides | Something outside the process | The process itself | | Fleet-wide reach | Must be arranged for every instance | Happens on each instance independently | | Silent-failure risk | Low: the send either happened or not | High: a dead watcher looks identical to no change | | Typical miss | An instance that was unreachable | A pointer swap the watch never saw | | Audit trail | Natural — the send is an event | Only if the process reports it | ## Choosing between them Neither is strictly better. A watch gives the shortest path from edit to effect with no coordination, which is what you want for a table edited during business hours. A signal gives an explicit, ordered, auditable moment, which is what you want for a change that must be correlated with something else. Many services support both, and the honest answer in an interview is that the trigger is the easy half: the hard half is that the reload is complete, atomic in memory, and visible from outside.
- Why does watching the individual file miss a change that watching its directory catches?Because the platform usually does not edit the file in place. It writes the new set somewhere else and then swaps the directory's pointer in one step, so the old object the watch was registered on still exists, unchanged, and simply is not what the path resolves to any more. A watch on the containing directory sees the swap, and re-opening by path after every event reads the current content.
- How would you detect that a replica's file watcher has died?Make the process publish the version it has adopted — a checksum of the bytes plus the time it adopted them — and export it. A dead watcher then shows up as a replica whose reported version stops moving while its peers advance. Without that, the failure is invisible: the process serves correctly and is simply frozen on an old table forever.
saying these in an interview costs you the question
- Believes a reload path exists because the configuration library supports one
- Treats a file watch as instant and guaranteed to fire
- Mutates the live configuration structure while re-reading it
- Keeps serving after a failed reload without logging the rejection
- Reloads the parsed values but not the structures derived from them
- Assumes a signal sent to one instance reaches the whole fleet