skip to content

A message consumer that read its broker password six weeks ago still runs a day after that password was replaced — why?

level: seniorimportance: should knowfreq 47%

answer

  1. checked once, at the door
  2. not re-checked per message
  3. the process still holds the old copy
  4. the next reconnect is the deadline
  5. force the reconnect, then withdraw

basics

~20 s

The password was checked when the connection was opened, and an established channel is not re-checked per message. The consumer still holds the six-week-old copy in memory, so it fails only at its next reconnect — which may be days after the change.

solid answer

~40 s

Connection-oriented systems normally authenticate at **connection establishment**, not per message. The consumer's channel was accepted six weeks ago and has carried that decision ever since, so replacing the password yesterday touched neither the channel nor the copy the process still holds. Nothing has withdrawn the old value either, so it remains a valid credential. The deadline is the next **reconnect** — a broker failover, a network interruption, the pool retiring an aged connection, a scale-out, an unrelated deploy. Then the process presents the value it resolved six weeks ago and is refused, at a time that looks unconnected to the change. "It is still running" is therefore not evidence that the consumer moved; only a successful reconnect after the replacement is.

go deeper

for a junior

Recall that a password is checked when a connection is opened, not on every message, so an already-open connection keeps working after the value is replaced.

for a middle

Explain both halves: the established channel is not re-checked, and the process still holds the copy it read at start-up. Name the next reconnect as the moment it breaks.

for a senior

Demonstrate the operating habit — force the reconnect and confirm the new value is accepted before withdrawing the old one, rather than treating "still running" as success and meeting the failure at 03:00.

for a principal

Own the estate-level version: how long two values must both be accepted, who decides that window, and how a fleet of long-lived consumers is cycled so replacements never depend on an unplanned network event.

## What was checked, and when The consumer presented its password once, when it opened the connection. The broker checked it then and accepted the channel. From that moment the channel carries the decision: messages flow over an already-authenticated connection and the credential is not presented again. Designs differ in the detail — some re-check on a periodic renewal, which shortens the window without changing the mechanism — but the common shape is a check at establishment and silence afterwards. So on the day the password was replaced, three things were true at once, and all three are ordinary: - the broker had already accepted this channel, six weeks earlier; - the consumer process still held, in memory, the copy of the password it resolved at start-up; - the old password had been **replaced**, not **withdrawn** — a new value now exists, and the old one is still accepted by the broker. Any one of those alone would explain a day of continued delivery. Together they explain six weeks of it. ## The reconnect is the real deadline The credential is presented again only when a new connection is opened. So the failure is scheduled by whatever opens one next: - the broker restarts, fails over, or drops idle clients; - a network interruption closes the channel and the client reconnects; - the pool retires a connection on age or idle time and opens a replacement; - the consumer scales out, and the new instance resolves the current value while the old instance keeps the earlier one — a fleet split across two values; - someone deploys something unrelated and the process restarts. None of those is scheduled by you, which is why the symptom arrives at 03:00 several days later, with no change deployed that day, and does not look related to a password replacement that has long since scrolled off the change log. ## What "still running" proves | observation | what it actually proves | |---|---| | the consumer is still delivering messages | a connection opened before the replacement is still accepted | | the store shows the new value as current | the write landed in the store; nothing about any consumer | | the configuration change was deployed successfully | the new value was delivered to the host; not that the process re-applied it | | the consumer reconnected after the change and came back up | the value it presented at that reconnect is accepted — this is the real proof | Only the last row is evidence. The credential is an unobservable copy inside the process until the process is made to present it. ## Doing the cutover deliberately 1. Put the new value in place, leaving the old one accepted. 2. Make the consumer re-resolve and then **recycle its connections** — or restart it outright, which does both at once. Re-resolving alone leaves the pool on the old value. 3. Force or observe a reconnect and confirm it comes back up authenticated. This is the step people skip, and it is the only one that turns a belief into an observation. 4. Withdraw the old value last, once the holder is demonstrably on the new one. ## The order, and when to invert it Step 4 comes last for a routine replacement because withdrawing first breaks every holder that has not moved. Under a live exposure the trade flips: if the old password is in someone else's hands, you withdraw first and accept the outage, because a consumer failing is cheaper than an attacker succeeding. Note that withdrawing is not guaranteed to be instant either — designs differ in whether withdrawing a credential drops sessions already established or only refuses the next one, so under a real exposure you should assume you may also have to end the existing connections explicitly. ## The sentence to leave the interviewer with A long-lived connection is a credential check frozen at the moment it was opened. Replacing the value changes what the *next* connection must present; it does nothing to the one that is already up, and it does nothing to the six-week-old copy in the consumer's memory. The gap between the change and the symptom is the whole problem, and closing it means forcing the reconnect yourself instead of waiting for the network to schedule it.

  • The consumer failed at 03:00 four days after the replacement, with nothing deployed that day. What happened?
    It reconnected. Something closed the channel authenticated six weeks earlier — a broker failover, a network interruption, an idle connection being retired — and the process presented the password it still held, which was refused. The change that caused it landed four days before the symptom, which is precisely why the incident does not look related to it.
  • Would revoking the old password rather than replacing it have made this visible sooner?
    Sooner, yes; safely, not necessarily. Designs differ in whether withdrawing a credential drops sessions already established, so the failure may still wait for a reconnect; where it does drop them, you take the outage at once. Revoke-first is the right order under a live exposure, not for a routine replacement.
  • Two instances of the consumer are running and only one has the new password. How did that happen?
    One instance was replaced or scaled out after the change and resolved the current value at start-up; the other has been up since before it and still holds the old copy. The fleet is split across two values, both accepted, and it stays split until the older instance reconnects or is cycled.

A door checked your pass when you arrived this morning. Changing the lock at lunchtime does not push you back outside — it decides what happens the next time you try to come in.

saying these in an interview costs you the question

  • Thinks the broker re-checks the password on every message
  • Believes replacing a value in the store reaches running clients
  • Treats continued delivery as proof the consumer moved
  • Says the old password stopped working the moment it was replaced
  • Waits for the symptom instead of forcing a reconnect
  • Assumes withdrawal always drops connections already established