skip to content

After a shared-secret rotation, some TACACS+ devices stop authenticating — how does a key mismatch present on the wire?

level: seniorimportance: should knowfreq 35%

answer

  1. no wrong-key status exists
  2. noise body, lengths do not add up
  3. discard the packet, signal ERROR
  4. ERROR means look elsewhere, not refused
  5. per-client keys bound the blast radius

basics

~20 s

A wrong shared secret yields a wrong pseudo_pad, so the body unmasks to noise. The receiver sums the component lengths inside it, finds they do not equal the header's length field, discards the packet and signals ERROR. No status means wrong key.

solid answer

~50 s

Nothing in TACACS+ verifies the key before it is used. Both peers derive a `pseudo_pad` and XOR; if the secrets differ, the pad differs and the body unmasks to noise. The only thing that catches it is the reconciliation RFC 8907 Section 4.5 requires: sum the component lengths declared inside the unmasked body and compare with the `length` in the cleartext header. On a mismatch the packet **MUST** be discarded and an ERROR signalled — the ERROR status of whichever exchange the packet belongs to, since authentication, authorization and accounting each have their own enumeration. That matters diagnostically, because ERROR means processing did not complete, so the device must behave as if the server were unreachable rather than telling anyone their credentials were refused. The symptom is a whole device failing over, not one user being rejected.

code

pseudocode · 13 lines
pseudocode
function receive(packet):
    pad  = pseudo_pad(packet.header.session_id, configured_key,
                      packet.header.version, packet.header.seq_no,
                      packet.header.length)
    body = mask(packet.body, pad)

    declared = sum of every component length field inside body
    if declared != packet.header.length:
        discard packet
        signal ERROR for this exchange   # no status means "wrong key"
        return

    process(body)

go deeper

for a junior

Remember that both sides derive the mask from a secret neither ever sends, so a mismatch cannot be reported as such — it simply produces unusable bytes.

for a middle

Describe the reconciliation: sum the component lengths inside the unmasked body, compare with the header length, then discard and signal ERROR on a mismatch.

for a senior

Use the ERROR-versus-FAIL distinction diagnostically — every user on one client failing identically is a key problem, one user failing everywhere is not — and keep per-client keys so a rotation's scope stays legible.

for a principal

Decide how key lifetimes are tracked and rotated across an estate where some devices are reachable only in a maintenance window, and what the specification obliges your tooling to support.

## There is no "wrong key" message TACACS+ has no handshake in which two peers demonstrate that they hold the same shared secret. The TCP connection opens, a packet arrives, and the receiver derives a `pseudo_pad` from its own configured key together with the `session_id`, version and `seq_no` it can read in the cleartext header. If the keys differ, the pad differs, and XORing it into the body produces bytes with no structure. So the failure is detected **after the fact, by structure**, and only by structure. There is no key identifier in the 12-byte header, no digest to compare, and no value in any of the three status enumerations that means "your key is wrong". ## The one check that fires 1. Unmask the body with the locally derived pad. 2. Read the component length fields that the body's own layout declares. 3. Sum them. 4. Compare that sum with the `length` field in the cleartext header. 5. If they differ, **discard the packet and signal an ERROR** — the ERROR status belonging to the exchange in question, since the authentication, authorization and accounting enumerations are three different sets of values. With a wrong key, step 2 is reading noise. The length fields become arbitrary numbers, so step 4 essentially always fails. The check is not a cryptographic verification and does not claim to be, but as a key-mismatch detector it is reliable enough that this is how a mismatch presents in practice. ## ERROR is not FAIL, and the difference drives the symptom - **FAIL**: processing completed and the answer is no. Someone typed the wrong password. - **ERROR**: processing did not complete, the result cannot be applied, and the device **MUST** behave as though the server were unreachable. A key mismatch is the second, which is why the user-visible symptom is not "access denied". It is a device whose exchanges with that server all abort, immediately and identically, regardless of who is logging in. Anyone reading the incident as a credential problem is looking at the wrong layer. ## Telling it apart from the things it resembles | Observation | Wrong shared secret | Wrong password | Server unreachable | |---|---|---|---| | Scope | every user on the affected client | one user, on every client | every user on every client that uses it | | Server-side record | a discarded packet, no decision | a completed decision of FAIL | no packet arrives at all | | Onset | immediate and total after the change | sporadic | immediate and total, fleet-wide | | Persistence | permanent until the key is corrected | cleared by typing it correctly | clears when the path or process returns | The distinguishing signal during a rotation is **which** clients broke. Because a server is required to allow a dedicated key per client, a rotation done client by client breaks exactly the clients whose two sides disagree — which is also why per-client keys make the blast radius legible instead of fleet-wide. ## What the specification asks of key management RFC 8907 Section 10.5.1 mixes strengths deliberately, and repeating them at the right strength is part of a good answer: - Servers **MUST NOT** expose shared secrets in logs. - Servers **MUST** allow a dedicated key per client. - Server management **MUST** provide a way to track key lifetimes and notify administrators. - Servers and clients **MUST** support keys of at least 32 characters. - Servers **MUST** support a minimum-complexity policy. - Administrators **SHOULD** change keys at regular intervals, and **SHOULD** configure at least 16 characters. - Servers **SHOULD** warn when keys are not unique per client. - Clients **SHOULD NOT** allow a server to be configured with no key, or with one under 16 characters. Keys should come from a good random source — the guidance on randomness for security is RFC 4086 — rather than from a memorable phrase, because offline guessing against a capture tests candidates as fast as MD5 computes. ## The redirection footnote One mechanism makes a single compromised secret travel: the deprecated FOLLOW redirection, in which a response tells the client to go and talk to a different server. The reason the specification singles it out as insecure is concrete — the redirection **could carry the shared key for another server to the client**, so breaking one session leverages into others. On an estate where one trackside cable run is far easier to reach than the rest, that is the difference between losing one site and losing the device-administration plane.

  • What does RFC 8907 require of shared-secret length and uniqueness, and at what strength?
    Servers and clients **MUST** support keys of at least 32 characters, servers **MUST** allow a dedicated key per client and **MUST** support a minimum-complexity policy. Administrators **SHOULD** configure at least 16 characters and change keys at intervals, and servers **SHOULD** warn when keys are not unique per client. Mixing those strengths up overstates the document.
  • Why is the deprecated FOLLOW redirection singled out as insecure?
    Because the redirection can hand the client the shared key for another server. That converts one broken session into access to others, so a single compromised link leverages across the estate. A client that honours it widens the blast radius of every other weakness in the obfuscation.
  • Can a receiver tell a wrong key from a corrupted packet?
    Not from the check itself — both present as a length sum that does not reconcile. What separates them is scope and persistence: corruption is sporadic and per-packet, while a key mismatch is total and permanent for that client until one side is corrected. Counters for errors received and connection failures make the pattern visible.

saying these in an interview costs you the question

  • Expects a status code meaning the shared secret is wrong.
  • Says the server returns FAIL so users see a rejected login.
  • Thinks the key is verified in a handshake before any body is sent.
  • Assumes one key for the whole fleet is fine if it is long.
  • Believes a mismatch is visible in the header before unmasking.
  • Treats a memorable passphrase as an adequate shared secret.