Why does a TACACS+ deployment open a fresh TCP connection for almost every command an engineer types?
answer
- one exchange, one session, one connection
- two sessions per command typed
- reuse must be agreed, not configured
- count connections against logins
- a single error reply means a broken connection
basics
~20 sBecause a TACACS+ session is one exchange and, unless single connection mode was established, a connection carries exactly one session. Each command typed can produce an authorization exchange and an accounting exchange, so each takes its own connection.
solid answer
~40 sThe churn falls out of two rules working together. A session is defined as one authentication sequence, one authorization exchange or one accounting exchange — and authorization and accounting sessions are always a single request and a single reply. Separately, a connection carries one session unless the two ends set `TAC_PLUS_SINGLE_CONNECT_FLAG := 0x04` on the connection's first two packets, in which case the server closes it when the session ends. So a device configured to authorize and record each command generates two sessions per command and, without an established single connection mode, two connections. Across several hundred devices that is a management plane made almost entirely of connection setup, which is what an operator is usually looking at when the server's connection counters climb far faster than its login counters.
go deeper
Remember that TACACS+ splits authentication, authorization and accounting into separate exchanges, and that each exchange is its own session.
Explain the arithmetic: two-packet exchanges, one session per connection by default, and therefore several connections per command unless reuse was negotiated.
Diagnose from a capture: separate healthy churn from a declined single-connection offer and from devices failing at the connection level and retrying. Say which two packets you read to decide.
Judge the trade the design offers — a stateless server and a disposable transport against setup cost at fleet scale — and note that reuse concentrates a device's management plane onto one connection whose failure takes the lot.
## The arithmetic behind the churn On a fabric of several hundred devices, an operator watching the device-administration server sees connections opening and closing far out of proportion to the number of people logged in. Nothing is broken. The count is the product of two separate rules in the protocol. 1. **A session is one exchange.** RFC 8907 defines a session as one authentication sequence, one authorization exchange or one accounting exchange. There is no session object covering a whole login. 2. **Authorization and accounting exchanges are single request/reply pairs.** They are two packets and then they are over. Only an authentication sequence can run long. 3. **A connection carries one session unless reuse was agreed.** Without single connection mode established on the first two packets, the server closes the connection when the session it carried ends. Put them together. A device configured to check each command and to record it produces, per command, one authorization session and one accounting session — and, with no reuse agreed, two TCP connections. An engineer running twenty commands on one device has generated forty connections. Multiply by the fabric. ## What the pattern looks like on the wire All of this is visible without the shared secret, because the header is unprotected: - Connection opens, one packet with `seq_no` 1 and `type` `TAC_PLUS_AUTHOR := 0x02` from the device, one packet with `seq_no` 2 from the server, connection closes. - Immediately afterwards, the same shape with `type` `TAC_PLUS_ACCT := 0x03`. - Each pair carries a different `session_id`. That is healthy churn. Two other shapes are not: - **The offer is being declined.** The device's first packet sets `TAC_PLUS_SINGLE_CONNECT_FLAG := 0x04` and the server's reply does not. The device wants reuse; the server is not agreeing, so every session pays for a new connection. - **The connection is failing before it does work.** A connection opens, a single reply comes back reporting an error, and the connection is finished. A connection-level ERROR — one arising from connection issues such as a mismatched shared secret rather than from the content of an exchange — means no further new sessions may be accepted on that connection, so a device in that state reconnects for every attempt and never gets past the first packet. The counter climbs, and the useful work is zero. Telling those three apart is the diagnosis. The first is the protocol behaving as designed, the second is a negotiation that is not being met, the third is a broken device that is quietly retrying. ## Why the design is like this at all | | One session per connection | Reused connection | |---|---|---| | State at the server between exchanges | none | an open connection per device | | Cost per exchange | a full connection setup and teardown | one round trip | | Blast radius of a transport failure | one exchange | the device's pending sessions on that connection | | Agreed by | nothing — it is the default | the flag on the first two packets | The disposable-transport default is what makes a TACACS+ server simple: between exchanges it holds nothing. That simplicity is paid for in setup cost, and the protocol offers the trade explicitly rather than assuming it. ## What the protocol does and does not let you fix The wire-level answer is the single connection negotiation, and it has to be acceptable to **both** ends: the device offers, the server confirms, and one side alone changes nothing. What sits beyond the protocol — how many commands a device sends for authorization at all, and whether it records each one — is a policy choice on the device rather than a property of the framing, and the framing will faithfully carry whatever that choice produces. ## Answering this in an interview Start from the definition of a session, not from the symptom. "A session is one exchange, and a connection carries one session unless reuse was negotiated" explains the observation in a sentence. Then show the diagnosis: read two packets to see whether the offer was accepted, and look for the short connection that ends in a single error reply, because that is the one that means something is actually wrong.
- The connection count is high but the counts of logins and commands are normal — where do you look first?At the first two packets of a sample of connections. If the device sets the single-connection flag and the server's reply does not, reuse is being declined and every exchange pays for a connection. If instead connections end after one reply reporting an error, the devices are failing at the connection level and retrying, which produces the same counter shape for a completely different reason.
- Why can a device not simply keep the connection open regardless of what the server said?Because the server is the end that closes it. Without the mode established, the server closes the connection at the end of the session, and the device cannot override that. Sending a second session's packets on a connection whose status was never established is exactly what the specification forbids.
- Does a long authentication sequence contribute to this churn?Barely. An authentication sequence may run to many packets, but it is still one session on one connection, so a login costs roughly one connection. The churn is driven by the exchanges that are two packets long and happen once or twice per command typed.
saying these in an interview costs you the question
- Blaming the transport instead of the one-session-per-connection default
- Thinking enabling reuse on the device alone is sufficient
- Assuming a high connection count always means something is failing
- Believing an authentication sequence needs a connection per packet
- Treating a connection-level error as an ordinary refusal of the request