skip to content

Why would a Kerberos service reject an AP-REQ whose service ticket is valid, unexpired and for that service?

level: seniorimportance: should knowfreq 44%

answer

  1. valid is not the same as fresh
  2. a ticket is meant to be reused
  3. two checks live outside the ticket
  4. skew window first, then replay cache
  5. KRB_AP_ERR_SKEW and KRB_AP_ERR_REPEAT

basics

~20 s

Two freshness checks sit outside the ticket. If the Authenticator's ctime falls outside the service's allowable clock skew, the answer is KRB_AP_ERR_SKEW; if that client name and timestamp pair is already in the service's Kerberos replay cache, it is KRB_AP_ERR_REPEAT.

solid answer

~40 s

A valid ticket says a KDC issued this credential, for this service, naming this client, within these times. None of that says the `AP-REQ` in front of the service is **fresh**. After decrypting the `Authenticator`, a service checks two more things. First, `ctime` and `cusec` must fall inside its allowable clock skew — RFC 4120 offers five minutes as an example, not as a required constant — and a client whose clock has drifted gets `KRB_AP_ERR_SKEW` while everything else about it is correct. Second, the service keeps a **Kerberos replay cache** of authenticators it has already accepted, keyed on the server name together with the client name and the `ctime`/`cusec` pair; a repeat gets `KRB_AP_ERR_REPEAT`. The cache only has to span the skew window, because anything older fails the clock check first.

code

pseudocode · 26 lines
pseudocode
function accept_ap_req(request):
    ticket_part = decrypt(request.ticket.enc-part, service_long_term_key)
    if decryption failed:
        reject with KRB_AP_ERR_MODIFIED

    session_key = ticket_part.key
    auth = decrypt(request.authenticator, session_key)
    if decryption failed:
        reject with KRB_AP_ERR_MODIFIED

    if auth.cname, auth.crealm differ from ticket_part.cname, ticket_part.crealm:
        reject  // ticket and Authenticator name the same principal or neither is used

    if now outside ticket_part validity interval:
        reject with KRB_AP_ERR_TKT_EXPIRED

    if absolute(now - auth.ctime) > allowable_skew:
        reject with KRB_AP_ERR_SKEW

    if replay_cache contains (server_name, auth.cname, auth.ctime, auth.cusec):
        reject with KRB_AP_ERR_REPEAT
    replay_cache.add(server_name, auth.cname, auth.ctime, auth.cusec)

    if MUTUAL-REQUIRED is set in request.ap-options:
        send AP-REP sealed under session_key
    return accepted

go deeper

for a junior

Remember that a service ticket is reused all day, so it cannot prove a request is recent. The Authenticator's timestamp is what carries freshness.

for a middle

Be able to name both checks and both error codes, and say why the cache is keyed on the client name and the ctime and cusec pair rather than on the ticket.

for a senior

This is the diagnosis question. Separate a drifting client clock from a stale service key from a duplicated request path, and say why widening the window is not a remedy.

for a principal

The judgment is where the replay guarantee lives once a service runs as many instances: a shared cache, or application messages bound by a sub-session key and sequence numbers, each with its own operational cost.

## Validity is not freshness The interesting thing about a service ticket is that it is meant to be used more than once. The KDC issues it with a lifetime measured in hours, and a plant-floor writer that opens a connection to the historian every thirty seconds presents the same ticket every time. So the ticket cannot carry freshness: if it did, it would be spent on first use. Freshness therefore lives entirely in the `Authenticator`, and the service enforces it with two checks that have nothing to do with the ticket's own validity interval. ## Check one: the clock The `Authenticator` carries `ctime` and `cusec`, the moment the client built it. The service compares that against its own clock and rejects anything outside an allowable window with `KRB_AP_ERR_SKEW`. Things worth being precise about here: - **The window is a local policy value.** RFC 4120 uses five minutes as an example when discussing it. Quoting five minutes as a constant the protocol mandates is a common and checkable mistake. - **The comparison is two-sided.** A client whose clock runs fast fails as surely as one that runs slow. - **The failure looks nothing like an authentication failure.** The credentials are correct, the principal is right, the ticket is unexpired, and the exchange still fails — which is why a drifting clock on one writer produces a fault report that reads as an intermittent permission problem. ## Check two: the Kerberos replay cache Within the skew window, an `Authenticator` captured off the wire would still pass the clock check. The service therefore keeps a **replay cache** — distinct from an identifier cache used against bearer credentials elsewhere — recording the authenticators it has accepted. A second presentation of the same one is answered with `KRB_AP_ERR_REPEAT`. What the entry is keyed on matters: - the **server name** the authenticator was presented to; - the **client name** from the Authenticator; - the **`ctime` and `cusec` pair**, whose microsecond component is what separates two authenticators built in the same second. Keying on the ticket instead would break normal operation immediately, because a legitimate client presents the same ticket all day. Keying on the session key would do the same. **How long entries must be kept** follows from the first check: only as long as the skew window. An authenticator older than the window is rejected by the clock check before the cache is ever consulted, so retaining it buys nothing. That bounds the cache by the window rather than by the ticket lifetime, which is what makes it affordable. ## The rejections side by side | Error returned | What triggered it | Where to look | |---|---|---| | `KRB_AP_ERR_SKEW` | The Authenticator's `ctime` is outside the service's allowable window | The two clocks, and the time source they follow | | `KRB_AP_ERR_REPEAT` | The same client name and `ctime`/`cusec` was already accepted for this server | A retransmit, a duplicated request path, or a genuine replay | | `KRB_AP_ERR_TKT_EXPIRED` | The ticket's own validity interval has passed | The client's credential renewal, not the exchange | | `KRB_AP_ERR_MODIFIED` | Decryption or integrity checking failed | The wrong key on the service side, or altered bytes | `KRB_AP_ERR_MODIFIED` is the one that most often means something mundane: the service is holding a key that no longer matches the one the KDC used to seal the ticket. ## Diagnosing it on a plant floor 1. Establish whether one writer fails or all of them. One failing writer with a correct credential points at that host's clock; all of them at once points at the service's key or its time source. 2. Read the error code actually returned rather than inferring it — the four above are distinguishable and they mean different things. 3. If it is skew, compare both clocks against the same reference and fix the drifting one. Widening the window is not a fix. 4. If it is repeat, find out whether the same request is being delivered twice: a retrying client that rebuilds its `AP-REQ` will get a new `ctime`, while one that re-sends identical bytes will not. ## Where the specification relaxes the cache The replay-cache obligation is scoped. It is stated for the case where the application has no other protection on the messages that follow; where the exchange establishes a sub-session key and uses sequence numbers to order and bind subsequent messages, that machinery carries part of the burden. This matters for a service that runs as several instances behind one service principal: a cache held per process or per host only guarantees uniqueness within itself, so the same authenticator can be accepted once per instance unless the instances share a cache or the application protects its own messages.

  • How long must a Kerberos replay cache retain an entry, and why not longer?
    Only as long as the allowable clock-skew window. An Authenticator older than the window already fails the skew check before the cache is consulted, so keeping the entry adds nothing. That is what bounds the cache by the window rather than by the ticket's lifetime.
  • A service runs as several instances behind one service principal — what happens to the replay guarantee?
    It holds only within whichever cache sees the request. With a per-process or per-host cache, one captured Authenticator can be accepted once by each instance. Closing that needs a cache the instances share, or application messages protected by a sub-session key and sequence numbers so that a bare replay of the first message achieves nothing.
  • Does widening the allowable skew to an hour make the exchange more tolerant or less safe?
    Both, and it is the same change. The window is exactly how long a captured `AP-REQ` remains usable to anything the replay cache does not see — another instance, or the same one after a restart that lost its cache. Correct the clocks rather than widening the window.

saying these in an interview costs you the question

  • Says a service rejects only expired or wrong-service Kerberos tickets.
  • Treats five minutes of clock skew as a value the specification mandates.
  • Thinks the Kerberos replay cache must hold entries for the ticket's whole lifetime.
  • Widens the skew window to stop skew errors and calls that a fix.
  • Believes the service asks the KDC whether an AP-REQ was already used.
  • Confuses the Kerberos replay cache with a bearer-token identifier cache.