skip to content

Your ACME automation renews several names in parallel and intermittently fails with badNonce — what is happening and what should the client do?

level: seniorimportance: nice to knowfreq 28%

answer

  1. one request, one server-issued value
  2. replay defence, not rate limiting
  3. concurrency exposes it, single runs do not
  4. the error hands back a fresh one
  5. bounded retry of the same request

basics

~20 s

Each signed ACME request consumes a one-time nonce, so parallel workers drawing from one pool race and lose. The badNonce error response carries a fresh Replay-Nonce, and the client should retry the same request with it rather than failing the run.

solid answer

~40 s

The anti-replay `nonce` in an ACME protected header is issued by the server — in a `Replay-Nonce` response header, or from `newNonce` — and accepted once. Two requests signed with the same value cannot both succeed, so concurrent renewals that share a nonce pool produce `urn:ietf:params:acme:error:badNonce` under load and never in a single-threaded test. The specification requires the server to include a fresh `Replay-Nonce` on a badNonce error, and says the client should retry using it. So the right behaviour is a bounded retry of that same signed request with the new nonce, and harvesting `Replay-Nonce` from every response instead of fetching one per request. A badNonce retry is expected protocol traffic, not an incident.

code

http · 9 lines
http
HTTP/1.1 400 Bad Request
Content-Type: application/problem+json
Replay-Nonce: IXVHDyxIRGcTE0VSblhPzw

{
  "type": "urn:ietf:params:acme:error:badNonce",
  "detail": "JWS has an invalid anti-replay nonce",
  "status": 400
}

go deeper

for a junior

Recall that each signed request carries a one-time value from the server and that reusing it is rejected. You are not expected to have hit this yet.

for a middle

Explain what the anti-replay value defends against and where a client gets one, including the response header that lets requests chain without an extra round trip.

for a senior

Diagnose it from the symptom shape: intermittent, concurrency-scaled, invisible in a single-threaded rerun — then fix it with a nonce pool and a bounded retry rather than by serialising everything.

for a principal

Set the expectation that retryable protocol conditions are absorbed by clients and never paged, and that the alert sits on the renewal outcome with runway to spare.

The nonce is the part of ACME that works perfectly in a test run of one certificate and starts failing the week the estate reaches fifty. Understanding it is understanding why. ## What the nonce is for A signed ACME request is not confidential. Anyone who can see it has a complete, validly signed object; if the server accepted it twice, a captured request could be replayed — a revocation, an order, a challenge-ready POST. So each request's protected header carries a `nonce` that the server issued and will accept **once**. Together with the `url` field, this pins a signature to one request, at one resource, at one time. Nonces arrive two ways: - from the `newNonce` resource, which a client can query directly; and - in the `Replay-Nonce` response header that the server attaches to responses, so a client can chain one request to the next without an extra round trip. ## Why parallelism breaks it A nonce is consumed by the request that uses it, not by the worker that took it. Renewing many names at once means several signed requests in flight, and the moment two of them are signed with the same value, one is rejected. Symptoms line up with this exactly: - the failure is **intermittent** and scales with concurrency; - it never reproduces in a single-threaded rerun; - the rejected request is well formed, so nothing else in the log looks wrong; - retrying "the whole renewal" usually succeeds, which hides the cause. A stale nonce fails the same way: a value harvested, then held while a long DNS propagation wait ran, may no longer be accepted. ## The contract on the error On a `badNonce` error the server **must** include a fresh `Replay-Nonce` header that it will accept on a retry, and the client **should** retry the request using it. Those are two different strengths of requirement and it is worth keeping them apart: the server is obliged to make the retry possible; the client is advised to take it. The error document is a problem object with the type `urn:ietf:params:acme:error:badNonce`. The retry is of the **same** request, re-signed. The payload does not change — only the nonce in the protected header, and therefore the signature. ## What a client should do 1. **Keep a nonce pool** and add the `Replay-Nonce` of every response to it, including error responses. Fetch from `newNonce` only when the pool is empty. 2. **Take a nonce per request, never per batch**, and never reuse one after a response has come back. 3. **Retry on badNonce with the nonce from that error response**, bounded — three to five attempts, not an unbounded loop, so a server that rejects every nonce surfaces as a failure instead of a spin. 4. **Do not alert on a badNonce retry.** It is expected traffic. Alert on the renewal not completing, which is the outcome that matters. 5. **Consider serialising per account** if the client cannot manage a pool safely; the throughput cost is trivial next to a renewal that silently stops. ## The operational shape | Symptom | Likely cause | Action | |---|---|---| | badNonce under concurrency only | one nonce used by two in-flight requests | pool nonces, one per request | | badNonce after a long wait step | the held value went stale | take the nonce immediately before signing | | badNonce on every attempt | clock or state problem at the server side | stop after the bounded retries and report | For a chain of pop-up sites renewing dozens of names on one nightly job, the dangerous version of this bug is not the noisy one. It is the job that treats a badNonce as a hard failure for that name, logs it, and moves on — so a handful of names quietly miss their renewal window every night while the job's exit status stays green.

  • Why is a nonce needed when the request is already signed?
    Because a signature proves who produced the bytes, not that they are new. A signed request is not confidential, so an observer holding a copy could submit it again — reordering an order, repeating a revocation. The one-time nonce makes each signed object valid exactly once, and the `url` field in the same header stops it being aimed at a different resource.
  • Should a badNonce retry be reported as a failure in renewal metrics?
    No. It is ordinary protocol traffic that a correct client absorbs, and alerting on it trains operators to ignore the channel. Count it if you want a concurrency signal, but page on the outcome — a name whose renewal has not completed while its remaining validity is shrinking. A client that treats badNonce as fatal turns a retryable condition into a missed renewal.

saying these in an interview costs you the question

  • Thinks badNonce means the account key is wrong
  • Fetches a new nonce and retries a freshly built request
  • Treats the nonce as a rate-limiting or throttling device
  • Retries badNonce in an unbounded loop
  • Says the nonce can be reused within one session
  • Alerts on every badNonce as an issuance incident