skip to content

An end-entity certificate under ACME automation still expired at 02:00 on a Saturday — what must the renewal schedule get right?

level: seniorimportance: must knowfreq 58%

answer

  1. no renewal verb, only a new order
  2. anchor to the certificate, not the job
  3. start with a third of life left
  4. check what is served, not the exit code
  5. page on runway, never on expiry

basics

~20 s

A renewal schedule must run against the served certificate's own notAfter with weeks of headroom, retry with jitter and backoff, and alert on renewal not having succeeded — while there is still runway — rather than on expiry itself.

solid answer

~40 s

ACME has no renewal operation: renewing is placing a new order for the same identifiers, so everything that could fail at first issuance can fail again. The schedule should be anchored to the certificate's `notAfter`, starting when roughly a third of the lifetime remains, which buys many attempts before anything is user-visible. Spread the start times with jitter so an estate does not converge on one hour and collect `rateLimited` errors. Most importantly, monitor the certificate **being served on the wire**, not the job's exit code: issuance and installation are separate steps, and a renewal that fetched a chain but never reloaded the terminator leaves the old certificate in place. Then page on "has not renewed for N days", which fires on a weekday.

code

pseudocode · 13 lines
pseudocode
for each certificate in served_certificates:
    lifetime = certificate.notAfter - certificate.notBefore
    renew_from = certificate.notAfter - (lifetime / 3)
    if now < renew_from + jitter(certificate.name):
        continue
    order = place_order(certificate.identifiers)
    if order.status is "valid":
        install(order.chain)
        reload_terminator()
    else:
        record_failure(certificate.name, order.error)
        if certificate.notAfter - now < runway_threshold:
            page_operator(certificate.name)

go deeper

for a junior

Recall that certificates are renewed on a schedule, well before they expire, and that renewing means getting a new certificate rather than extending the old one.

for a middle

Explain why the schedule is anchored to the certificate's notAfter with a fraction of lifetime as headroom, and why jitter and backoff belong in it.

for a senior

Show that you monitor the certificate served on the wire rather than the job, and that the page fires on remaining runway so the incident lands on a weekday.

for a principal

Own the estate-level bet: how much runway is enough, whether renewal is centralised or per-service, and how a name's retirement is driven through the same lifecycle as its creation.

Expiry outages happen at organisations that automated issuance years ago. The protocol did its job; the timer around it did not. Automating issuance removes the manual step and replaces it with a scheduled job that has its own failure modes, and the schedule is the actual availability control. ## There is no renewal, only a new order ACME has no verb that extends a certificate. A renewal is a fresh order for the same identifiers, which means every dependency of first issuance is a dependency of every renewal: - the challenge must still be answerable — the same port open, the same path reachable, the same zone writable; - the account must still be in good standing with the authority; - the automation must still hold its account key; - the new chain must still be installable where the old one lives. A server may reuse a recent still-valid authorization and skip the challenge, but that is a server-side optimisation a client cannot count on. Design for the challenge running every time. ## Anchor the schedule to notAfter Schedule against the certificate's own `notAfter`, read from the certificate, not against the date the job last ran: 1. Compute the lifetime as `notAfter - notBefore`. 2. Begin attempting renewal when about **one third of the lifetime remains**. 3. Retry on a modest interval through that window rather than once. 4. Add per-name jitter so a hundred names do not all wake at the same minute. The headroom is the whole point. Starting weeks out converts "the renewal failed" from an outage into a ticket: there is room for repeated attempts, for a DNS change to propagate, for an authority incident to end, and for a person to look at it on a weekday. ## Rate limits are part of the design Authorities limit issuance per account and per name, and exceeding them surfaces as an error in the `urn:ietf:params:acme:error:` namespace — `rateLimited` among them. Two habits cause it: retrying a hard failure in a tight loop, and renewing an entire estate on the same cron minute. Back off on repeated failure, honour any wait the response asks for, and spread the schedule. ## Issuance is not deployment This is the failure behind most "but the renewal job is green" incidents. | What the job did | What the client sees | |---|---| | ordered and fetched a new chain | nothing, until it is installed | | installed it but did not reload the terminator | the old certificate, still being served | | reloaded one node of several | some connections fine, some expired | | renewed a name nobody serves any more | a permanently failing job and no impact | The check that catches all four is the same: connect to the service and read the `notAfter` of the certificate actually presented. A renewal job's exit code is a statement about the job; the served certificate is a statement about the outage. ## Alert on runway, not on expiry - **Page** when the served certificate has less than a set runway left — comfortably more than one renewal window. - **Ticket** when a renewal has not succeeded for several days while runway remains. - **Never** rely on the expiry alert as the first signal. By then there is no time left, and the hour is not yours to choose. ## Names that go away For a chain of pop-up stores this last point is not academic. A storefront that closes leaves a name that no longer resolves, and a renewal for it fails forever — noisily enough to train everyone to ignore the channel, which is how a real failure gets missed. Retiring a name is part of the lifecycle: drop it from the order's identifier set, and where the certificate should not remain usable, submit it to the `revokeCert` resource, which is the protocol's automated way to ask the issuer to revoke. The certificate's shrinking lifetime does the rest. ## Why short lifetimes are the design, not a burden Shorter certificates mean the renewal path is exercised constantly instead of annually, so a broken path is discovered in days rather than on the one night it matters. The corollary is that the timer, not the certificate, is the thing to monitor: a renewal mechanism that runs weekly and is watched is a stronger availability guarantee than a long-lived certificate that nobody thinks about until it is gone.

  • Why is renewal in ACME a new order rather than an update to the existing one?
    Because the protocol issues certificates rather than maintaining them. An order is a record of one issuance decision, and its authorizations record control proven at that moment. Extending it would mean re-using a control proof that may no longer hold. A new order re-checks the identifiers and produces a new certificate with a new serial and a new validity window; the previous one continues until it expires or is revoked.
  • A storefront closes and its name is released. What does the automation do with the certificate?
    Remove the identifier from the renewal set first, so the job stops failing on a name nobody serves. If the certificate should not remain usable — the name may be picked up by someone else, or the key was on a machine now out of your control — submit it to the `revokeCert` resource, signed either by the account that ordered it or by the certificate's own key. That is a request to the issuer, not the mechanism relying parties consult.
  • Two hundred names renew nightly and the job starts returning rateLimited. What changes?
    Spread and back off. Jitter the start time per name across the whole renewal window instead of a single cron minute, so attempts are distributed over days rather than minutes. Back off exponentially on repeated failure for the same name instead of retrying tightly, and stop retrying a name whose failure is not transient. Because the window opens with a third of the lifetime left, spreading costs nothing in safety.

saying these in an interview costs you the question

  • Renews a fixed number of days after the last successful run
  • Treats the renewal job's exit code as proof of the served certificate
  • Believes ACME extends the existing certificate in place
  • Schedules every name in the estate on one cron minute
  • Relies on the expiry alert as the first warning
  • Keeps retrying names for storefronts that no longer exist