An end-entity certificate under ACME automation still expired at 02:00 on a Saturday — what must the renewal schedule get right?
answer
- no renewal verb, only a new order
- anchor to the certificate, not the job
- start with a third of life left
- check what is served, not the exit code
- page on runway, never on expiry
basics
~20 sA renewal schedule must run against the served certificate's own notAfter with weeks of headroom, retry with jitter and backoff, and alert on renewal not having succeeded — while there is still runway — rather than on expiry itself.
solid answer
~40 sACME has no renewal operation: renewing is placing a new order for the same identifiers, so everything that could fail at first issuance can fail again. The schedule should be anchored to the certificate's `notAfter`, starting when roughly a third of the lifetime remains, which buys many attempts before anything is user-visible. Spread the start times with jitter so an estate does not converge on one hour and collect `rateLimited` errors. Most importantly, monitor the certificate **being served on the wire**, not the job's exit code: issuance and installation are separate steps, and a renewal that fetched a chain but never reloaded the terminator leaves the old certificate in place. Then page on "has not renewed for N days", which fires on a weekday.
code
pseudocode · 13 linesfor each certificate in served_certificates:
lifetime = certificate.notAfter - certificate.notBefore
renew_from = certificate.notAfter - (lifetime / 3)
if now < renew_from + jitter(certificate.name):
continue
order = place_order(certificate.identifiers)
if order.status is "valid":
install(order.chain)
reload_terminator()
else:
record_failure(certificate.name, order.error)
if certificate.notAfter - now < runway_threshold:
page_operator(certificate.name)go deeper
Recall that certificates are renewed on a schedule, well before they expire, and that renewing means getting a new certificate rather than extending the old one.
Explain why the schedule is anchored to the certificate's notAfter with a fraction of lifetime as headroom, and why jitter and backoff belong in it.
Show that you monitor the certificate served on the wire rather than the job, and that the page fires on remaining runway so the incident lands on a weekday.
Own the estate-level bet: how much runway is enough, whether renewal is centralised or per-service, and how a name's retirement is driven through the same lifecycle as its creation.
Expiry outages happen at organisations that automated issuance years ago. The protocol did its job; the timer around it did not. Automating issuance removes the manual step and replaces it with a scheduled job that has its own failure modes, and the schedule is the actual availability control. ## There is no renewal, only a new order ACME has no verb that extends a certificate. A renewal is a fresh order for the same identifiers, which means every dependency of first issuance is a dependency of every renewal: - the challenge must still be answerable — the same port open, the same path reachable, the same zone writable; - the account must still be in good standing with the authority; - the automation must still hold its account key; - the new chain must still be installable where the old one lives. A server may reuse a recent still-valid authorization and skip the challenge, but that is a server-side optimisation a client cannot count on. Design for the challenge running every time. ## Anchor the schedule to notAfter Schedule against the certificate's own `notAfter`, read from the certificate, not against the date the job last ran: 1. Compute the lifetime as `notAfter - notBefore`. 2. Begin attempting renewal when about **one third of the lifetime remains**. 3. Retry on a modest interval through that window rather than once. 4. Add per-name jitter so a hundred names do not all wake at the same minute. The headroom is the whole point. Starting weeks out converts "the renewal failed" from an outage into a ticket: there is room for repeated attempts, for a DNS change to propagate, for an authority incident to end, and for a person to look at it on a weekday. ## Rate limits are part of the design Authorities limit issuance per account and per name, and exceeding them surfaces as an error in the `urn:ietf:params:acme:error:` namespace — `rateLimited` among them. Two habits cause it: retrying a hard failure in a tight loop, and renewing an entire estate on the same cron minute. Back off on repeated failure, honour any wait the response asks for, and spread the schedule. ## Issuance is not deployment This is the failure behind most "but the renewal job is green" incidents. | What the job did | What the client sees | |---|---| | ordered and fetched a new chain | nothing, until it is installed | | installed it but did not reload the terminator | the old certificate, still being served | | reloaded one node of several | some connections fine, some expired | | renewed a name nobody serves any more | a permanently failing job and no impact | The check that catches all four is the same: connect to the service and read the `notAfter` of the certificate actually presented. A renewal job's exit code is a statement about the job; the served certificate is a statement about the outage. ## Alert on runway, not on expiry - **Page** when the served certificate has less than a set runway left — comfortably more than one renewal window. - **Ticket** when a renewal has not succeeded for several days while runway remains. - **Never** rely on the expiry alert as the first signal. By then there is no time left, and the hour is not yours to choose. ## Names that go away For a chain of pop-up stores this last point is not academic. A storefront that closes leaves a name that no longer resolves, and a renewal for it fails forever — noisily enough to train everyone to ignore the channel, which is how a real failure gets missed. Retiring a name is part of the lifecycle: drop it from the order's identifier set, and where the certificate should not remain usable, submit it to the `revokeCert` resource, which is the protocol's automated way to ask the issuer to revoke. The certificate's shrinking lifetime does the rest. ## Why short lifetimes are the design, not a burden Shorter certificates mean the renewal path is exercised constantly instead of annually, so a broken path is discovered in days rather than on the one night it matters. The corollary is that the timer, not the certificate, is the thing to monitor: a renewal mechanism that runs weekly and is watched is a stronger availability guarantee than a long-lived certificate that nobody thinks about until it is gone.
- Why is renewal in ACME a new order rather than an update to the existing one?Because the protocol issues certificates rather than maintaining them. An order is a record of one issuance decision, and its authorizations record control proven at that moment. Extending it would mean re-using a control proof that may no longer hold. A new order re-checks the identifiers and produces a new certificate with a new serial and a new validity window; the previous one continues until it expires or is revoked.
- A storefront closes and its name is released. What does the automation do with the certificate?Remove the identifier from the renewal set first, so the job stops failing on a name nobody serves. If the certificate should not remain usable — the name may be picked up by someone else, or the key was on a machine now out of your control — submit it to the `revokeCert` resource, signed either by the account that ordered it or by the certificate's own key. That is a request to the issuer, not the mechanism relying parties consult.
- Two hundred names renew nightly and the job starts returning rateLimited. What changes?Spread and back off. Jitter the start time per name across the whole renewal window instead of a single cron minute, so attempts are distributed over days rather than minutes. Back off exponentially on repeated failure for the same name instead of retrying tightly, and stop retrying a name whose failure is not transient. Because the window opens with a third of the lifetime left, spreading costs nothing in safety.
saying these in an interview costs you the question
- Renews a fixed number of days after the last successful run
- Treats the renewal job's exit code as proof of the served certificate
- Believes ACME extends the existing certificate in place
- Schedules every name in the estate on one cron minute
- Relies on the expiry alert as the first warning
- Keeps retrying names for storefronts that no longer exist