skip to content

You run Caddy with automatic HTTPS in containers and scale it from one instance to six behind a network load balancer. Certificate issuance starts failing against the certificate authority's rate limits. What is going wrong and how do you fix it?

level: seniorimportance: should knowfreq 42%

answer

  1. state lives in storage, not in config
  2. six instances, six ACME clients
  3. containers lose it unless mounted
  4. shared storage means one obtains, all load
  5. or terminate TLS in exactly one place

basics

~20 s

Each Caddy instance has its own private storage, so all six independently order certificates for the same names — and container restarts lose that storage and re-order again. The fix is one shared, persistent storage backend that every instance reads and writes.

solid answer

~60 s

Caddy's certificate management is driven by its storage. By default that is the local file system under the data directory, which in a container is ephemeral unless you mount it. So six replicas means six independent ACME clients ordering certificates for the same hostnames, and every redeploy starts over — that is what burns the CA's per-domain issuance limits. Two fixes, and you pick by topology. If instances can share state, point them all at one storage backend: with shared storage, Caddy coordinates through storage-level locks so exactly one instance obtains a certificate and the others load it, and renewals happen once. The stock binary's storage is the file system, so a shared volume works for a single instance or a small set; a networked backend such as Redis or Consul is a third-party module you compile in with `xcaddy`. If sharing state is not on the table, stop terminating TLS in six places — terminate once at the tier in front and let Caddy serve behind it with automatic HTTPS off.

code

bash · 5 lines
bash
docker run -d --name caddy \
  -p 80:80 -p 443:443 \
  -v caddy_data:/data \
  -v "$PWD/Caddyfile:/etc/caddy/Caddyfile" \
  caddy

go deeper

for a junior

Know that Caddy writes its certificates and account key to a data directory on disk, and that a container without a volume for it loses them on every restart.

for a middle

Explain that certificate management is driven by storage, so instances with separate storage each run their own ACME client, and that the container image expects a volume at /data.

for a senior

Show the diagnosis: compare served certificate serials across replicas, read the logs for repeated issuance, and choose between shared storage with lock-based coordination and terminating TLS once upstream.

for a principal

Own the topology decision — whether every edge replica may hold private keys at all, what a custom binary for a networked storage backend costs your release process, and how you keep issuance volume observable before a limit is hit.

## The mechanism you have to name Caddy's certificate lifecycle is not held in memory and it is not held in the config. It lives in **storage**. Storage is where issued certificates and keys, the ACME account key, OCSP staples and the coordination locks all live. The default implementation is the local file system, rooted at the data directory (`$XDG_DATA_HOME/caddy`, typically `~/.local/share/caddy`, or the service user's equivalent). Everything about scaling automatic HTTPS follows from one sentence: **an instance manages the certificates it can see in its own storage.** Two instances with two private storages are two unrelated ACME clients that happen to be configured with the same hostnames. ## The two failure shapes, which usually arrive together **Ephemeral storage.** The official container image expects a volume at `/data`. Without it, every container start has empty storage, so Caddy orders fresh certificates for every name it serves. A deploy loop, a crashlooping container or an autoscaler that churns instances turns that into repeated issuance for the same names. Public CAs limit certificates per registered domain and duplicate certificates per period; you can exhaust those in an afternoon and then be locked out of issuing for your own domain while your existing certificates keep aging. **Horizontal replicas.** Six instances with six private storages multiply everything by six: six ACME accounts, six orders per name, six renewal loops. It also means clients see different certificates depending on which replica they land on — functionally fine, but it makes any certificate pinning, any log correlation on serial number, and any "is the new cert live yet" check confusing. ## The fix: one storage, many instances Give every instance the same storage and the behaviour becomes correct without any leader election of your own. When Caddy needs a certificate it takes a lock in storage; the instance that wins does the ACME order and writes the result; the others wait, then load what was written. Renewal is the same path, so it happens once per certificate rather than once per instance. Instances also notice certificates that appeared in storage after they started. What backend depends on what you can share: - **A persistent volume** works when the volume can genuinely be shared or when there is one instance. A per-replica volume is not shared storage — it fixes the restart problem and leaves the replica problem untouched. - **A networked backend** (Redis, Consul, an object store) is what a real horizontally scaled deployment uses. These are **third-party storage modules**, and because Caddy is a single statically compiled binary they must be built in with `xcaddy` and then selected with the `storage` global option in the Caddyfile. That is a real cost: your edge tier now ships a custom binary you build and version yourself. ``` { storage file_system { root /data } } ``` ## The other fix: stop issuing in six places Sometimes shared storage is the wrong answer because six proxies each holding the same private key is a topology you do not want. Then terminate TLS once — at the load balancer, the CDN or a dedicated ingress tier — and run Caddy behind it with `auto_https off` (or with plain `http://` site addresses). You have given up the headline feature at that tier, which is a legitimate outcome: Caddy's automatic HTTPS pays off where Caddy is the thing facing the internet. ## How you would confirm the diagnosis Caddy's logs name it directly: repeated "obtaining certificate" for the same hostname across instances or across restarts, followed by CA errors mentioning rate limits. Compare certificate serial numbers served by different replicas; if they differ, storage is not shared. Check whether the data directory is a mount or a container-layer path. And check whether the ACME account key is being regenerated, which is the tell that storage was empty at start. ## The judgment an interviewer is listening for The naive answer is "raise the rate limit" or "use a staging CA". Both dodge the problem. The real answer is that automatic HTTPS makes certificate state an instance-local resource by default, and horizontal scale demands that you decide, explicitly, where that state lives — shared storage, or one termination point. Deciding not to decide is what produces the rate-limit outage.

  • If two instances share storage, what stops both from ordering the same certificate at the same moment?
    Caddy takes a lock in the storage backend before starting an ACME order for a name. The instance that acquires it does the order and writes the certificate; the others block on the lock, then read what was written rather than ordering their own. That is why the coordination requires the storage to be genuinely shared — locks in a per-replica volume coordinate nothing.
  • Would pointing at the ACME staging environment be a reasonable workaround while you are locked out?
    Only as a temporary diagnostic. Staging has far looser limits, so it proves the plumbing works, but staging certificates are not publicly trusted, so you cannot serve production traffic with them. The real remedy is to fix storage so issuance stops repeating, then let the limit window elapse.
  • You cannot share storage for security reasons. What is the next best design?
    Terminate TLS once, in front of the Caddy fleet, and run Caddy with automatic HTTPS off behind it. Certificate state then belongs to one component, and the replicas become stateless proxies. You lose Caddy's headline feature at that tier, which is an acceptable trade when key custody is the constraint.

saying these in an interview costs you the question

  • Blames the certificate authority instead of duplicated issuance
  • Assumes replicas coordinate over the network without shared storage
  • Mounts a separate volume per replica and calls it shared
  • Treats the data directory as a cache that can be discarded
  • Suggests switching to the staging CA as a permanent fix

context