skip to content

Caddy's on-demand TLS obtains a certificate during the TLS handshake for a hostname it has never been configured with. What does that make possible, what is the abuse risk, and what control contains it?

level: seniorimportance: nice to knowfreq 32%

answer

  1. certificate obtained during the handshake
  2. the trigger is a stranger's connection
  3. custom customer domains, no reload
  4. an endpoint must approve each name
  5. fail closed, and scope it narrowly

basics

~20 s

On-demand TLS lets customers point their own domains at your service with no config change, because Caddy issues per hostname at handshake time. Unrestricted, anyone aiming DNS at you triggers issuance attempts, so it must be gated by an ask endpoint that approves each name.

solid answer

~50 s

Normal automatic HTTPS manages certificates for names that appear in your config. On-demand TLS inverts that: when a handshake arrives for a name Caddy has no certificate for, it obtains one right then. That is what makes multi-tenant custom domains practical — a customer CNAMEs `shop.theirbrand.com` at you and it just works, with no reload and no config listing thousands of hostnames. The risk is the same inversion: an unrestricted on-demand server will attempt issuance for any name anyone points at your IP, which burns the CA's rate limits, fills storage, and adds handshake-time work an attacker controls. The containment is the `ask` endpoint: before attempting issuance, Caddy calls a URL you host with the candidate domain, and proceeds only if you answer with a success status. Current Caddy versions require that permission check rather than relying on the older rate-limiting knobs. Scope on-demand to the site block that serves tenant domains, never to your own names.

code

bash · 2 lines
bash
# What Caddy asks before issuing for an unrecognised hostname
curl -i "http://127.0.0.1:9000/check-domain?domain=shop.theirbrand.com"

go deeper

for a junior

Know the distinction: ordinary automatic HTTPS handles names listed in the config, while on-demand handles names that show up in a handshake, and the second one needs a gate.

for a middle

Explain the mechanics — the site-level on_demand setting, the global ask endpoint receiving the domain query parameter, and the success status that authorises issuance.

for a senior

Demonstrate the threat model: unauthenticated strangers can trigger issuance, so the real risk is rate-limit exhaustion and handshake-time load, and the approval endpoint must be fast, cached and fail-closed.

for a principal

Own the platform question — whether tenant custom domains should be lazily issued at all versus driven by a control plane that provisions ahead of time, and what limits, monitoring and pruning you need before the feature is customer-visible.

## Two different modes of the same feature Caddy's ordinary automatic HTTPS is **config-driven**: the set of names is known when the config loads, and Caddy obtains certificates for exactly those names in the background. On-demand TLS is **traffic-driven**: the set of names is unknown, and a certificate is obtained during the TLS handshake for whatever hostname the client asked for. In the Caddyfile it is enabled per site with the `tls` directive: ``` https:// { tls { on_demand } reverse_proxy localhost:8080 } ``` and the permission check is configured once in the global options block with `on_demand_tls { ask <url> }`. ## The problem it solves The motivating case is a SaaS with **customer custom domains**. Thousands of tenants each want their storefront on their own hostname. Without on-demand TLS you have to enumerate every tenant hostname in config, obtain certificates for all of them ahead of time, and reload every time a tenant is added or removed — a control-plane job with its own queue, its own failures and its own latency between "customer added a domain" and "HTTPS works". On-demand collapses that to: the customer points DNS at you, and the first handshake produces a certificate. The same shape shows up in white-label products, vanity domains, and platforms that host user-supplied sites. ## Why it is dangerous unrestricted Everything that makes it convenient makes it attackable, because the trigger is **an inbound handshake from an unauthenticated stranger**: - **Issuance amplification.** Anyone can point any DNS name they control at your IP and cause an issuance attempt. Those attempts count against the CA's limits for whatever domains are involved and against your account's order limits — so a stranger can get you rate-limited, which then blocks issuance for your *legitimate* tenants. - **Handshake-time cost.** Each unknown name means a synchronous decision and potentially a full ACME order inside a connection setup. That is CPU and latency an attacker chooses the volume of. - **Storage growth.** Every certificate obtained is written to storage and then maintained and renewed forever unless something prunes it. - **Failed-attempt churn.** Names that will never validate still cost you the attempt, repeatedly, as clients retry. Note what is *not* on this list: an attacker cannot obtain a certificate for a domain they do not control. Validation is still validation. The damage is to your issuance budget and your capacity, not to anyone else's identity. ## The control: the ask endpoint The `ask` setting names an HTTP URL that Caddy consults **before** it attempts issuance for an unrecognised name. Caddy makes a request carrying the candidate hostname as the `domain` query parameter; a success response permits issuance, anything else denies it. In the JSON configuration this is expressed as an on-demand permission module rather than a bare setting, and modern Caddy requires a permission check to be configured before it will enable on-demand at all — the older approach of bounding the damage with issuance interval and burst limits was not a real defence, because a rate limit still lets the attacker consume the whole budget, just more slowly. The endpoint is where your actual policy lives, and it should be cheap and strict: - look the hostname up in the tenant database and approve only domains a paying, verified tenant has registered; - verify the tenant has completed whatever domain-ownership step you require, so the DNS pointing at you is intentional; - answer fast and cache, because this call sits inside a handshake; - fail closed on your own errors — a database outage should deny new issuance, not approve everything. ## Scoping On-demand belongs on the site block that serves tenant traffic, and only there. Your own hostnames should stay config-driven so that their certificates are obtained proactively at startup rather than lazily on first request — you do not want your primary domain's certificate to depend on a handshake arriving, or on the ask endpoint being up. ## What to say about operating it Monitor the ratio of approvals to denials and the absolute rate of issuance; a spike in denials is someone probing you, and a spike in approvals is either growth or a bug in your tenant lookup. Alert on approaching CA limits rather than discovering them during an outage. And have a plan for pruning certificates for tenants who left, because on-demand adds names to storage and nothing removes them for you.

  • Can an attacker use an unrestricted on-demand server to obtain a certificate for a domain they do not own?
    No. The certificate authority still validates control of the name, so a name the attacker does not control will simply fail. What they gain is the ability to make you spend issuance attempts, storage and handshake CPU, and to push your account into rate limits that then block your real tenants. The harm is availability and budget, not misissuance.
  • Why not just cap the issuance rate instead of running an ask endpoint?
    A rate limit bounds the speed of the abuse, not its total. An attacker still consumes the entire issuance budget, just spread out, and legitimate tenants still get blocked behind the cap. An approval endpoint answers a different question — is this name one of ours — which is the only check that separates real traffic from probing.
  • What should the ask endpoint do when its own tenant database is unavailable?
    Deny. Failing open means the first database outage turns into unrestricted issuance and possibly a rate-limit lockout that outlasts the outage itself. Denying means new custom domains stop activating for a few minutes, which is a far cheaper failure. Cache recent approvals so already-known domains keep working while the lookup is down.

saying these in an interview costs you the question

  • Thinks an attacker can obtain certificates for domains they do not own
  • Treats interval and burst limits as sufficient protection
  • Enables on-demand across all sites including primary domains
  • Puts an expensive database join inside the handshake path
  • Fails open when the approval endpoint errors

context