Every gate on the mountain still admits a certificate the authority revoked last week - why is that the designed default, and what actually closes it?
answer
- silence is read as consent
- availability beat enforcement
- the attacker controls the check too
- opt in per certificate, not globally
- short lifetimes beat better policing
basics
~20 sBecause clients soft-fail: when a status check cannot be completed they proceed. An attacker positioned to intercept the connection can also drop the status query, so live checking stops an accident but not an adversary. What closes it is a certificate that demands a stapled answer, pre-distributed revocation data, or short lifetimes.
solid answer
~50 sA revocation check is a second network operation that can fail for entirely innocent reasons - a responder outage, a captive portal, a gate that came online over a thin link. Hard-failing on that turns the issuer's responder into a single point of failure for every site beneath it, so clients overwhelmingly treat an unanswered check as a pass. The consequence is structural rather than accidental: anyone able to sit in the connection's path and use a revoked certificate is also able to drop the status query, and the check becomes a request the attacker chooses to answer. Three things genuinely close it: the TLS Feature X.509v3 extension of RFC 7633 (commonly called Must-Staple), which makes a missing stapled answer fatal for that certificate; revocation data pushed to clients ahead of time so no live fetch is needed; and short-lived certificates, which shrink the window instead of policing it.
code
pseudocode · 15 linesfunction decide(chain, delivered_status, now):
if path_validation_fails(chain, now):
return REJECT
status = evaluate(delivered_status, chain.leaf.certid, now)
if status is REVOKED:
return REJECT
if status is GOOD:
return ACCEPT
// no usable answer: timed out, unsigned, stale, or never delivered
if chain.leaf has tls_feature_extension(status_request):
return REJECT // the certificate promised an answer
return ACCEPT // soft-fail: the default nearly everywherego deeper
Recall that if a client cannot reach the revocation answer it usually carries on and accepts the certificate, so revocation is not an instant off switch.
Explain why availability drove that default and name the mechanism that reverses it for one certificate: an extension in the certificate demanding a stapled answer.
Show the threat-model argument - the attacker who can use the revoked certificate can also silence the check - and pair it with re-keying and shorter lifetimes as the real response.
Own the trade across an estate: how much availability you will spend on enforcement, whether you accept a per-certificate hard-fail with a monitored pipeline, or whether you shorten lifetimes so the question stops mattering.
## Soft-fail is a decision, not a bug Status checking adds a second dependency to a connection. If a client hard-fails whenever it cannot get an answer, then the issuer's responder becomes a availability dependency for every site under that issuer: one responder outage takes them all down, a captive portal blocks the check before the user can log in, and a mountain gate that has just come online over a thin link fails every pass on the queue. Faced with that, clients chose availability. **An unanswered check is treated as a pass.** That is soft-fail, and it is why a revoked certificate is routinely still accepted. ## Why that is worse than it sounds The threat model is where this stops being a tolerable trade-off: - To use a revoked certificate against a victim, an attacker must be in a position to serve traffic for the name - in the network path, or controlling resolution. - That is the *same* position needed to make the status query fail: drop the packets, black-hole the responder's address, or simply return nothing. - So the check reliably catches the case where nobody is attacking - an operator who revoked and forgot - and reliably fails to catch the case it exists for. Stated plainly: **live revocation checking with soft-fail raises the cost of an attack by roughly nothing.** A candidate who can say that, without treating it as scandal, is showing the judgment the question is asked for. ## The three things that actually change the outcome | mechanism | how it closes the gap | what it costs | |---|---|---| | **TLS Feature extension (RFC 7633, "Must-Staple")** | the certificate itself declares that a stapled status will be present, so a conforming client that sees none must fail the connection | a broken stapling pipeline becomes a self-inflicted outage for that certificate, with no way to soften it remotely | | **pre-distributed revocation data** | the client is shipped an aggregated set of revocations ahead of time, so the decision needs no live fetch and an attacker has nothing to block | coverage is curated by whoever compiles the set, not complete; it is a client-vendor mechanism, not something a site operator controls | | **short-lived certificates** | the window in which a stolen key is useful shrinks toward the certificate's own lifetime, so revocation matters less | issuance and deployment must be fully automated; a renewal failure now takes the service down on a schedule | The third is the honest answer, and the direction the ecosystem has taken. If a certificate lives for days, the gap between compromise and expiry is comparable to the freshness window of a status answer anyway, and a mechanism that half-works matters much less. ## Hard-fail, and why it is not simply switched on It is tempting to argue that clients should just hard-fail. The objections are concrete: 1. **Correlated failure.** Every certificate under one issuer shares one responder. Its outage is an outage of everything beneath it - the exact concentration risk operators try to avoid elsewhere. 2. **Networks that intercept by design.** A captive portal must be reachable before any status query can succeed, so a client that hard-fails cannot get onto the network that would let it check. 3. **Blame is misassigned.** The user sees a site failing when the site is fine and a third party is down, and the pressure lands on the operator who cannot fix it. This is why Must-Staple is per-certificate and opt-in: it lets an operator choose hard-fail for one certificate, where they control both the stapling pipeline and the consequences, instead of imposing it globally. ## What to do about it operationally - After a key compromise, **revoke and re-key**. Treat the revocation as bookkeeping and the new certificate as the actual fix, and keep the old certificate's remaining lifetime as your exposure estimate. - **Shorten lifetimes** before you invest in policing revocation - it is the only change that improves the worst case. - If you adopt Must-Staple, **operate the stapling refresh as a monitored job** with alerting on staleness; the extension converts a silent degradation into an outage, which is the point and also the risk. - Do not present revocation checking to stakeholders as a control that stops an adversary. It is hygiene, and the honest statement of what it does is part of the answer.
- Why do clients not simply hard-fail on an unanswered revocation check?Because it concentrates availability on the issuer's responder: one outage would take down every site beneath that issuer, and a captive portal would block the check before a user could join the network that allows it. The failure would also be blamed on the site operator, who cannot fix a third party. Hard-fail is therefore offered per certificate, as something an operator opts into knowingly.
- What is the operational risk of putting the TLS Feature extension on a production certificate?It converts a silent degradation into a hard outage. A conforming client that receives no stapled answer must fail the connection, so any break in the refresh job - a responder outage lasting past `nextUpdate`, a renewal that forgets to re-fetch, a cache that was never warmed - takes the service down for those clients. The extension is in a signed certificate, so it cannot be turned off without replacing the certificate.
- If revocation barely works, why revoke at all?Because it does work against the non-adversarial case and it is the record that matters afterwards: clients that do check and do get an answer stop trusting the certificate, pre-distributed revocation sets can pick the entry up, and the issuer's own records show the certificate was withdrawn and when. What it does not do is contain an attacker who already holds the key, which is why re-keying and a short remaining lifetime are the actual response.
A gate attendant phones head office to ask whether a pass was cancelled, and waves the skier through when the line is dead - because a lift queue halted by a dead phone line costs the resort more than one cancelled pass. Anyone who can cut the line has also decided the answer.
saying these in an interview costs you the question
- Says a revoked certificate stops working everywhere immediately
- Treats soft-fail as an implementation bug rather than a deliberate trade
- Thinks stapling alone makes a missing status answer fatal
- Claims hard-fail everywhere is obviously the right default
- Says short-lived certificates make revocation checking stronger
- Believes an attacker cannot influence whether the status query succeeds