skip to content

Users of your DNSSEC-validating resolver report that one domain is down, yet it resolves through other resolvers; how do you confirm a validation failure rather than an outage?

level: seniorimportance: should knowfreq 22%

answer

  1. same resolver, two queries
  2. Checking Disabled as the control
  3. read the extended error
  4. why does elsewhere work?

basics

~20 s

Repeat the query to the same resolver with CD set: SERVFAIL without CD but records with CD means validation failed, not reachability. Extended DNS Errors such as DNSSEC Bogus or Signature Expired, and another validator's result, confirm it.

solid answer

~50 s

Send the same query to the same resolver twice, once normally and once with `CD` (Checking Disabled) set. If the plain query gets SERVFAIL and the `CD` query gets `NOERROR` with records, the resolver has the data and is refusing them because they failed validation (RFC 4035 §3.2.2, §5.5). If both fail, look at reachability or delegation instead; an RFC 8914 Extended DNS Error of 22 (No Reachable Authority) points that way. Read any EDE on the failure: 6 DNSSEC Bogus, 7 Signature Expired, 8 Signature Not Yet Valid, 9 DNSKEY Missing, 10 RRSIGs Missing, and 13 Cached Error if you are seeing a stored failure. Then explain "works elsewhere": the other resolver may not validate, or may differ in clock or trust anchors. The fix belongs to the zone operator; on your side, a negative trust anchor scoped to that domain (RFC 7646) is safer than turning validation off.

go deeper

for a junior

Recall that a DNSSEC failure shows up as SERVFAIL, and that repeating the query with Checking Disabled set makes the resolver return the data if signatures were the problem.

for a middle

Explain the plain-versus-CD comparison and what each result pair means, and know the Extended DNS Error codes for DNSSEC Bogus, Signature Expired and No Reachable Authority.

for a senior

Diagnose why the domain works elsewhere: a non-validating resolver, a skewed clock, different trust anchors or cached state, and respond with a scoped negative trust anchor rather than disabling validation.

for a principal

Set the policy in advance: who may add a negative trust anchor, how long it may stay, how you notify the zone operator, and what your monitoring treats as a validation incident.

## The symptom A **validating resolver** checks DNSSEC signatures and, when the signatures on a signed zone fail, answers `SERVFAIL` (`RCODE 2`) with no records (RFC 4035 §5.5). Users see a domain that "is down". Meanwhile someone tries a different resolver and the domain works. Both observations are true, and the protocol gives you the tools to tell which story they tell. The goal is to separate three cases: a **validation failure** (the data exist but are bogus), a **reachability or delegation failure** (the resolver cannot get the data at all), and a **local problem** in your resolver. ## Step 1: the Checking Disabled comparison The `CD` header bit tells a validating resolver that the querier will do its own checking. RFC 4035 §3.2.2 says the resolver should then return the data even if its own policy would reject them, and §5.5 says it returns the full response only when `CD` is set. That makes `CD` a clean control experiment: same resolver, same question, one variable. | Plain query | Same query with `CD` set | Reading | |---|---|---| | `SERVFAIL` | `NOERROR` with records | the resolver has the data and rejects them: **validation failure** | | `SERVFAIL` | `SERVFAIL` | the resolver cannot get an answer: reachability, delegation or server trouble | | `NOERROR`, `AD` clear although the query set `DO` | same | insecure (unsigned or outside any chain of trust): not a validation problem | If the resolver keeps a **BAD cache** (RFC 4035 §4.7, which RFC 6840 §3.1 says resolvers should implement), the `CD` query may be answered from that cache; §3.2.2 says data from it go back to a `CD` query and a `CD`-clear query gets SERVFAIL. Either way, the comparison shows that the refusal is about validation. ## Step 2: read the Extended DNS Error RFC 8914 lets a resolver attach an **Extended DNS Error (EDE)** option, a numeric `INFO-CODE` plus optional text, to any response. Useful codes here: | `INFO-CODE` | Name | What it points at | |---|---|---| | 6 | DNSSEC Bogus | validation ended in the bogus state | | 7 | Signature Expired | no signature currently valid, some or all expired | | 8 | Signature Not Yet Valid | no signature currently valid, some not yet valid | | 9 | DNSKEY Missing | a `DS` exists at the parent but no supported matching `DNSKEY` in the child | | 10 | RRSIGs Missing | signatures were expected and not found | | 13 | Cached Error | the SERVFAIL is coming from the resolver's cache | | 22 | No Reachable Authority | the authoritative servers could not be reached: an outage, not DNSSEC | Two cautions apply. EDE is **unauthenticated** and, per RFC 8914, must not change protocol processing, so treat it as a lead, not proof. And a resolver is not required to send it, so its absence tells you nothing. ## Step 3: explain why it works elsewhere A domain that fails on one validator and succeeds on another resolver is not a contradiction. Likely reasons: - **The other resolver does not validate.** It never checks signatures, so it returns the same bogus data your resolver refused. Rule this out first: it is also the most dangerous one to "fix" by switching users over. - **The clocks differ.** Signatures are valid only between their `Inception` and `Expiration` times, compared with the validator's own notion of the current time (RFC 4035 §5.3.1). RFC 4035 §3.2.2 names an incorrectly set clock on the recursive server as a reason validation can fail there but not elsewhere. - **The trust anchors differ.** A resolver can only validate from the trust anchors it holds; a difference there changes the verdict. - **Cached state differs.** One resolver may still hold an older, valid answer; another may hold a cached failure (`EDE 13`, a BAD cache entry, or a server failure that RFC 2308 lets it keep for up to five minutes). A useful heuristic: if **one** signed domain fails, suspect the zone; if **many unrelated** signed domains start failing at once with `Signature Expired` or `Signature Not Yet Valid`, suspect your own resolver's clock. ## Step 4: respond without giving up validation 1. **Hand the evidence to the zone's operator.** The plain/`CD` pair and the EDE code are concrete. Why the signatures went bad, such as a missed re-signing or a stale `DS` at the parent, is a key-management problem on their side. 2. **If the outage must be relieved first, scope the exception.** RFC 7646 defines **negative trust anchors**, described in RFC 9364 as a way to mitigate validation failures by disabling validation at specified domains. One domain stays unvalidated; every other domain keeps its protection. As an operating practice, remove it once the zone is fixed. 3. **Do not turn validation off globally**, and do not move users to a non-validating resolver: both remove protection from every domain to work around one. 4. **Expect a short tail.** After the zone is fixed, cached failures can persist until their TTL runs out; a BAD cache entry carries a TTL that RFC 4035 says should be small.

  • Many unrelated signed domains start failing on your DNSSEC-validating resolver at once, with Signature Expired or Signature Not Yet Valid; what do you suspect?
    The resolver's own clock. RFC 4035 §5.3.1 compares each signature's `Inception` and `Expiration` with the validator's notion of the current time, and §3.2.2 names an incorrectly set clock on the recursive server as a reason validation fails there but not at a party that validates for itself. One failing zone points at the zone; many at once, with time-window codes, point at the validator.
  • Why is a negative trust anchor a better emergency response than disabling DNSSEC validation?
    A negative trust anchor (RFC 7646) disables validation only at the named domain, so every other signed zone your users reach keeps its protection against forged answers. Turning validation off, or pointing users at a non-validating resolver, removes that protection everywhere to work around one zone's mistake. Remove the exception once the zone's operator has fixed the signatures.
  • After the zone's operator fixes the signatures, some users still get SERVFAIL for a few minutes; why?
    They are being served a cached failure. RFC 2308 lets a resolver cache a server failure for up to five minutes, and a validating resolver should keep a BAD cache of answers that failed validation (RFC 6840 §3.1) with a small TTL (RFC 4035 §4.7). RFC 8914's Extended DNS Error 13, Cached Error, marks such an answer when the resolver sends it.

saying these in an interview costs you the question

  • It resolves on another resolver, so our resolver must be the broken one.
  • SERVFAIL on one domain means its authoritative servers are unreachable.
  • An Extended DNS Error of DNSSEC Bogus proves someone is attacking us.
  • The quick fix is to turn validation off until the domain recovers.
  • Without an Extended DNS Error the failure cannot be DNSSEC-related.