A DNSSEC-signed zone suddenly fails on validating resolvers while plain DNS lookups still work; which re-signing and rollover mistakes cause this, and how do you prevent them?
answer
- signatures carry an absolute deadline
- who re-signs, and is it running
- a secondary serving an old copy
- a DS that matches nothing
basics
~20 sUsually signatures expired because re-signing stopped or a secondary served a stale copy, or the parent's DS points at a removed key-signing key. Prevent it by monitoring signature expiry on every server and matching DS to DNSKEY before removing keys.
solid answer
~50 sNon-validating resolvers ignore DNSSEC, so "works without validation, fails with it" points at the signatures or the chain, not the records. Three operational causes dominate. **Expired signatures:** every `RRSIG` carries an absolute expiration time, and if the signer stops re-signing (broken automation, an unreachable private key) the zone turns bogus as the oldest signatures pass it. **A stale secondary:** a secondary that lost zone transfers keeps serving its last copy until its SOA expire timer fires, which can be long after the signatures inside that copy expired. **A dangling DS:** the key-signing key the parent's `DS` matches was removed, by a rollover finished too early or an operator change. Prevent them by alerting on time-to-expiry at every authoritative server, re-signing days before expiry, keeping SOA expire well below signature validity, and checking the DS against the `DNSKEY` RRset before any key removal.
go deeper
Recall that DNSSEC signatures have an expiration time and that the parent's DS must match a published key; if either breaks, validating resolvers fail while plain DNS still works.
Explain the three operational causes, stopped re-signing, a stale secondary and a dangling DS, and why a non-validating resolver notices none of them.
Show the monitoring you would run on every authoritative server, the re-sign and SOA expire margins you would set, and the fastest recovery for each cause.
Treat signature validity, re-sign margin and on-call coverage as one design: the validity window is how long the organisation has to notice and fix a broken signer.
## Why only validating resolvers fail A **validating resolver** checks each answer's `RRSIG` signature against the zone's keys and the parent's `DS` record; a non-validating resolver does not. When a zone breaks only for validators, the records are fine and the DNSSEC material around them is not. Validators report SERVFAIL, sometimes with an Extended DNS Error from RFC 8914 such as 7 (Signature Expired) or 9 (DNSKEY Missing). This answer is about what the operator did to cause it. A quick triage, querying each authoritative server directly, separates the usual causes: - every server fails and the earliest signature expiration is already in the past: re-signing stopped; - only some servers fail, and their SOA serial and signature times lag the others: a stale secondary; - signatures are fresh everywhere, but no DS digest at the parent matches a key in the `DNSKEY` RRset: a dangling DS. ## Cause 1: re-signing stopped DNSSEC brings **absolute time** into DNS (RFC 6781, Section 4.4). Each signature is valid only between its inception and expiration times, and RFC 4035 requires a validator to reject it once its clock is past the expiration. TTLs are relative and forgiving; signature expiry is not. Typical triggers: - the signing job or its automation silently stopped; - the signer lost access to the private key (moved files, changed permissions, a hardware module offline); - an old or unsigned copy of the zone was loaded over the signed one. The zone keeps working until the first signatures expire, so the failure surfaces days or weeks after its cause. ## Cause 2: a secondary serving a stale copy A secondary that cannot reach its primary keeps answering from its last transfer until the SOA **expire** timer runs out. RFC 6781 points out that nothing couples that timer to signature expiration, so the secondary can serve expired signatures long before it gives up. Only the queries that reach that server fail, which makes the outage intermittent and hard to reproduce. ## Cause 3: a DS left pointing at a removed key The parent's `DS` names the child's key-signing key (KSK). Remove that key, by finishing a rollover too early, changing DNS operator or regenerating keys, and every cached `DS` matches nothing. RFC 6781 (Section 4.3.3) calls a DS that points at a nonexistent key **security lameness**; it is harmless only while another DS still matches a published key. RFC 7344 is blunt: a child that removes every key its DS set represents can only repair the damage by restoring those keys or by contacting the parent. ## Recovery | Cause | Fastest fix | Why it is still slow | |---|---|---| | Expired signatures | re-sign and push to every server | validators may have cached the bogus result | | Stale secondary | restore transfers, or take the server out of the NS set | the NS change has its own TTL | | Dangling DS | put the old KSK back into the `DNSKEY` RRset | a DS change waits for the parent's registration delay and then the DS TTL | Turning DNSSEC off, by asking the parent to delete the DS, is the last resort; RFC 8078 warns that re-establishing trust in a delegation is hard and takes time. ## Prevention 1. **Monitor time-to-expiry** of the earliest-expiring signature by querying every authoritative server directly, not only the primary. 2. **Re-sign well before expiry.** RFC 6781 suggests ending a signature's publication at least one maximum zone TTL, preferably a few days, before it expires, enough to survive a long weekend. 3. **Keep SOA expire below signature validity.** RFC 6781 suggests roughly a third or a quarter of the validity period, so a cut-off secondary stops answering before its signatures lapse. 4. **Size validity periods** to the time needed to notice and fix a broken signer (a few days at minimum), with TTLs a few times shorter than the validity period. 5. **Gate every key removal** on a check that each DS the parent publishes still matches a key that stays in the zone. 6. **Automate parent updates** with CDS and CDNSKEY so a forgotten manual DS change cannot leave a DS dangling. The common thread: every one of these outages is silent until a deadline passes, so the defence is measuring the distance to that deadline, not waiting for errors.
- Why can a DNSSEC zone keep working for days after its signer stopped, then fail almost all at once?Signatures already published stay valid until their expiration times, which are typically days or weeks away. Nothing breaks until the oldest pass expiry; then RRsets fail as their signatures lapse, often clustered because the zone was signed in one run. The delay hides the cause, which is why time-to-expiry monitoring beats waiting for errors.
- The parent's DS points at a DNSSEC key-signing key you deleted an hour ago; why is restoring the key faster than changing the DS?Restoring the old key is a change in your own zone: once it reaches your servers and cached `DNSKEY` RRsets without it expire, within about one DNSKEY TTL, the old DS matches again. A DS change has to pass the parent's registration delay and then the DS TTL before every cache sees it.
saying these in an interview costs you the question
- If plain DNS answers correctly, the zone's DNSSEC must be fine.
- Signatures stay valid as long as the records' TTLs keep being refreshed.
- A secondary stops answering as soon as the signatures it serves expire.
- Removing the old key-signing key is safe once the new one is published.
- The fastest fix for a dangling DS is always to ask the parent to remove it.