Why does swapping a DNSSEC zone-signing key in one step - new key in, old key out, zone re-signed - break validation for some resolvers?
answer
- caches outlive your zone edit
- key set and signatures cached separately
- each copy lives for its own TTL
- publish before signing, keep after stopping
basics
~20 sValidating resolvers cache the DNSKEY set and the signatures separately, each for its own TTL. After a one-step swap some pair new signatures with an old key set, or old signatures with a new key set missing the old key; both mismatches are bogus.
solid answer
~50 sA validator needs a signature and the key that made it **at the same moment**, but it may hold them from different points in time: the `DNSKEY` RRset and each `RRSIG` are cached independently, each for its own TTL. Delete the old zone-signing key and re-sign in one step and two mismatches appear. A resolver with the old `DNSKEY` RRset cached receives new signatures made by a key it does not have; a resolver with old signatures cached fetches a new `DNSKEY` RRset that no longer holds the old key. Both end as bogus data, and that resolver's users get SERVFAIL. The fix is ordering plus waiting: publish the new key for at least the propagation delay plus the DNSKEY TTL before signing with it, and keep the old key until every signature it made has expired from caches.
go deeper
Recall that resolvers cache keys and signatures separately for their TTLs and that a validator needs a matching pair. State the rule: publish a new key before signing with it, keep an old key until its signatures are gone.
Explain both failure directions, a new signature against an old cached key set and an old cached signature against a new key set, and show that each wait is a propagation delay plus a TTL.
Show that you would compute the waits from the zone's actual TTLs, confirm every authoritative server has the change before the next step, and treat key removal as the riskiest step.
Frame the tradeoff: long TTLs lower query load but stretch every rollover, so TTL policy and rollover cadence have to be decided together rather than by separate teams.
## What a validator must hold at once DNSSEC proves that DNS data came from the zone owner by attaching a signature, an `RRSIG` record, to every set of records of one name and type (an **RRset**). To check a signature, a **validating resolver** also needs the public key that made it, which the zone publishes in its `DNSKEY` RRset. The key that signs the zone's ordinary records is conventionally called the **zone-signing key (ZSK)**. The trap is that the two pieces travel separately: - an `RRSIG` arrives with the answer it covers and is cached for that answer's TTL (RFC 4034 requires an `RRSIG`'s TTL to match the TTL of the RRset it covers); - the `DNSKEY` RRset is fetched on its own and cached for the DNSKEY TTL; - a resolver can therefore hold a signature from today's zone and a key set from yesterday's, or the reverse. Nothing tells a cache that the zone changed. Until a cached copy's TTL runs out, the resolver keeps using it. ## The two ways a one-step swap fails Suppose the zone `example.net` replaces its ZSK by deleting the old key, adding the new one and re-signing every RRset in the same edit. RFC 6781 (Section 4.1) describes the two failures that follow: 1. **New signature, old key set.** A resolver that cached the `DNSKEY` RRset an hour ago still holds only the old key. It now receives a fresh answer signed by the new key, finds no matching key in its cached set, and marks the answer **bogus**. 2. **Old signature, new key set.** A resolver that cached an answer with its old signature, but whose cached `DNSKEY` RRset has just expired, fetches the new key set. The old key is gone from it, so the cached signature cannot be verified, and that answer is bogus too. Either way the users of that resolver get SERVFAIL for names in the zone, while users of a non-validating resolver see nothing wrong. That is the classic "it only fails on some networks" report. ## The rule that prevents it Every safe rollover method follows two ordering rules: - **Publish before you sign.** A new key must sit in the `DNSKEY` RRset long enough that every cached copy of that RRset already contains it before any signature made with it appears. - **Keep after you stop.** An old key must stay in the `DNSKEY` RRset until every signature it made has expired from every cache. How long is long enough? RFC 7583 writes the waits as formulas. The publication wait is `Ipub = Dprp + TTLkey`: the **propagation delay** (the time for a change to reach every authoritative server) plus the DNSKEY TTL. The retire wait is `Iret = Dsgn + Dprp + TTLsig`: the time to re-sign the whole zone, plus propagation, plus the largest TTL of any signature the old key made. | Step | One-step swap | Pre-publish rollover | |---|---|---| | New key enters the key set | in the same edit as signing | first, unused | | Signing switches to the new key | immediately | after propagation plus the DNSKEY TTL | | Old key leaves the key set | immediately | after re-signing, propagation and the largest signature TTL | | Resolvers that can see bogus data | any with a cache from before the edit | none | ## Why this is a DNSSEC problem and not a plain DNS one Plain DNS tolerates stale caches: an old address record is merely out of date for a while. DNSSEC turns staleness into failure because a validator must find a **matching pair** of signature and key. That is why every rollover method in RFC 6781 and RFC 7583 (pre-publish and double-signature for the ZSK, and the methods for the key-signing key that also involve the parent's `DS` record) is at heart a schedule that keeps old and new material overlapping for at least one cache lifetime. The same logic explains why removal is the dangerous step. Publishing a key early costs little; signing with it too early or deleting the old one too early is what breaks resolvers, and the damage lasts until the stale copies age out. ## What an interviewer is listening for - That you name **caching** as the cause, not slow servers or slow zone transfers. - That you can state **both** failure directions, not only "the new key is not there yet". - That the waits come from **TTLs plus propagation**, not from a rule of thumb such as "a few hours". - That you treat deleting a key as the riskiest moment of any rollover.
- Does a DNSSEC key-signing key suffer the same cache mismatch when it is replaced?Not in the same way. The key-signing key signs only the `DNSKEY` RRset, and that signature travels with the key set, so a resolver never holds one without the other. The risk moves to trust: the parent's `DS` record points at the key-signing key, so replacing it has to be coordinated with the parent and with the DS TTL instead.
- Could you skip the waits by lowering the DNSKEY TTL just before the swap?Only if you lower it well in advance. Resolvers that cached the key set under the old, longer TTL keep it for that long, so a lowered TTL protects you only once the old TTL has itself run out. RFC 7583 assumes constant TTLs in its timelines and warns that changing them around a rollover changes the timings.
A bank changing its official seal first sends every branch a specimen of the new seal, waits until all of them have it on file, and only then starts stamping documents with it; it keeps the old specimen on file until every document stamped with the old seal has been dealt with. Throw the old specimen away early and old documents are rejected; stamp before the specimen arrives and new ones are.
saying these in an interview costs you the question
- Resolvers fetch the DNSKEY set fresh for every validation, so a key swap takes effect instantly.
- The old key can go as soon as nothing in the zone is signed with it any more.
- Only resolvers that have never seen the zone can fail during a key change.
- A rollover's waiting time is a fixed few hours, whatever the zone's TTLs are.