Lookups under a delegated DNS subdomain are intermittently slow, and one of its listed name servers answers REFUSED for the zone; what is a lame delegation, and why does it hurt?
answer
- listed but not serving
- parent and child NS sets disagree
- resolver retries another server
- an abandoned hosted zone
basics
~20 sA lame delegation is an NS record naming a server that does not serve the zone. Resolvers that pick it lose a round trip or a timeout before retrying elsewhere, and lookups fail if every listed server is lame.
solid answer
~50 sRFC 9499 quotes the classic definition: a server "delegated responsibility for providing nameservice for a zone (via NS records) but is not performing nameservice for that zone". It happens when a team moves to new servers but the parent still lists an old one, a secondary is removed without updating the `NS` set, or the zone is deleted from a hosting server. A resolver chooses among the listed servers; when it picks the lame one it gets `REFUSED`, a non-authoritative or upward referral, or silence, and must retry elsewhere, so a fraction of lookups is slow. RFC 8906 §3.1.1 gives the test: every delegated server should answer an `SOA` query for the zone with an `SOA` record. If the lame server sits at a shared hosting service where anyone can create zones, it is also a takeover risk.
go deeper
Recall that a lame delegation is an NS record pointing at a server that does not actually serve the zone.
Explain what a resolver gets from a lame server (REFUSED, a useless referral or silence) and why it retries elsewhere, making lookups intermittently slow.
Diagnose by querying the zone's SOA directly at every server named by the parent and by the child, and recognise a delegation to a deleted hosted zone as a takeover risk.
Own delegation hygiene as a process: periodic parent-versus-child NS audits, and removing delegations as a mandatory step whenever a zone or hosting account is retired.
## Definition A **lame delegation** exists when an `NS` record names a server for a zone, but that server does not serve the zone. RFC 9499 records the classic wording from operational guidance: a nameserver "delegated responsibility for providing nameservice for a zone (via NS records) but is not performing nameservice for that zone (usually because it is not set up as a primary or secondary for the zone)". A second quoted definition adds that "sometimes these hosts (if they exist!) don't even run name servers". The lame `NS` record may sit in the parent (the delegation itself), in the child's own apex `NS` set, or in both. ## How delegations go lame - **Servers moved, parent not updated.** A team moves `dev.example.com` to new servers and fixes its own apex `NS` set, but the parent `example.com` still delegates to the old ones. - **Secondary retired.** A server is removed from service or stops transferring the zone, while the `NS` records still list it. - **Zone deleted at a hosting service.** The subdomain is no longer used, its zone is removed from a shared authoritative service, and the delegation is left pointing there. - **Delegation published too early.** The parent's `NS` records were added before the child zone was loaded on every listed server. - **Typos.** A misspelled server name in the `NS` RRset. RFC 1034 §4.2.2 asks both administrators to keep the `NS` records on each side of the cut consistent, and RFC 8906 advises parent operators "to regularly check that the delegating NS records are consistent with those of the delegated zone"; most lame delegations are the drift those checks would catch. ## What the resolver experiences A recursive resolver that follows the delegation picks one of the listed servers. From a lame one it may get: | Response from the lame server | Why it is useless | |---|---| | `REFUSED` | the server is reachable but not configured for the zone | | a referral upward, or a non-authoritative answer with `AA` clear | the server is not authoritative for the name | | no response at all | the host is gone or runs no DNS service; the resolver waits for a timeout | In each case the resolver has to try another listed server. The effects on users: 1. **Intermittent latency.** Lookups that happen to choose the lame server first pay an extra round trip, or a full timeout if it is silent. Many resolvers track per-server response times and avoid slow servers for a while, but that is implementation behaviour, and each fresh cache pays again. 2. **Reduced redundancy.** With four servers listed and one lame, the zone runs on three while looking like four. 3. **Outage when all are lame.** If every listed server is lame, resolution fails, typically surfacing as `SERVFAIL` from the recursive resolver to its clients. ## Finding one RFC 8906 §3.1.1 gives the core check: "If a zone is delegated to a server, that server should respond to a SOA query for that zone with an SOA record", and "responding with anything other than an SOA record in the answer section indicates a bad delegation". So: - list the servers the **parent** delegates to, and the servers in the child's own apex `NS` set; - query each one directly, without recursion, for the zone's `SOA`; - expect an authoritative answer (`AA` set) with the same serial from every one; - anything else, including a timeout, marks a lame server. ## When lame becomes dangerous: dangling delegations A lame delegation to a server that no longer exists costs latency. A lame delegation to a **shared authoritative hosting service** can cost the subdomain: - the delegation still points at the service's name servers; - the zone was deleted there, so those servers typically answer `REFUSED`; - if the service lets any customer create a zone with that name on those same servers, someone else can create `dev.example.com` and answer **authoritatively** for it, with the parent's delegation sending the world to them. A related case is a name server whose own domain has expired: whoever registers it next controls the servers your delegation names. What an attacker does with such a foothold belongs to security material; the protocol lesson here is that an `NS` record is a standing grant of authority, and removing it is part of decommissioning. ## Fixing - Make every listed server serve the zone, or remove it from both the parent's delegation and the child's apex `NS` set. - Change the child first and the parent second when adding servers; reverse the order when removing them. - Delete delegations for zones you no longer run, rather than leaving them pointing at an empty hosting account.
- Why does a lame DNS server usually cause slowness rather than an outright failure?A delegation lists several servers and a recursive resolver retries another one when a server refuses, gives a useless referral or times out. Only the lookups that try the lame server first pay the extra round trip or timeout. Resolution fails outright only when every listed server is lame or unreachable.
- Why can fixing the child zone's apex NS set leave the delegation lame?Resolvers reach the child through the parent's referral, which uses the parent's own NS records for the cut. If those still list the retired server, fresh resolvers keep being sent to it. Both sides of the cut have to change, and RFC 1034 §4.2.2 makes keeping them consistent a joint responsibility of both administrators.
saying these in an interview costs you the question
- A lame delegation shows up as NXDOMAIN for every name in the zone.
- One lame server among several has no effect users can notice.
- Updating the child's apex NS set fixes what the parent delegates.
- A delegation to a deleted hosted zone is harmless once nobody uses it.
- Lame means the name server is merely slow to respond.