skip to content

After a record is changed on a DNS zone's primary server, the secondaries keep serving the old data; how do SOA serial, NOTIFY and AXFR/IXFR normally propagate a change, and what likely broke?

level: seniorimportance: should knowfreq 28%

answer

  1. secondaries compare one number
  2. serial not advanced
  3. NOTIFY triggers an early SOA check
  4. full versus incremental transfer
  5. REFRESH, RETRY, EXPIRE

basics

~20 s

Secondaries transfer a zone only when the primary's SOA serial is newer: NOTIFY or the REFRESH timer triggers an SOA check, then AXFR or IXFR copies the change. Usually the serial was not advanced, NOTIFY was lost, or the transfer was refused.

solid answer

~50 s

A secondary never compares zone contents; it compares the `SOA` **serial**. When NOTIFY (RFC 1996) arrives, or its `REFRESH` interval passes, it queries the primary for the zone's `SOA`; if the serial is newer by sequence-space arithmetic it pulls the zone with **AXFR** (whole zone, TCP only, RFC 5936) or **IXFR** (only the differences, RFC 1995). If a check fails it retries every `RETRY` seconds, and after `EXPIRE` without success it must discard the zone. So I would query the `SOA` directly at the primary and each secondary. Same serial on all: the edit was made without advancing the serial, so nobody transfers. Newer serial on the primary: NOTIFY did not arrive (a server outside the notify set, a firewall), the transfer was refused by access control, or TCP to the primary is blocked. A serial moved backwards (by less than half the 32-bit range) looks older, so no secondary fetches it.

go deeper

for a junior

Recall that a DNS zone has a primary where it is edited and secondaries that copy it, and that the SOA serial number tells them a new version exists.

for a middle

Explain the pull cycle: REFRESH or NOTIFY triggers an SOA check, a newer serial triggers AXFR or IXFR, and RETRY and EXPIRE govern failures.

for a senior

Diagnose by comparing serials server by server, then separate an unadvanced or backwards serial from a lost NOTIFY, refused transfer or blocked TCP path.

for a principal

Design replication so failures are loud: monitor serial lag per secondary, alert well before EXPIRE, and keep transfer access control explicit as servers are added.

## The roles RFC 1034 §4.3.5 describes the model. One server is the **primary** for the zone: the zone is edited there, from a master file or through dynamic updates. The **secondaries** hold copies obtained by **zone transfer** and answer authoritatively from them. RFC 1996 §2.1 defines a secondary simply as "an authoritative server which uses zone transfer to retrieve the zone". Secondaries may also feed further secondaries, forming a dependency graph rooted at the primary. ## The serial is the only change detector A secondary never diffs record contents. It compares one number: the `SERIAL` field of the zone's `SOA` record, an unsigned 32-bit value (RFC 1035 §3.3.13). RFC 1034 §4.3.5 says the serial "is always advanced whenever any change is made to the zone", whether by a simple increment or a date-based scheme. Comparison uses **sequence space arithmetic**: the number wraps, and one serial counts as newer than another only if it is ahead by less than half of the 32-bit range. Two consequences follow: - a serial edited *downwards* by less than half the range, for example when switching from a date-based scheme to a counter, looks **older** to every secondary, which then does not transfer; - the quick recovery is to raise the primary's serial above the value the secondaries hold; genuinely reaching a smaller number means wrapping around the 32-bit space in steps smaller than half the range, letting every secondary catch up after each step. Dynamic update (RFC 2136) takes care of this automatically: if an update leaves the serial unchanged, the server must increment it before the change becomes visible. ## The refresh timers Three `SOA` fields drive polling (RFC 1034 §4.3.5, RFC 1035 §3.3.13): | Field | Meaning for a secondary | |---|---| | `REFRESH` | seconds to wait after loading the zone before checking the primary's serial | | `RETRY` | seconds between further attempts when a check fails | | `EXPIRE` | if no check succeeds for this long, the copy is obsolete and must be discarded | The check itself is a plain query for the zone's `SOA`. An equal serial restarts the `REFRESH` wait. `EXPIRE` is the dangerous one: a secondary cut off from its primary for longer than `EXPIRE` stops serving the zone, so an unnoticed transfer failure eventually becomes an outage on that server. ## NOTIFY: not waiting for REFRESH Polling alone means changes can take up to `REFRESH` seconds to spread. **NOTIFY** (RFC 1996) lets the primary push a hint: 1. After a change, the primary sends a NOTIFY (opcode 4) for the zone to its **notify set**, by default every server in the zone's `NS` RRset except the one named in the `SOA` `MNAME` field. 2. NOTIFY normally travels over UDP and is retransmitted until answered; RFC 1996 suggests a 60-second interval and at most 5 retransmissions as reasonable defaults, not requirements. 3. The secondary responds, then behaves as if its `REFRESH` timer had expired: it queries its primary for the `SOA` and, if the serial increased, starts a transfer. NOTIFY carries no authority of its own. Any records in its answer section are only an "unsecure hint", and the decision is always made on the serial. Servers left out of the `NS` RRset, such as stealth secondaries, are not in the default notify set and must be added explicitly, or they only learn of changes at `REFRESH`. ## AXFR and IXFR | | AXFR (RFC 5936) | IXFR (RFC 1995) | |---|---|---| | Sends | the whole zone | only the differences since the client's serial | | Transport | TCP only; "AXFR sessions over UDP transport are not defined" | UDP if the entire reply fits one message, otherwise TCP | | Framing | begins and ends with the zone's `SOA` | a server may fall back to sending the full zone | An AXFR client must serve only a completely transferred copy (RFC 5936 §6), so a transfer that breaks halfway leaves the old copy in place. RFC 5936 says an implementation should let operators restrict AXFR to specific clients and should not default to "open to all". A newly added secondary whose address is not on the primary's allow list will therefore be refused. Transaction signatures (TSIG) are one of the access-control mechanisms RFC 5936 recommends. ## Diagnosing stale secondaries Query each server directly for the zone's `SOA`, without recursion, and compare serials: 1. **All serials equal, data differs on the primary.** The edit was loaded without advancing the serial, or the primary never reloaded the edited file. Advance the serial and reload. 2. **Primary newer, secondaries older.** Look at the push and pull paths: is NOTIFY reaching them (firewall, notify set), is the transfer refused by access control, is TCP to the primary blocked? 3. **Serial on the primary lower than on secondaries.** It went backwards; raise it above the secondaries' value. 4. **Secondary not answering the zone at all.** It may have hit `EXPIRE`; its logs will show repeated failed refresh checks.

  • Why is DNS zone replication driven by the SOA serial rather than by comparing the zone's records?
    One 32-bit number is cheap to query over plain DNS and gives a total order between versions, so a secondary can decide with a single SOA lookup whether it is behind. Diffing contents would need the whole zone every time. The cost is discipline: every change must advance the serial, and a serial moved backwards looks older, so secondaries ignore the new version.
  • When would an IXFR request still result in the whole zone being sent?
    RFC 1995 lets the server choose to send the full zone, for example when it no longer holds the history back to the client's serial, or when the differences would be larger than the zone. The reply then looks like an AXFR, beginning and ending with the SOA. A server that does not recognise IXFR at all leads the client to fall back to AXFR.
  • Why is exceeding EXPIRE worse than simply serving stale data?
    Past EXPIRE, RFC 1034 §4.3.5 says the secondary must assume its copy is obsolete and discard it, so it stops answering for the zone. A silent transfer failure therefore turns, after EXPIRE, into a server that fails every query for the zone while still being listed in the NS set.

saying these in an interview costs you the question

  • Secondaries detect changes by diffing zone contents with the primary.
  • NOTIFY carries the changed records and the secondary applies them.
  • Lowering the serial is fine because any difference triggers a transfer.
  • AXFR can run over UDP when the zone is small.
  • A secondary that loses its primary keeps serving its copy forever.
  • NOTIFY reaches every authoritative server, listed or not, by default.