skip to content

A hostname resolves to an address you know is wrong. Using `dig`, how would you establish whether the bad answer is being served from a caching resolver or is what the authoritative nameservers actually publish?

level: seniorimportance: should knowfreq 45%

answer

  1. ask the same question in three places
  2. choose the server you query
  3. the aa flag in the header
  4. TTL is what remains, not what was set
  5. +trace walks from the root

basics

~20 s

Compare three answers with dig: your normal resolver, a walk from the root using +trace, and a direct query to the zone's authoritative servers with @ns1.example.com. Matching answers mean the record itself is wrong; a stale answer only at the resolver, with a counting-down TTL, means cache.

solid answer

~40 s

Ask the same question at three different places and compare. First `dig +short api.example.com` for the answer your host currently gets. Then `dig +trace api.example.com`, which walks the delegation from the root and queries the authoritative servers itself, bypassing your resolver's cache. Finally get the zone's nameservers with `dig NS example.com +short` and query one directly: `dig @ns1.example.com api.example.com`, checking for the `aa` flag in the header to confirm the answer is authoritative. If the authoritative answer is correct and yours is not, you are looking at a cached record — repeat the query and watch the TTL count down, which proves you are being served a remaining lifetime rather than a fresh lookup. If the authoritative answer is also wrong, stop troubleshooting resolvers and go fix the zone.

go deeper

for a junior

Know that dig lets you pick which server answers by writing @server, and that this is how you compare what different servers say about the same name.

for a middle

Explain what +trace does — iterating from the root through the delegation — and why the TTL in a recursive resolver's answer is a countdown that reveals a cached record.

for a senior

Show a disciplined sequence that establishes authoritative ground truth before touching anything, checks every nameserver for a lagging secondary, and separates cache from record so the fix matches the fault.

for a principal

Own the change process: TTL policy before migrations, how long propagation realistically takes across caches you do not control, and why designs that depend on instant DNS changes for failover are fragile.

## Three places an answer can come from When a name returns the wrong address, the answer originated in one of three places, and each demands a different fix: 1. **The authoritative nameservers for the zone** publish it — the record is genuinely wrong and someone must edit the zone. 2. **A recursive resolver's cache** holds an old copy that has not expired yet — you wait out the TTL or flush. 3. **Something local to this host** overrides it — a different class of problem, and one you rule in or out by seeing whether other hosts agree. `dig` is the instrument that tells them apart, because unlike an application it lets you choose *which server you ask*. ## Step one: what am I getting now? ```bash dig +short api.example.com ``` `+short` prints just the data. When you want the full picture, drop it and read the header line, which looks like `;; flags: qr rd ra;` plus a status such as `NOERROR`, `NXDOMAIN` or `SERVFAIL`. `NXDOMAIN` means the name does not exist; `SERVFAIL` from a recursive resolver, when the authoritative servers answer fine, usually points at a validation or upstream failure rather than a missing record. ## Step two: ask a different recursive resolver ```bash dig @9.9.9.9 api.example.com +short ``` If a public resolver returns the correct address and yours does not, the divergence is on your side and cache is the leading suspect. If both are wrong the problem is further out. ## Step three: bypass caches entirely with +trace ```bash dig +trace api.example.com ``` `+trace` makes dig do the iterative work itself: it starts from the root servers, follows the delegation to the TLD servers, follows that to the zone's nameservers, and prints each referral. Because dig is querying those servers directly, no recursive resolver's cache is involved. It answers the question "what does the delegation chain currently say?" — including whether the delegation itself is pointing at the wrong nameservers, which is a failure mode people forget exists. Caveats: `+trace` needs outbound access to arbitrary nameservers, which some networks block, and it resolves from *your* vantage point, so a zone that returns different answers by geography or by view will show you your own view. ## Step four: ask the authoritative servers directly ```bash dig NS example.com +short dig @ns1.example.com api.example.com ``` The header of that reply should contain `aa` — authoritative answer — meaning this server owns the zone rather than repeating what someone told it. If the authoritative servers publish the bad address, your investigation is over: the record is wrong, and no amount of flushing will help. Query *each* listed nameserver, because a common and confusing fault is one secondary that failed to transfer the latest zone and keeps serving the old record to a fraction of the internet. ## Reading cache evidence from the TTL The TTL in a recursive resolver's answer is not the zone's configured TTL; it is the *remaining* lifetime of the cached copy. Run the same query twice with a gap: ```bash dig api.example.com | grep -E '^api' sleep 30 dig api.example.com | grep -E '^api' ``` A TTL falling from 280 to 250 proves the answer came out of cache and tells you exactly how long until it expires. A TTL that resets to the zone's full value each time means you are getting fresh lookups. You can also ask a resolver to answer only from its cache with `dig +norecurse @resolver api.example.com`, which many resolvers honour by returning the cached record or nothing. ## Putting the combinations together - Authoritative correct, resolver wrong, TTL counting down → cache; wait or flush. - Authoritative wrong → fix the zone; everything downstream is behaving correctly. - Nameservers disagree with each other → a broken zone transfer to one secondary. - `+trace` shows a delegation to nameservers you do not recognise → the delegation at the parent is wrong, which is a registrar-level problem. ## What the interviewer is listening for Not flag recitation — the reasoning. The good answer establishes ground truth at the authoritative servers first, uses the TTL as physical evidence of caching rather than guessing, and knows that low TTLs are what make a planned change fast to propagate, which is why you lower them *before* a migration rather than during the incident.

  • The authoritative servers are correct, your resolver is wrong, and the TTL still has an hour left. What are your options?
    Wait, or flush. Flushing works only on caches you control; anything downstream of you keeps the old copy until it expires, and so do resolvers belonging to your users. That is the whole argument for lowering the record's TTL well before a planned change and raising it again afterwards — during an incident it is already too late, because the low value itself has to propagate.
  • Two of a zone's four nameservers return the old address. What does that tell you?
    Zone replication has partly failed: one or more secondaries never received or applied the latest zone. Clients get correct or stale answers depending on which server their resolver happened to pick, which is why the symptom looks intermittent. Compare the SOA serial on each server — the laggards will show an older serial — and fix the transfer or notify path.
  • When would +trace mislead you?
    When the zone serves different answers by client location or by view, since +trace resolves from your machine and shows your view rather than the affected user's. It also fails outright on networks that block direct queries to external nameservers. In both cases, querying named authoritative servers directly, or running the trace from a host near the affected user, is the more trustworthy step.

saying these in an interview costs you the question

  • Flushing the cache always fixes a wrong answer
  • dig's TTL is the value configured in the zone
  • One authoritative server's answer speaks for all of them
  • +trace still goes through my configured resolver
  • A SERVFAIL means the record does not exist

context