A JVM service keeps opening connections to an instance address that was replaced twenty minutes ago, even though `dig` run on the same host already returns only the new address and the record's TTL is 60 seconds. What is still holding the old address, and how would you fix it?
answer
- dig proves the resolver, not the app
- the runtime caches on its own clock
- a reused connection never re-resolves
- networkaddress.cache.ttl, not the record's TTL
- cap connection age as well
basics
~20 sThe runtime and the connection pool are holding the old address, not DNS. A JVM caches lookups under its own networkaddress.cache.ttl setting, ignoring the record's TTL, and an already-open pooled connection never re-resolves at all.
solid answer
~60 s`dig` proves the resolver path is fine, so the stale copy lives above it, inside the process. Two layers do that. First, the JVM keeps its own address cache: successful lookups are held for whatever the `networkaddress.cache.ttl` security property says — historically `-1`, meaning forever, under a security manager — and that cache never looks at the DNS record's TTL. Second, and usually the bigger effect, an HTTP or gRPC connection pool is bound to an *address*, not a name: while a keep-alive connection is being reused, no lookup happens at any TTL. A client that resolved once when it was constructed behaves the same way. I would confirm it by listing the process's established peers and comparing them with the current answer, then fix it by setting a short positive cache TTL and a small negative TTL, capping the pool's maximum connection lifetime so connections are periodically rebuilt, and — for anything needing fast failover — putting a virtual IP or a registry in front so the caller's target address stops changing.
code
bash · 3 lines# compare what the process is actually talking to with what DNS says now
ss -tnp state established '( dport = :8443 )'
dig +short api.internal.example Ago deeper
Know that a program does not look up a hostname on every request — it caches the answer and reuses connections — so a DNS change does not take effect the moment you make it.
Explain the two caches by name: the runtime's own address cache with its configurable TTL, and the connection pool that binds a socket to an address so no lookup happens at all while it is reused.
Demonstrate the diagnosis: prove where the stale copy lives by comparing live peers with the current answer, then fix it with a bounded cache TTL and a bounded connection lifetime rather than a restart.
Argue the structural point — DNS TTLs are not a traffic-withdrawal mechanism. Decide where the platform needs push-based membership or health-checked virtual IPs so failover time is a property of the platform, not of each client's caching defaults.
## Read the evidence before guessing `dig` opens its own socket and asks the resolver directly. A correct `dig` answer therefore proves exactly one thing: the zone data and the resolver path are healthy. It says nothing about the application, because the application never asks the resolver when it already has an answer of its own. That is the whole diagnosis in one sentence — **the stale copy is above the resolver, inside the process.** ## Layer one: the runtime's address cache The JVM caches name lookups inside `InetAddress`. Two security properties in `java.security` govern it: - `networkaddress.cache.ttl` — how many seconds a *successful* lookup is retained. `-1` means cache for the lifetime of the JVM; `0` disables caching. - `networkaddress.cache.negative.ttl` — how long a *failed* lookup is retained. The critical property of this cache is that **it does not read the DNS record's TTL**. Whatever the zone publishes, the JVM keeps the answer for its own configured period. Historically the positive default depended on whether a security manager was installed: with one, it cached forever; without one, a short implementation-defined period of about 30 seconds. Since security managers are deprecated and disabled by default in recent JDKs, the short default is what most services get — but base images, security-hardening scripts and framework defaults all edit `java.security`, so the effective value is worth checking rather than assuming. You can override it at launch without touching the file: ```bash java -Djava.security.properties=/app/dns.properties -jar app.jar ``` The negative cache deserves its own attention. If a lookup fails during a registry or DNS blip and the negative TTL is high, an instance that has come back stays invisible for that whole period, which looks exactly like the positive-cache symptom in reverse. ## Layer two: the connection pool — usually the real culprit A pooled keep-alive connection is a socket to a *specific address*. Once established, reusing it involves no name lookup whatsoever, so a TTL of zero and an address cache of zero would both change nothing. A service that keeps a warm pool to a busy dependency can keep talking to a decommissioned host for as long as that host still accepts connections and the pool never evicts the entry. This is why the symptom so often outlives every TTL in sight. The fix is a bounded connection lifetime: most HTTP clients expose a maximum connection age (independent of idle timeout) that forces the pool to discard and rebuild connections periodically, and the rebuild is what triggers a fresh lookup. Idle eviction alone is not enough for a continuously busy pool, because the connection is never idle. ## Layer three: clients that resolve once A third pattern hides in library code: a client constructed with a hostname that resolves it eagerly and stores the resulting addresses for the object's lifetime. Nothing later re-resolves. The tell is that a restart fixes it permanently until the next change. ## Confirming which layer it is Compare the *live* peers against the *current* answer. If the process's established connections point at the old address and no new connections are being opened, it is the pool. If new connections are still going to the old address, it is the runtime cache or a resolve-once client. Checking the effective value of the cache property inside the running image, rather than the value in your repository, separates the last two. ## Fixes, from tactical to structural 1. Set a short positive cache TTL (tens of seconds) and a very small negative TTL. 2. Cap the pool's maximum connection lifetime so addresses are revisited on a bounded schedule. 3. Make failure handling re-resolve: on a connection error, drop the cached address rather than retrying the same dead peer. 4. Structurally, stop asking DNS to express membership. Put a health-checked virtual IP or proxy in front of the fleet so the caller's address stops changing, or use a registry that pushes membership changes. Both replace "wait for a cache to expire" with "be told". ## The general lesson A DNS TTL is a request to cooperative caches, not a mechanism for withdrawing traffic. Every layer between the zone and the socket keeps its own copy on its own schedule: the recursive resolver, any host or sidecar stub cache, the runtime, the client library, and the open connection itself. Fast, predictable failover comes from a push-based membership source or a proxy that health-checks its own backends — not from a smaller number in the zone file.
- Why doesn't setting the DNS record's TTL to zero solve this?Because neither layer that is holding the address consults it. The JVM's address cache runs on its own configured period and never reads the record's TTL, and a connection already open to an address performs no lookup at all while it is being reused. A zero TTL only raises query volume.
- What is `networkaddress.cache.negative.ttl` for, and why does leaving it high hurt?It controls how long a failed lookup is remembered. If a name fails to resolve during a brief outage and the negative TTL is long, the caller keeps treating the dependency as non-existent well after it has recovered — an outage that outlives its cause. Values of zero to a few seconds are typical.
- How does putting a virtual IP in front of the dependency change the picture?The caller resolves one address that essentially never changes, so caching at every layer stops mattering for membership. Instance churn moves behind the proxy, which detects failures with its own health checks in seconds rather than waiting for caches to expire. The cost is an extra hop and a component to operate.
saying these in an interview costs you the question
- The DNS TTL controls when the application re-resolves the name
- dig returning the new address proves the app will use it
- The only remedy is restarting the service after every change
- Keep-alive connections re-resolve their hostname periodically
- Java always caches DNS forever and nothing can change it