skip to content

How do you decide whether to pin every Go service to the pure-Go DNS resolver fleet-wide?

level: principalimportance: should knowfreq 26%

answer

  1. parity against host features
  2. the sleeper cost is somebody else's cache
  3. two teams own half the decision each
  4. prefer the lever you can pull back
  5. decide from facts about the environments

basics

~20 s

Weigh environment parity against what the host's resolver stack provides. Pin Go's own resolver when services need identical behaviour everywhere and use nothing beyond files and DNS; do not pin when the platform relies on host caching or features Go cannot reproduce.

solid answer

~50 s

The gain from pinning is determinism: every service resolves the same way on a laptop, in CI and in production, so a class of environment-only failures disappears. The cost is opting out of everything the host's name-service stack provides — any mechanism beyond the hosts file and DNS, mDNS-style names, the C-only environment knobs, and any caching the host does for you. So decide from facts about your environments: enumerate what the C path gives you today, check whether a caching resolver sits between your services and the upstream nameservers, and price the query volume you would send without it. It is also not solely your call — the platform team owns the base image and the node's resolver configuration and can overrule you. Agree the policy, write it down, and prefer the reversible lever, the run-time environment setting, over baking the choice into every binary.

go deeper

for a junior

Know that the resolver a Go binary uses can be fixed deliberately rather than left to the machine, and that this is a deployment decision rather than something you change per call site.

for a middle

Be able to list the mechanisms — run-time environment setting, build-level choice, a resolver value in code — and say which of them survives a rebuild and which can be reverted without one.

for a senior

Show the migration: force it in staging with debug output on, compare resolver decisions and error classifications, watch query volume at the resolver, and move service by service with a cheap reversal.

for a principal

Own the trade and the stakeholders. State what the host stack gives you today, price the lost caching, agree a written default with the platform team, and choose a mechanism whose blast radius matches how confident you are.

## What is actually being decided "Pin the resolver" means removing the run-time choice: every service resolves names with Go's own DNS client, regardless of what the machine it lands on would have chosen. The alternative is the default, where each environment decides and the same binary can take a different path on a laptop, a CI runner and a production node. This is a policy decision, not a code decision, and it has an owner problem attached: the service team writes the code and the build, but the platform team owns the base image and the node's resolver configuration. Either side can invalidate the other's assumption. ## The case for pinning **Parity.** The failure this prevents is the expensive one: a service that works everywhere except the one environment nobody tested, failing for a reason invisible in the source. Pinning collapses the matrix. **Predictable concurrency cost.** Go's own resolver holds a goroutine per in-flight lookup instead of an operating-system thread. On a service that fans out to many names under load, that difference is the gap between a bounded memory profile and thread growth when the resolver is slow. **Simpler artefacts.** A resolution path with no C dependency is one fewer thing that has to match between the build machine and the run machine. ## The case against **You lose the host's features.** Anything the machine's name-service configuration routes somewhere other than the hosts file and DNS stops working. If service discovery in your environment leans on such a mechanism, pinning breaks it — quietly, for some names only. **You lose the host's cache.** This is the one teams most often miss. Go's resolver keeps no answer cache, so if the C path was being served by a caching resolver on the node, pinning sends your full query volume upstream. On a busy fleet that can be a step change in load on a resolver somebody else operates, and the first symptom is their rate limiting, not your latency graph. **You may be overruling the platform.** If the base image and node configuration are maintained centrally, a per-service pin is a local decision with fleet-wide consequences, and it is exactly the kind of thing that gets reverted in an image update the service team never sees. ## How to decide, concretely 1. **Enumerate what the C path provides today** in each environment — what the name-service configuration actually routes, whether any C-only environment knobs are set, whether `.local` names are used. 2. **Find out whether a caching resolver is in the path**, and if so, estimate the query volume you would send upstream without it. Agree that number with whoever runs DNS before you change anything. 3. **Check what your images already force.** A fleet already shipping minimal images without C linkage has made this decision implicitly; the work is then to make it explicit and consistent rather than accidental. 4. **Decide the blast radius you want.** A run-time environment setting is per-deployment and reversible. A build-level choice is guaranteed in the artefact and cannot be undone by whoever runs it. A `PreferGo` resolver in code is library-local and does not give you a fleet policy at all. 5. **Write it down and give it an owner.** A policy that lives in one team's deployment templates is not a fleet policy. Put it where base images are defined, with the reason attached, so a future image change has to consider it. ## Sequencing the change Roll it out the way you would any behaviour change: force the resolver in a staging environment with debug output on, compare the resolver decisions and the failure classifications before and after, watch query volume at the resolver rather than only in your service, then move service by service. Keep the reversal cheap — if the lever is an environment setting, reversal is a redeploy of configuration, not a rebuild of every binary. ## What a strong answer sounds like A strong answer names the trade rather than picking a side: parity and predictable concurrency cost against lost host features and lost caching; identifies caching as the sleeper risk; makes the decision from facts about the actual environments; and recognises that the platform team is a stakeholder who can overrule it, so the outcome should be a written, jointly-owned policy with a reversible mechanism — not a flag one team quietly adds to its own build.

  • What would make you decide against pinning Go's own resolver?
    Evidence that the environment depends on something Go cannot reproduce: a name-service configuration routing lookups beyond files and DNS, mDNS-style names in use, C-only resolver environment knobs set deliberately, or a node-level caching resolver whose absence would push a step change in query volume onto a shared upstream. Any one of those turns the pin from a simplification into an outage waiting for a specific name.
  • How do you keep the decision reversible during an incident?
    Choose the run-time environment setting as the mechanism rather than baking the resolver into the binary, so switching back is a configuration redeploy instead of a rebuild and re-release of every service. Keep both paths exercised somewhere — a build that never runs the other resolver is a reversal you have never tested.
  • Who owns the policy when the platform team maintains the base image?
    Jointly, and it has to be written down where images are defined. The service team owns its build, but the base image and node resolver configuration are the platform's, and an image change can silently invalidate a per-service pin. The workable outcome is a documented default with an explicit opt-out path, plus a check that flags services diverging from it.

saying these in an interview costs you the question

  • Pins the resolver everywhere without checking what the host provides
  • Ignores the query volume that disappearing host caching creates
  • Treats it as a purely technical call with no platform-team stakeholder
  • Bakes the choice into every binary with no run-time reversal
  • Decides from a blog post rather than from the actual environments