Using Amazon Route 53 weighted records, you want roughly 5% of traffic for api.example.com to reach a newly deployed stack. How do weighted records produce that split, and why will the real share of users differ from 5%?
answer
- relative weights, not percentages
- chosen per query, not per user
- weight 0 parks a record
- all-zero weights means equal, not off
- one cached answer serves a whole resolver
basics
~20 sWeighted records share a name, and Route 53 picks one per query with probability equal to its weight divided by the total — for example 5 against 95. The real user share drifts because each cached answer serves many clients for the whole TTL.
solid answer
~50 sYou create two records with the same name and type, each with its own `SetIdentifier` and a `Weight`; Route 53 chooses one at random per query, with probability weight over the sum of all weights. Weights run 0–255, so 5 and 95 gives about a twenty-to-one split, and setting a record's weight to 0 takes it out of rotation without deleting it. The split is per DNS query, not per user or per request: a caching resolver asks once and serves that one answer to everybody behind it until the TTL expires, and a client that keeps a connection open never re-resolves at all. So the traffic share is approximate and lumpy, especially with large corporate or ISP resolvers. If you need precise, per-request splitting, do it at the load balancer with weighted target groups rather than in DNS.
go deeper
Know that several records can share one name, each with a weight, and Route 53 picks one at random in proportion. Say plainly that weights are relative numbers, not percentages.
Explain selection as weight over the sum of healthy weights, what weight 0 does, and the all-zero special case. Be ready to explain why DNS splits queries rather than users.
Demonstrate the operational sequence: lower the TTL first, wait it out, then shift weights, and watch real request metrics rather than assuming the configured ratio. Say when you would push the split down to the load balancer instead.
Own the choice of layer. Argue when a coarse, globally reachable DNS shift is the right instrument — cross-region, cross-account, hybrid cutovers — and when precision, instant rollback and per-request control justify doing it in the load balancer or the application.
## The mechanism A weighted record set is several records sharing a name and type. Each has: - a `SetIdentifier` — a label unique within the group, so the records can coexist; - a `Weight` — an integer from 0 to 255; - optionally a `HealthCheckId`. When a query arrives, Route 53 sums the weights of the records it considers healthy and picks one at random with probability *weight ÷ sum*. Weights are relative, not percentages: 5 and 95 behaves the same as 1 and 19, and 50 against 50 is an even split whatever the numbers. For a 5% canary, 5 and 95 is the readable choice because the numbers read as percentages even though Route 53 never treats them that way. ```json [ { "Name": "api.example.com", "Type": "A", "SetIdentifier": "stable", "Weight": 95, "TTL": 60, "ResourceRecords": [{ "Value": "203.0.113.10" }] }, { "Name": "api.example.com", "Type": "A", "SetIdentifier": "canary", "Weight": 5, "TTL": 60, "ResourceRecords": [{ "Value": "203.0.113.20" }] } ] ``` Two special cases are worth memorising. A weight of **0** removes a record from selection while leaving it configured — the usual way to park a canary or drain a stack instantly without deleting anything. But if **every** record in the group has weight 0, Route 53 treats them all as equally weighted and distributes across all of them, which is the opposite of what someone expects who set them all to zero hoping to stop traffic. ## Why the observed split is not the configured split The weight governs the *answer to a DNS query*, and there is no fixed relationship between queries and users or requests. **One answer serves many clients.** A recursive resolver — an ISP's, a corporate one, a public one — queries Route 53 once, caches the answer for its TTL, and hands that same address to every client behind it. If a resolver serving 40% of your users happens to draw the canary, the canary briefly receives far more than 5% of traffic. Conversely, a resolver that draws the stable answer sends the canary nothing at all for the life of that cache entry. The larger the resolver population per answer, the coarser the split becomes. **One answer serves many requests.** A client that resolves once and then reuses a keep-alive connection or a connection pool sends thousands of requests over that first choice. Weighted DNS splits *lookups*; your dashboards count *requests*. **Clients cache too.** Application runtimes and containers hold resolved addresses for their own lifetimes, sometimes ignoring the TTL entirely. The practical consequence is that a weighted split is a statistical trend that converges over hours and a large, diverse client base — not a control knob you can trust for a five-minute canary or a small user population. ## TTL is the knob that controls responsiveness Because every cached answer freezes a decision, the record's TTL sets how quickly a weight change reaches users. During a canary or a migration you lower the TTL (60 seconds or less) *before* you start, wait out the old TTL so the long-lived caches expire, and only then begin shifting weights. Raise it back afterwards, since a low TTL means more queries and Route 53 bills per query. ## Health checks change the arithmetic Attaching a health check to each weighted record makes Route 53 exclude unhealthy records from the sum. If the canary is unhealthy, the stable record receives 100% — its weight is now the entire sum. This makes weighted routing a reasonable blue/green mechanism as well as a canary one: two equally weighted stacks, each health-checked, with traffic automatically concentrating on whichever survives. As with all Route 53 groups, if every record is unhealthy it answers as though they were all healthy rather than returning nothing. ## When to use something else Weighted DNS is right when the two targets are genuinely separate endpoints — different regions, different accounts, a legacy data centre alongside AWS — because DNS is the only layer that sits above all of them. When both targets are behind the same load balancer, split there instead: a listener rule with weighted target groups decides per request, takes effect immediately, and is unaffected by caching. Reserve weighted records for the cases where nothing below DNS can see both destinations.
- What happens if you set the weight of every record in the group to 0?Route 53 treats them all as equally weighted and keeps answering with all of them. Zero means "exclude this one from a group that still has weight elsewhere"; it is not a global off switch. To stop answering with a name you delete the records or point them somewhere else — an all-zero group behaves exactly like an evenly weighted one.
- You need the canary to take exactly 5% of API requests for a controlled experiment. Would you still use weighted DNS?No. DNS splits queries, and caching means the request-level share can be far off. For an exact per-request split, do it below DNS — a load balancer rule with weighted target groups, or a feature flag in the application, both of which decide per request and change instantly. Weighted records are for coarse, cross-endpoint shifts.
- How do you prepare the TTL before shifting weights during a migration?Lower the record's TTL well in advance — 60 seconds or less — and wait at least one full old TTL so existing caches expire at the short value. Only then start moving weights, so each change propagates within a minute. Raise the TTL again when the migration is finished, since short TTLs multiply query volume and Route 53 charges per query.
saying these in an interview costs you the question
- Thinks weights must add up to 100
- Says the split is exact per user or per request
- Believes setting all weights to 0 stops traffic
- Expects a weight change to take effect instantly
- Ignores TTL and caching when planning a canary