A team runs one canary instance beside twenty stable instances and calls it a 5% canary. Why might the canary receive far more or far less than 5% of requests, and what actually controls the real share?
answer
- fleet share is not request share
- the split happens at connect time
- long-lived connections pin clients
- affinity means the same users, repeatedly
- weights or a header rule, then measure
basics
~20 sInstance count only approximates traffic share when balancing happens per request. With long-lived connections, session affinity or client-side balancing, clients are pinned at connect time, so the canary's real share is set by the routing layer — not by arithmetic on instance counts.
solid answer
~60 sOne in twenty-one instances is 5% of the fleet, which is only 5% of requests if every request is independently balanced across all instances. That assumption breaks in several ordinary ways. If the balancer works at the connection level and clients hold long-lived connections — multiplexed HTTP, gRPC streams, WebSockets — the split is decided once at connect time and then frozen, so the canary's share is whatever fraction of *connections* it won, and a client that never reconnects never reaches it. Session affinity pins the same users repeatedly. Client-side balancing and DNS round-robin distribute per client, not per request. And least-connections algorithms send a disproportionate burst to a freshly started instance with zero open connections. What controls the real share is the routing layer: explicit weights at a request-level proxy, or a routing rule keyed on a header or user identifier, which gives you a deterministic cohort instead of a random sample. Whatever you choose, verify the share by measuring request rate per version rather than inferring it from instance counts.
code
nginx · 12 linesupstream app {
server app-stable:8080 weight=95;
server app-canary:8080 weight=5;
}
server {
listen 80;
location / {
proxy_pass http://app;
proxy_http_version 1.1;
}
}go deeper
Understand the distinction being drawn: the fraction of instances running new code is not automatically the fraction of requests it serves, because something else decides where each request goes.
Explain how the balancing layer changes the answer — a request-level proxy can enforce a weight, while a connection-level balancer decides once per connection and then keeps sending everything to the same backend.
Name the concrete distorters you would check: long-lived multiplexed connections, session affinity, client-side or DNS-based balancing, least-connections bursts onto a cold instance, and uneven request cost. Then verify the real share by measuring request rate per version.
Own the platform choice between a weighted random split and a deterministic cohort split, and what each implies for multi-step flows, for which customers bear the risk, and for whether a clean canary result actually generalises to the whole population.
## Instances are not requests "One canary out of twenty-one instances" is a statement about capacity allocation. "5% of traffic reaches the canary" is a statement about request routing. The two coincide only under a specific set of conditions: every request is balanced independently, all instances are weighted equally, connections are short-lived, there is no affinity, and request cost is roughly uniform. Real systems violate at least one of these most of the time. ## Where the split actually happens The decisive question is *what unit the load balancer distributes*. **A request-level (layer 7) proxy** parses each request and picks a backend for it. Here instance count really does approximate request share, and explicit weights make it exact. A weighted upstream can send five requests in a hundred to the canary regardless of how many instances back each pool. **A connection-level (layer 4) balancer** picks a backend once, when the connection is established, and every byte on that connection thereafter goes to the same place. The split is then over connections, not requests. If a client opens one connection and sends a million requests over it, that client contributes a million requests to whichever backend it happened to land on. Modern protocols make this worse rather than better: multiplexed HTTP and gRPC deliberately keep one long-lived connection and run many concurrent streams over it, precisely so they do not pay for handshakes. The practical consequence: with a fleet of long-lived clients, adding a canary instance can leave it receiving almost nothing for hours, because nobody reconnects. The canary looks perfectly healthy — it has barely served anything. ## The other distorters **Session affinity.** If the balancer pins a user to an instance by cookie or source address, the canary serves the same small cohort over and over. Your sample is not 5% of traffic, it is 100% of a handful of users — which may be unrepresentative in either direction, and which concentrates the damage of a bad release onto those specific users for its whole duration. **Client-side balancing and DNS.** When clients resolve a name and choose an endpoint themselves, distribution is a property of every client's own cache and algorithm. Resolvers cache for the record's lifetime and some clients pin the first address they get forever. You are not in control of the split at all. **Least-connections and other adaptive algorithms.** A freshly started canary has zero open connections, so a least-connections policy considers it the least loaded backend and routes a burst to it — the opposite of a cautious 5%. Slow-start or warmup support in the proxy exists specifically to blunt this. **Uneven request cost.** Even a perfect 5% of requests is not 5% of *work*. If the canary happens to catch a share of expensive endpoints, it will show worse resource metrics than the stable pool for reasons that have nothing to do with the new code — and you will chase a phantom regression. **Autoscaling underneath you.** If the stable pool scales from twenty to forty during the canary, the canary's fleet-share halves without anyone touching the release. ## What actually controls the share Two mechanisms, with different properties: **Explicit weights at a request-level proxy.** You state the ratio and the proxy enforces it per request: ``` upstream app { server app-stable:8080 weight=95; server app-canary:8080 weight=5; } ``` This gives a random sample of traffic. It is simple, it is what most canary tooling does under the hood, and it is unstable for any individual user — the same person may hit stable and canary on consecutive requests, which is fine for a stateless read and awful for a multi-step flow whose steps disagree. **Rule-based routing on a request attribute.** Route on a header, an account identifier, or a hash of a user id. This yields a *deterministic cohort*: the same user always gets the same version, so multi-step flows are coherent, internal staff can be routed first, and you can exclude your largest customers. The tradeoff is that a cohort is not a random sample — a hash bucket of users may have systematically different behaviour than the population, so a clean canary does not prove the change is safe for everyone. ## Making the number true Whatever mechanism you pick, do not trust the arithmetic. Measure request rate broken down by version, at the proxy, and confirm it matches what you asked for before you interpret any other metric. If it does not, the usual fixes are: put the split at a request-level proxy rather than a connection-level one; bound connection lifetime with a maximum connection age or maximum requests per connection so long-lived clients are periodically redistributed; disable affinity for the canary window; and enable proxy slow-start so a cold canary is not handed a burst. Comparing latency between versions is meaningless until the denominator is real.
- When would you route a canary by user cohort rather than by random weight?When the change spans multiple requests. A random weight can send steps of one checkout to different versions, so a user meets old and new behaviour in one flow. Hashing on user id keeps each person on one version for the whole canary, which also lets you seed internal staff first and exclude your largest accounts. The cost is that a cohort is not a random sample of behaviour.
- Your canary shows higher CPU per request than the stable pool. Why might that not indicate a regression?Because the two pools may not be doing the same work. An equal share of requests is not an equal share of cost if the canary caught expensive endpoints, and a freshly started instance is also paying cold-start costs — empty caches and unoptimised hot paths — that the long-running stable pool paid off hours ago. Compare per-endpoint metrics after a warmup period before concluding anything.
- How do you keep long-lived clients from starving a canary of traffic?Bound connection lifetime so clients are periodically redistributed — a maximum connection age or a cap on requests per connection at the proxy. Alternatively, move the split to a request-level proxy that picks a backend per request rather than per connection. Without one of these, the canary's share is frozen at whatever fraction of existing connections it happened to win, which may be near zero.
saying these in an interview costs you the question
- One instance out of twenty always means 5% of traffic
- The canary is healthy, so it must be serving its share
- Sticky sessions do not affect a canary because the split is random
- A cold canary showing higher latency proves the new code is slower
- Adding instances to the canary is the only way to increase its traffic share