An nginx `upstream` block uses `ip_hash;` and one backend receives most of the traffic from a large corporate customer while the others sit idle. What exactly does nginx hash, and why does that produce the skew?
answer
- deterministic by address, not by user
- IPv4 is not hashed in full
- one office, one egress address
- removing a server reshuffles the map
- hash on a cookie distributes per user
basics
~10 sNginx's ip_hash keys on the first three octets of the client's IPv4 address, or the whole IPv6 address. Everyone behind one corporate NAT egress address therefore hashes identically and lands on a single backend.
solid answer
~50 s`ip_hash` is deterministic by design: nginx hashes the client address and maps it to one server so the same client keeps returning to the same place. For IPv4 it uses only the **first three octets**, and for IPv6 the full address. The consequence is that a whole office behind one NAT gateway, or a mobile carrier's CGNAT range, presents as a single key and pins to a single backend — the traffic distributes by network, not by user. The second surprise is that the mapping is over the currently live server set, so adding or removing a server reshuffles many clients at once; the `down` parameter exists precisely to take one out while preserving the hashing for everyone else. If nginx sits behind a proxy, note that `ip_hash` uses the address after the realip module has replaced it, so `set_real_ip_from` is what stops every client hashing to the same value.
code
nginx · 7 linesupstream backend {
hash $cookie_sessionid consistent;
server 10.0.0.11:8080;
server 10.0.0.12:8080;
server 10.0.0.13:8080;
keepalive 16;
}go deeper
Know that ip_hash sends the same client address to the same backend every time, and that this is affinity rather than even distribution.
State the key precisely — first three IPv4 octets, whole IPv6 address — and connect it to NAT and CGNAT as the cause of a hot backend. Know that down preserves the map while removing a server.
Show the operational reasoning: affinity keyed on the network fights both uneven client topology and elastic capacity, realip must be in place behind a CDN, and a consistent cookie hash is the targeted fix when affinity is genuinely required.
Question whether affinity is needed at all before tuning it: pushing session state out of the backend removes the constraint entirely, and keeping it makes rolling deploys, autoscaling and capacity planning harder for every service behind the proxy.
## What the directive does ```nginx upstream backend { ip_hash; server 10.0.0.11:8080; server 10.0.0.12:8080; server 10.0.0.13:8080; } ``` With `ip_hash`, nginx derives a key from the client address and uses it to select a server, so requests from that address consistently reach the same backend. The documented key is the **first three octets of an IPv4 address**, or the **entire IPv6 address**. That IPv4 detail is the whole answer to the skew question: addresses within the same /24 hash to the same value. ## Why one backend gets the traffic Address-based mapping assumes addresses are distributed roughly like users. Real networks break that assumption comprehensively: - A corporate office of thousands of employees leaves through one or a few NAT addresses. - Mobile carriers place enormous user populations behind CGNAT ranges. - An automated client, scraper or partner integration is one address making a large share of the requests. Any of those is a single hash key. Nginx is behaving exactly as specified; the distribution is uneven because the input is uneven. Adding backends does not help — the heavy key still resolves to exactly one of them. ## The second trap: changing the server set The mapping is computed over the live servers in the group. Remove one and a substantial share of clients are remapped to different backends, losing whatever local state made affinity worth having in the first place. Nginx provides `down` for this: ```nginx upstream backend { ip_hash; server 10.0.0.11:8080; server 10.0.0.12:8080 down; server 10.0.0.13:8080; } ``` Marking a server `down` takes it out of rotation while preserving the current hashing of client addresses for the rest, which is what you want during maintenance. Scaling the group up or down is a different matter: the reshuffle is inherent, and it is why affinity and elastic capacity pull against each other. ## The generic hash directive When you want affinity keyed on something other than the network address, the `hash` directive takes an arbitrary key built from variables, and the optional `consistent` parameter switches to a ketama-style ring so that adding or removing a server remaps far fewer keys: ```nginx upstream backend { hash $cookie_sessionid consistent; server 10.0.0.11:8080; server 10.0.0.12:8080; } ``` Hashing a session cookie distributes per user rather than per network, which fixes the NAT skew directly, at the cost of requiring that the key exists on every request — requests without the cookie fall back to hashing an empty key and pile onto one server, so make sure the key is always present. ## Realip changes the input If nginx sits behind a CDN or another load balancer, the peer address is that hop's. The realip module rewrites the peer address early in request processing, so once `set_real_ip_from` and `real_ip_header` are configured, `ip_hash` keys off the recovered client address rather than the front hop's. Without that configuration the symptom is extreme: every request shares a handful of addresses and effectively one backend receives everything. ## Ordering inside the block One mechanical rule that trips people: when a group uses a balancing method other than the default round robin, the method directive must appear **before** `keepalive` in the block. Config that mixes `ip_hash` and upstream keepalive in the wrong order does not behave as intended. ## What to say in an interview Name the key precisely — three octets for IPv4, the full address for IPv6 — then connect it to NAT and CGNAT as the cause of the skew, then mention the two operational consequences: `down` for preserving the map during maintenance, and a cookie-based `hash ... consistent` when you need per-user rather than per-network distribution. That sequence shows you have read the directive's behaviour rather than assumed it hashes the full address.
- Why does removing a server from an ip_hash group disturb clients that were not using it?The selection is computed over the set of live servers, so shrinking the set changes the mapping for a large share of keys, not only those pointing at the removed server. For planned maintenance use the `down` parameter, which takes the server out while preserving the existing hashing. For genuine scaling, the reshuffle is inherent — `hash ... consistent` limits how much of it happens.
- How would you get per-user affinity instead of per-network affinity in nginx?Use the generic `hash` directive with a key that identifies the user rather than the network, such as `hash $cookie_sessionid consistent;`. The `consistent` parameter uses a ketama ring so a membership change remaps only a fraction of keys. The requirement is that the key is present on every request; requests missing it hash on an empty value and concentrate on one server.
- Nginx sits behind a CDN and ip_hash sends practically everything to one backend. What is missing?The realip configuration. Without `set_real_ip_from` and `real_ip_header`, the peer address is the CDN node's, so only a handful of distinct addresses exist and the hash has almost no entropy. Once realip replaces the peer address early in request processing, ip_hash keys off the recovered client address and the distribution recovers.
saying these in an interview costs you the question
- Assumes ip_hash hashes the complete IPv4 address
- Thinks adding backends will relieve a hot hash key
- Expects the mapping to survive a change in the server set
- Uses ip_hash behind a CDN without configuring realip
- Believes address-based affinity distributes evenly across users