skip to content

An nginx `upstream` block uses `ip_hash;` and one backend receives most of the traffic from a large corporate customer while the others sit idle. What exactly does nginx hash, and why does that produce the skew?

level: middleimportance: nice to knowfreq 36%

answer

  1. deterministic by address, not by user
  2. IPv4 is not hashed in full
  3. one office, one egress address
  4. removing a server reshuffles the map
  5. hash on a cookie distributes per user

basics

~10 s

Nginx's ip_hash keys on the first three octets of the client's IPv4 address, or the whole IPv6 address. Everyone behind one corporate NAT egress address therefore hashes identically and lands on a single backend.

solid answer

~50 s

`ip_hash` is deterministic by design: nginx hashes the client address and maps it to one server so the same client keeps returning to the same place. For IPv4 it uses only the **first three octets**, and for IPv6 the full address. The consequence is that a whole office behind one NAT gateway, or a mobile carrier's CGNAT range, presents as a single key and pins to a single backend — the traffic distributes by network, not by user. The second surprise is that the mapping is over the currently live server set, so adding or removing a server reshuffles many clients at once; the `down` parameter exists precisely to take one out while preserving the hashing for everyone else. If nginx sits behind a proxy, note that `ip_hash` uses the address after the realip module has replaced it, so `set_real_ip_from` is what stops every client hashing to the same value.

code

nginx · 7 lines
nginx
upstream backend {
    hash $cookie_sessionid consistent;
    server 10.0.0.11:8080;
    server 10.0.0.12:8080;
    server 10.0.0.13:8080;
    keepalive 16;
}

go deeper

for a junior

Know that ip_hash sends the same client address to the same backend every time, and that this is affinity rather than even distribution.

for a middle

State the key precisely — first three IPv4 octets, whole IPv6 address — and connect it to NAT and CGNAT as the cause of a hot backend. Know that down preserves the map while removing a server.

for a senior

Show the operational reasoning: affinity keyed on the network fights both uneven client topology and elastic capacity, realip must be in place behind a CDN, and a consistent cookie hash is the targeted fix when affinity is genuinely required.

for a principal

Question whether affinity is needed at all before tuning it: pushing session state out of the backend removes the constraint entirely, and keeping it makes rolling deploys, autoscaling and capacity planning harder for every service behind the proxy.

## What the directive does ```nginx upstream backend { ip_hash; server 10.0.0.11:8080; server 10.0.0.12:8080; server 10.0.0.13:8080; } ``` With `ip_hash`, nginx derives a key from the client address and uses it to select a server, so requests from that address consistently reach the same backend. The documented key is the **first three octets of an IPv4 address**, or the **entire IPv6 address**. That IPv4 detail is the whole answer to the skew question: addresses within the same /24 hash to the same value. ## Why one backend gets the traffic Address-based mapping assumes addresses are distributed roughly like users. Real networks break that assumption comprehensively: - A corporate office of thousands of employees leaves through one or a few NAT addresses. - Mobile carriers place enormous user populations behind CGNAT ranges. - An automated client, scraper or partner integration is one address making a large share of the requests. Any of those is a single hash key. Nginx is behaving exactly as specified; the distribution is uneven because the input is uneven. Adding backends does not help — the heavy key still resolves to exactly one of them. ## The second trap: changing the server set The mapping is computed over the live servers in the group. Remove one and a substantial share of clients are remapped to different backends, losing whatever local state made affinity worth having in the first place. Nginx provides `down` for this: ```nginx upstream backend { ip_hash; server 10.0.0.11:8080; server 10.0.0.12:8080 down; server 10.0.0.13:8080; } ``` Marking a server `down` takes it out of rotation while preserving the current hashing of client addresses for the rest, which is what you want during maintenance. Scaling the group up or down is a different matter: the reshuffle is inherent, and it is why affinity and elastic capacity pull against each other. ## The generic hash directive When you want affinity keyed on something other than the network address, the `hash` directive takes an arbitrary key built from variables, and the optional `consistent` parameter switches to a ketama-style ring so that adding or removing a server remaps far fewer keys: ```nginx upstream backend { hash $cookie_sessionid consistent; server 10.0.0.11:8080; server 10.0.0.12:8080; } ``` Hashing a session cookie distributes per user rather than per network, which fixes the NAT skew directly, at the cost of requiring that the key exists on every request — requests without the cookie fall back to hashing an empty key and pile onto one server, so make sure the key is always present. ## Realip changes the input If nginx sits behind a CDN or another load balancer, the peer address is that hop's. The realip module rewrites the peer address early in request processing, so once `set_real_ip_from` and `real_ip_header` are configured, `ip_hash` keys off the recovered client address rather than the front hop's. Without that configuration the symptom is extreme: every request shares a handful of addresses and effectively one backend receives everything. ## Ordering inside the block One mechanical rule that trips people: when a group uses a balancing method other than the default round robin, the method directive must appear **before** `keepalive` in the block. Config that mixes `ip_hash` and upstream keepalive in the wrong order does not behave as intended. ## What to say in an interview Name the key precisely — three octets for IPv4, the full address for IPv6 — then connect it to NAT and CGNAT as the cause of the skew, then mention the two operational consequences: `down` for preserving the map during maintenance, and a cookie-based `hash ... consistent` when you need per-user rather than per-network distribution. That sequence shows you have read the directive's behaviour rather than assumed it hashes the full address.

  • Why does removing a server from an ip_hash group disturb clients that were not using it?
    The selection is computed over the set of live servers, so shrinking the set changes the mapping for a large share of keys, not only those pointing at the removed server. For planned maintenance use the `down` parameter, which takes the server out while preserving the existing hashing. For genuine scaling, the reshuffle is inherent — `hash ... consistent` limits how much of it happens.
  • How would you get per-user affinity instead of per-network affinity in nginx?
    Use the generic `hash` directive with a key that identifies the user rather than the network, such as `hash $cookie_sessionid consistent;`. The `consistent` parameter uses a ketama ring so a membership change remaps only a fraction of keys. The requirement is that the key is present on every request; requests missing it hash on an empty value and concentrate on one server.
  • Nginx sits behind a CDN and ip_hash sends practically everything to one backend. What is missing?
    The realip configuration. Without `set_real_ip_from` and `real_ip_header`, the peer address is the CDN node's, so only a handful of distinct addresses exist and the hash has almost no entropy. Once realip replaces the peer address early in request processing, ip_hash keys off the recovered client address and the distribution recovers.

saying these in an interview costs you the question

  • Assumes ip_hash hashes the complete IPv4 address
  • Thinks adding backends will relieve a hot hash key
  • Expects the mapping to survive a change in the server set
  • Uses ip_hash behind a CDN without configuring realip
  • Believes address-based affinity distributes evenly across users

context