skip to content

In Kubernetes Gateway API, how do HTTPRoute backendRefs weights give a search-autocomplete canary a 0.5% share, and which traffic do those weights never reach?

level: seniorimportance: should knowfreq 42%

answer

  1. share no longer equals pod count
  2. one Service per track
  3. weight over sum of weights
  4. default 1, zero parks it
  5. only gateway traffic is split

basics

~20 s

List a stable and a canary Service in one HTTPRoute rule with weights 995 and 5; the gateway sends 0.5% of requests to the canary whatever its pod count. Pod-to-pod calls via the Service ClusterIP bypass those weights.

solid answer

~50 s

Replica ratios tie share to pod count, so a 0.5% canary would need 200 pods. With Gateway API you give each track its own Service, `autocomplete-stable` and `autocomplete-canary`, and list both in one HTTPRoute rule's `backendRefs`. The share for each is `weight` divided by the sum of weights in that list, so 995 and 5 give 0.5%. Weights are not percentages, need not add up to 100, default to 1 and may be 0, which sends that backend nothing. The canary can now run two pods for redundancy while taking 0.5% of requests, and the gateway proxy balances per request. The catch is scope: the weights apply only to traffic that passes through that Gateway. Other pods calling the `autocomplete` Service by its ClusterIP go through kube-proxy and get whatever the Service selector gives them. The plain `networking.k8s.io/v1` Ingress has no weight field at all, so there a split depends on a controller-specific annotation.

code

bash · 2 lines
bash
kubectl patch httproute autocomplete --type=json -p '[{"op":"replace","path":"/spec/rules/0/backendRefs/1/weight","value":0}]'
kubectl get httproute autocomplete -o jsonpath='{.spec.rules[0].backendRefs[*].weight}'

go deeper

for a junior

Recall that a Gateway API HTTPRoute can list two Services with weights, and each gets weight over the total.

for a middle

Explain the weight rules: default 1, zero means no traffic, no need to sum to 100, and one Service per track.

for a senior

Show where the split stops: in-cluster ClusterIP callers, invalid backends returning errors, and how the gateway is exposed on bare metal.

for a principal

Decide whether routing moves to Gateway API for every service and how internal traffic is covered, since the gateway alone splits only north-south requests.

## Why move the split out of the Service In a replica-ratio canary, a Service balances new connections evenly across Ready pods, so the canary's share is its fraction of the pods. That ties two things together that should be separate: **how much traffic** the canary gets and **how many pods** it runs. A search-autocomplete service with 47 pods can go no lower than about 2%, and a 0.5% canary would need 200 pods. **Weighted routing** fixes this by letting a Layer 7 proxy decide the split per request, from numbers you set. ## How HTTPRoute weights work In Gateway API, an `HTTPRoute` rule has a `backendRefs` list. Each entry names a backend, usually a Service and port, and may carry a `weight`. The rules, from the API's own field documentation: - A backend's share is its `weight` divided by the **sum of all weights in that `backendRefs` list**. - A weight is **not a percentage**, and the weights do not have to add up to 100. - An omitted `weight` **defaults to 1**; the allowed range is 0 to 1,000,000. - A weight of **0** means that entry gets no traffic, which is how you park a backend without deleting it. - Support for weights is **Core**, so every conformant implementation must honour them, with some rounding allowed. - A backend Service with no ready endpoints should get 503 responses for its share, not a silent shift to the other backend. For the autocomplete canary: | backendRef | weight | Share | Pods | |---|---|---|---| | `autocomplete-stable` | 995 | 99.5% | 47 | | `autocomplete-canary` | 5 | 0.5% | 2 | Each Service selects only its own track (`track: stable` or `track: canary`), so the gateway sends each request to one track's pods on purpose. The canary runs two pods so one failure does not take it down, and scaling it to five pods changes nothing about its share. ## Stepping the canary 1. Create `autocomplete-canary` with 2 replicas and its own Service. 2. Set weights 995 and 5 and watch the canary's error rate and latency for the bake period. 3. Move to 950 and 50, then 750 and 250, adding canary pods as its share grows so it has the capacity. 4. Promote by setting the new image on the stable Deployment, then set the canary weight to 0 and scale it down. 5. Abort at any step by setting the canary weight to 0. That is one field change, and the canary pods can stay up for debugging. ## Traffic the weights never see This is the part people miss. An HTTPRoute attached to a Gateway governs requests that **enter through that Gateway**. Internal callers work differently: | Caller | Path | Does the 995/5 split apply? | |---|---|---| | Browser via the Gateway | Gateway proxy to backendRefs | Yes | | Another pod calling `autocomplete` ClusterIP | kube-proxy to the Service's selected pods | No | | Another pod calling `autocomplete-canary` directly | kube-proxy to canary pods | No, it gets 100% canary | If the internal `autocomplete` Service selects only `track: stable`, internal callers never test the canary. If it selects `app: autocomplete`, they get the replica ratio, 2 of 49 pods, about 4%, not 0.5%. Say which one you mean. Splitting east-west traffic by weight needs a mesh, which is a different subject. ## Ingress and the bare-metal cluster - The `networking.k8s.io/v1` **Ingress** resource has no weight field. Controllers that support canaries do it through their own annotations, which differ from controller to controller and do not move between them. That is one reason to put new routing on Gateway API. - On a **9-node bare-metal cluster with no cloud load balancer**, the Gateway's data plane still needs a way in. A Service of type `LoadBalancer` stays pending unless something like MetalLB hands out addresses; otherwise you expose the gateway through NodePort or host networking. None of that changes how weights work, but it decides whether the Gateway is on the request path at all. - The gateway balances **per request** across the pods behind each backend, so keep-alive clients no longer skew the split the way they do with kube-proxy. ```yaml apiVersion: gateway.networking.k8s.io/v1 kind: HTTPRoute metadata: name: autocomplete spec: parentRefs: - name: search-edge rules: - backendRefs: - name: autocomplete-stable port: 80 weight: 995 - name: autocomplete-canary port: 80 weight: 5 ```

  • You set the canary weight to 5 but forgot to create the autocomplete-canary Service. What do users see?
    An HTTPRoute backendRef that points to a missing Service is invalid. Gateway API says requests that would have gone to an invalid backend must get a 500, so about 0.5% of requests fail instead of shifting to stable. The route's status conditions report the unresolved reference, so check `kubectl describe httproute` before raising the weight.
  • Why does a gateway-level split behave better with keep-alive clients than a replica-ratio canary?
    kube-proxy chooses a pod once per TCP connection, so every request on a long-lived connection goes to the same pod. A gateway proxy ends the client connection itself and chooses a backend per request, using the weights, so a single busy client is spread across both tracks in the configured ratio.
  • The canary gets 0.5% at the gateway but 0% of calls from the checkout service. Is that a bug?
    Not in the gateway. Checkout calls the internal `autocomplete` Service by its ClusterIP, which kube-proxy handles using that Service's selector. If the selector is `track: stable`, checkout never reaches the canary. Internal callers need their own plan: a Service that includes canary pods, a mesh route, or accepting that internal paths are tested only after promotion.

saying these in an interview costs you the question

  • HTTPRoute weights are percentages and must add up to 100
  • An omitted weight gives that backend no traffic
  • The gateway's weights also split pod-to-pod calls to the Service
  • The standard Ingress resource has a weight field for canaries
  • Raising canary replicas raises its share under weighted routing