skip to content

On a 140-node Kubernetes cluster shared by 22 teams, other teams' controllers start getting HTTP 429s after one team deploys a new operator. How do you confirm API Priority and Fairness is rejecting them and contain the noisy client?

level: seniorimportance: should knowfreq 40%

answer

  1. server 429 or client wait
  2. PF headers carry UIDs
  3. rejections labelled by level and reason
  4. dump_requests shows the flow
  5. own level, precedence below 9000

basics

~20 s

Confirm that the 429s come from Priority and Fairness by checking the X-Kubernetes-PF response headers and apiserver_flowcontrol_rejected_requests_total. Find the flow in the APF debug dumps, then move the operator into its own small priority level while its team fixes the call pattern.

solid answer

~50 s

First, rule out client-side throttling: those waits are logged by client-go and never reach the server. A real APF rejection is a `429` whose response carries `X-Kubernetes-PF-FlowSchema-UID` and `X-Kubernetes-PF-PriorityLevel-UID`. `kubectl -v=8` shows them, and they map to object UIDs. Next, check `apiserver_flowcontrol_rejected_requests_total` by `priority_level`, `flow_schema` and `reason`, along with the in-queue and wait-duration metrics. Use `kubectl get --raw /debug/api_priority_and_fairness/dump_requests` to find the flow distinguisher that fills the queues. On a default setup, every ServiceAccount lands in `workload-low`, so a hot-looping operator that runs unpaginated LISTs holds many seats and pushes its neighbours into `queue-full` and `time-out` rejections. To contain it, create a dedicated PriorityLevelConfiguration with few shares and `lendablePercent: 0`, and a FlowSchema that matches its ServiceAccount at a precedence below 9000. Then get its team to use a watch cache, pagination and backoff. Don't exempt the operator, and don't blindly raise the inflight limits.

code

bash · 3 lines
bash
kubectl get flowschemas -o custom-columns=NAME:.metadata.name,UID:.metadata.uid,LEVEL:.spec.priorityLevelConfiguration.name
kubectl get --raw /debug/api_priority_and_fairness/dump_priority_levels
kubectl get --raw '/debug/api_priority_and_fairness/dump_requests?includeRequestDetails=1'

go deeper

for a junior

Remember that Kubernetes API 429s can be traced: the response headers and the apiserver_flowcontrol metrics say which priority level rejected the request.

for a middle

Explain how a ServiceAccount's requests land in workload-low, and why expensive LISTs can fill a level even though fair queuing is on.

for a senior

Walk through the triage: headers, rejection reasons, debug dumps, each replica separately. Then contain the client with its own level, and push the owning team to use caching, paging and backoff.

for a principal

Decide how APF is split across 22 teams: which tenants get their own levels, how much lending to allow, which alerts to run, and when to add control-plane capacity.

## Start by finding where requests are held A "throttled cluster" report can mean three different things, and each needs a different fix: - **Client-side waits** inside client-go, logged with the reason `client-side throttling, not priority and fairness`. No request reached the server. - **APF rejections**: kube-apiserver answered `429 Too Many Requests` with `Retry-After`. - **Plain overload**: requests are admitted, but they run slowly because kube-apiserver or etcd is struggling. Ask an affected team for a client log, or for a request run with `kubectl -v=8`, which prints response headers. A 429 from **API Priority and Fairness (APF)** carries `X-Kubernetes-PF-FlowSchema-UID` and `X-Kubernetes-PF-PriorityLevel-UID`. Match those UIDs against `metadata.uid` on the FlowSchema and PriorityLevelConfiguration objects. Also line up the time the 429s started with the operator's rollout. If rejections jump at the moment of the deploy, that is strong evidence, though it does not prove the cause on its own. ## Read the APF signals APF publishes metrics under the `apiserver_flowcontrol_` prefix: | Metric | What it tells you | |---|---| | `apiserver_flowcontrol_rejected_requests_total` | Rejections by `priority_level`, `flow_schema` and `reason` (`queue-full`, `time-out`, `concurrency-limit`) | | `apiserver_flowcontrol_current_inqueue_requests` | How many requests are waiting in each level right now | | `apiserver_flowcontrol_request_wait_duration_seconds` | How long requests waited in a queue before running or being rejected | | `apiserver_flowcontrol_current_executing_seats` | Seats in use; a jump here shows expensive requests holding many seats | | `apiserver_flowcontrol_dispatched_requests_total` | Requests that were admitted, for comparison with rejections | For the live picture, kube-apiserver serves debug dumps. Reading them needs RBAC access to those non-resource URLs. `dump_priority_levels` shows seats and queue totals for each level. `dump_queues` shows each queue. `dump_requests` lists the waiting requests with their FlowSchema, their flow distinguisher and their estimated seats, and `?includeRequestDetails=1` adds the username, verb and path. Every kube-apiserver replica has its own queues, so check all of them, not only the one your load balancer happened to route you to. ## Why one client hurt everyone In this scenario, the new operator is the document-OCR pipeline's dispatcher, and it runs as a ServiceAccount. With the default objects, the `service-accounts` FlowSchema (precedence 9000) sends it to `workload-low`, the same level as every other team's controllers and CI jobs. The `ByUser` distinguisher makes the operator a separate flow, and fair queuing limits how much of the level it can take. Fairness does not make the level any bigger, though. The dispatcher was hot-looping: on every reconcile it listed all Jobs across the cluster without `limit`. The server charged each of those LISTs several seats, and each held them for seconds. The level stayed full, queues grew past `queueLengthLimit` and waits ran out, so other flows started getting 429s with reasons `queue-full` and `time-out`. ## Contain it with a dedicated level The short-term fix is an APF change, made carefully: 1. Create a **PriorityLevelConfiguration** for the operator with small `nominalConcurrencyShares`, `lendablePercent: 0` and short queues. 2. Create a **FlowSchema** that matches the operator's ServiceAccount with a `matchingPrecedence` below 9000, for example 8500, so it takes priority over `service-accounts`. 3. Watch `rejected_requests_total` move from `workload-low` to the new level. Now only the noisy client queues and gets 429s. 4. Remember that new shares slightly shrink every other level's nominal seats, because the server limit is divided by the new share total. Avoid three moves. Putting the operator in `exempt` lets it bypass every limit. Turning off APF with `--enable-priority-and-fairness=false` removes the protection for everyone. Raising `--max-requests-inflight` and `--max-mutating-requests-inflight` without checking kube-apiserver memory lets more expensive LISTs sit in memory at the same time. ## Fix the client, then set guardrails The APF change only buys time, because the operator still behaves badly. Ask its team to: - stop re-listing on every reconcile and read from a local list-then-watch cache; - page any LIST they cannot avoid with `limit` and `continue`; - back off on errors and honour `Retry-After` instead of retrying at once; - set client-go `QPS` and `Burst` to match what the operator really needs. Then turn the incident into standing rules. Alert on sustained `rejected_requests_total` for each level. Document a FlowSchema pattern that new operators can copy. Add a review step for any client that raises its QPS. Keep the operator's dedicated level after the fix, sized for its real peak, so the next regression hits only that team. With 22 teams on one control plane, the default `workload-low` level is shared by all of them, and deciding who gets their own level is part of running the platform.

  • Why is moving the noisy operator into the exempt priority level the wrong fix?
    Exempt requests are never queued or limited. That removes the only thing stopping the operator from filling kube-apiserver, so the load moves from `workload-low` onto the whole server, including node heartbeats and leader election. Exempt is for `system:masters` and health probes, where APF must never block access. A misbehaving client needs a smaller level, not an unlimited one.
  • The team says the operator only needs more capacity. When would raising the server's inflight flags be justified?
    Only when kube-apiserver has spare CPU and memory while the levels are saturated with legitimate work. Every extra seat allows another request to run at the same time, and expensive LISTs are what exhaust memory. Check apiserver memory and latency under peak load first, raise the limits in small steps, and prefer giving shares to the level that needs them. More seats do not fix a client that loops.
  • The 429s appear on only one of three kube-apiserver replicas. What does that tell you?
    Each replica runs its own APF queues and seat budget, so load that is spread unevenly shows up unevenly. A client with long-lived connections can stay pinned to one replica, and a load balancer can send it a disproportionate share. Compare the flowcontrol metrics for each instance before concluding that the configuration is wrong. Spreading the connections can matter as much as tuning shares.

saying these in an interview costs you the question

  • Put the noisy operator in the exempt level so it stops failing.
  • Any 429 a client sees must have come from API Priority and Fairness.
  • Fair queuing means one flow can never affect other flows in its level.
  • Just double --max-requests-inflight; more seats always helps.
  • Disable API Priority and Fairness during incidents to let traffic through.
  • APF queue state is shared across all kube-apiserver replicas.