How does Kubernetes API Priority and Fairness decide which priority level an incoming API request lands in, and what happens when that level is full?
answer
- classify first, then limit
- lowest precedence number wins
- shares divide the summed inflight limits
- flow equals schema plus distinguisher
- shuffle-sharded queues, then 429
basics
~20 skube-apiserver checks each request against the FlowSchema objects, lowest matchingPrecedence first, and sends it to the PriorityLevelConfiguration the matching schema names. If that level has no free seats, the request waits in a fair queue by flow or gets HTTP 429.
solid answer
~50 sAPI Priority and Fairness adds the `--max-requests-inflight` and `--max-mutating-requests-inflight` limits together into one server budget of seats (600 by default). It divides that budget among PriorityLevelConfigurations in proportion to their `nominalConcurrencyShares`. Each request is checked against FlowSchemas. Among the ones that match its user, groups, verb and resource, the one with the lowest `matchingPrecedence` wins and names the level. Inside a `Limited` level with free seats, the request runs right away. Otherwise it is assigned to a **flow**: the schema plus a `ByUser` or `ByNamespace` distinguisher. The flow is shuffle-sharded into one of a few queues, and fair queuing serves those queues, so one noisy flow cannot starve the others. When the queue is full or the wait runs too long, the server returns `429` with `Retry-After`. Two objects are mandatory. `exempt` covers `system:masters` and is never limited. `catch-all` matches anything left over, has a tiny share, and rejects rather than queues.
code
yaml · 38 linesapiVersion: flowcontrol.apiserver.k8s.io/v1
kind: PriorityLevelConfiguration
metadata:
name: ocr-batch
spec:
type: Limited
limited:
nominalConcurrencyShares: 10
lendablePercent: 0
limitResponse:
type: Queue
queuing:
queues: 16
handSize: 4
queueLengthLimit: 50
---
apiVersion: flowcontrol.apiserver.k8s.io/v1
kind: FlowSchema
metadata:
name: ocr-pipeline
spec:
priorityLevelConfiguration:
name: ocr-batch
matchingPrecedence: 8500
distinguisherMethod:
type: ByUser
rules:
- subjects:
- kind: ServiceAccount
serviceAccount:
name: ocr-dispatcher
namespace: doc-ocr
resourceRules:
- verbs: ["*"]
apiGroups: ["*"]
resources: ["*"]
namespaces: ["*"]
clusterScope: truego deeper
Remember the chain: a FlowSchema matches the request and names a priority level, the level has limited seats, and extra requests wait or get HTTP 429.
Explain how precedence picks the FlowSchema, how shares divide the summed inflight limits, what makes a flow, how shuffle sharding works, and the three rejection reasons.
Show you can read the default objects, place a new client with a precedence below 9000, and predict how extra shares shrink the other levels.
Discuss APF as a tenancy policy: which identities get their own level, how much lending to allow, and why putting a client in exempt is a last resort.
## Why the API server needs it kube-apiserver can serve only so many requests at once. Before **API Priority and Fairness (APF)**, two flags capped concurrency: `--max-requests-inflight` (default 400) and `--max-mutating-requests-inflight` (default 200). Those caps did not care who was asking, so one hot-looping controller could fill the server and starve node heartbeats and leader-election renewals. APF is on by default and controlled by `--enable-priority-and-fairness`. It keeps the two flags but **adds them together** into one server-wide concurrency limit (600 **seats** by default) and divides that limit among priority levels. The point is isolation more than throughput. Node heartbeats and leader-election renewals should keep working while a misbehaving client is slowed down. ## Classification: the FlowSchema A **FlowSchema** (API group `flowcontrol.apiserver.k8s.io`, version `v1`) is a classification rule. Its spec has four fields: - `rules`: each rule lists `subjects` (a `User`, `Group` or `ServiceAccount`) plus `resourceRules` (verbs, API groups, resources, namespaces, `clusterScope`) and/or `nonResourceRules` (verbs and URLs such as `/healthz`). - `matchingPrecedence`: an integer from 1 to 10000, default 1000. If several FlowSchemas match, the one with the **lowest** value wins. - `priorityLevelConfiguration.name`: the level the request goes to. - `distinguisherMethod`: `ByUser`, `ByNamespace`, or not set. It decides how requests are split into **flows**. The default objects kube-apiserver creates show the pattern: | FlowSchema | Precedence | Priority level | |---|---|---| | `exempt` (group `system:masters`) | 1 | `exempt` | | `probes` (`/healthz`, `/readyz`, `/livez`) | 2 | `exempt` | | `system-leader-election` | 100 | `leader-election` | | `system-node-high` | 400 | `node-high` | | `kube-controller-manager`, `kube-scheduler` | 800 | `workload-high` | | `service-accounts` (all ServiceAccounts) | 9000 | `workload-low` | | `global-default` (all users) | 9900 | `global-default` | | `catch-all` | 10000 | `catch-all` | ## Priority levels and seats A **PriorityLevelConfiguration** has `type: Exempt` or `type: Limited`. - **Exempt** requests are never queued or limited. - A **Limited** level gets a nominal share of the server limit in proportion to its `nominalConcurrencyShares`: `NominalCL = ceil(ServerCL × shares / sum of shares)`. With the default objects, the shares add up to 245. So `workload-low` (100 shares) nominally gets `ceil(600 × 100 / 245)` = 245 seats, and `global-default` (20 shares) gets 49. - `lendablePercent` lets a level lend its idle seats to other levels. `borrowingLimitPercent` limits how much a level may borrow, and leaving it unset means no limit. - Most requests take one seat. A LIST that the server expects to load many objects is estimated at several seats. - A long-running WATCH holds its seats only while the server sends its initial events, not for the whole life of the stream. ## Queuing and fairness Here is what happens when a Limited level with `limitResponse.type: Queue` has no free seat: 1. The request's **flow** is the pair (FlowSchema name, distinguisher value). With `ByUser`, the distinguisher value is the username. 2. The flow is hashed and **shuffle-sharded**. The server deals a hand of `handSize` queues from the level's `queues`, and the request joins the shortest queue in that hand. 3. The queues are served by **fair queuing**, which limits how much of the level a heavy flow can take while others are waiting. 4. The request is rejected if its queue already holds `queueLengthLimit` requests, or if it waits too long. Each kube-apiserver instance keeps its own queues. In a highly available control plane, every replica applies these limits separately. ## Rejection: HTTP 429 A rejected request gets **`429 Too Many Requests`** with a `Retry-After` header. A level with `limitResponse.type: Reject` skips queuing and rejects as soon as it is full. The metric `apiserver_flowcontrol_rejected_requests_total` records the level, the FlowSchema and a reason: `queue-full`, `time-out` or `concurrency-limit`. client-go retries these responses on its own after the `Retry-After` delay, so a short burst of rejections often never reaches application code. ## The mandatory exempt and catch-all objects `exempt` and `catch-all` are **mandatory**. kube-apiserver recreates them if they are missing and resets their spec, except for `nominalConcurrencyShares` and `lendablePercent` on the `exempt` level. The `catch-all` FlowSchema matches every authenticated and unauthenticated request at precedence 10000, so every request gets classified. Its level has only 5 shares and rejects instead of queuing, because it is a safety net and not meant for real traffic. The other default objects are called **suggested**. kube-apiserver updates them only while their `apf.kubernetes.io/autoupdate-spec` annotation is `true`, so an operator who customizes one can stop the server from overwriting it.
- Is API Priority and Fairness related to a Pod's PriorityClass?No. The two only share the word "priority". A PriorityClass ranks Pods for kube-scheduler, which uses it to order pending Pods and to preempt lower-priority ones. APF ranks HTTP requests to kube-apiserver and limits how many run at once for each level. A Pod's priorityClassName has no effect on how its controller's API calls are classified. Only the requester's identity and the request's verb and resource matter to a FlowSchema.
- Adding the ocr-batch level with 10 shares changes the other levels. By how much?The share total rises from 245 to 255, so every Limited level's nominal seats shrink a little. `workload-low` drops from `ceil(600 × 100 / 245)` = 245 to `ceil(600 × 100 / 255)` = 236. The new level gets `ceil(600 × 10 / 255)` = 24. The server limit stays at 600, so adding shares moves capacity from one level to another rather than creating it.
- Why does APF shuffle-shard flows into a hand of queues instead of giving each flow its own queue?The number of distinct users or namespaces has no upper limit, while the number of queues is fixed by the `queues` field. Hashing each flow to a small hand of `handSize` queues and joining the shortest one means a heavy flow fills only its own hand. Most other flows have at least one queue in their hand that the heavy flow does not share, so they keep being served.
saying these in an interview costs you the question
- The FlowSchema with the highest matchingPrecedence number wins the match.
- API Priority and Fairness is the same mechanism as Pod PriorityClass preemption.
- Requests that match no custom FlowSchema bypass all limits.
- Each API request always costs exactly one seat, whatever it asks for.
- Deleting the catch-all FlowSchema removes it permanently from the cluster.
- Fair queuing gives every flow its own dedicated queue.