skip to content

Kubernetes lets you add API endpoints either with a CustomResourceDefinition or through the API aggregation layer and an APIService object. Explain how aggregation works inside kube-apiserver and when you would choose it over a CRD.

level: seniorimportance: should knowfreq 30%

answer

  1. APIService = group/version → Service; aggregator proxies
  2. requestheader CA + X-Remote-User headers = front proxy trust
  3. delegated authn/authz: TokenReview + SubjectAccessReview
  4. metrics-server: in-memory, never etcd
  5. broken APIService slows discovery for everyone → CRD by default

basics

~20 s

An APIService registers a group/version and points at a Service; kube-apiserver then proxies every request for that group/version to that extension server, which implements storage and semantics itself. Choose it when you need non-etcd storage or custom verbs.

solid answer

~50 s

With a **CRD**, kube-apiserver itself serves the new resource: it stores objects in etcd, validates them against your OpenAPI schema, and gives you watch, RBAC and kubectl support for free. You write no server. With **aggregation**, you register an `APIService` naming a group/version and a backing `Service`. The aggregator inside kube-apiserver adds that group/version to discovery and **proxies** matching requests to your extension API server, which does everything itself — storage, validation, verbs. It authenticates the front proxy through a requestheader CA and delegates authorisation back to the main API server via `SubjectAccessReview`, so RBAC still applies. Choose aggregation when the data does not belong in etcd (metrics-server keeps a rolling in-memory window; persisting it would wreck etcd), when you need verbs or subresources CRDs cannot express, or when object churn is too high for etcd. Otherwise choose CRDs: no server to run, no availability risk. A broken aggregated API degrades discovery for the whole cluster.

code

yaml · 15 lines
yaml
apiVersion: apiregistration.k8s.io/v1
kind: APIService
metadata:
  name: v1beta1.metrics.k8s.io
spec:
  group: metrics.k8s.io
  version: v1beta1
  groupPriorityMinimum: 100
  versionPriority: 100
  insecureSkipTLSVerify: false
  caBundle: <base64 CA cert>
  service:
    name: metrics-server
    namespace: kube-system
    port: 443

go deeper

for a junior

Know that CRDs are the normal way to add resources, and that metrics-server is a different mechanism that plugs a group/version into the API.

for a middle

Explain the proxy model — APIService points at a Service, kube-apiserver forwards matching paths — and that CRD data lives in etcd while an extension server owns its own storage.

for a senior

Cover the requestheader trust chain, TokenReview/SubjectAccessReview delegation, the discovery blast radius of an unhealthy APIService, and give concrete criteria for choosing aggregation.

for a principal

Treat every APIService as a control-plane availability dependency with its own certificate lifecycle and HA requirements, and set an organisational default of CRDs with a named exception bar.

## Two extension mechanisms, one API surface Both mechanisms make `kubectl get <thing>` work and both are governed by ordinary RBAC, but the machinery behind them is completely different. **CustomResourceDefinition.** You submit a CRD describing a group, versions, kind, scope and an OpenAPI v3 structural schema. kube-apiserver dynamically creates handlers, stores objects in the same etcd it uses for built-ins, validates against your schema (with CEL validation rules if you supply them), and supports watch, the `/status` and `/scale` subresources, printer columns, and server-side apply. You run no additional server; typically you also run a controller, but the controller is just another API client. **Aggregation.** You run a second API server — one built on the same `k8s.io/apiserver` library the core uses — expose it as a Service, and register it: ```yaml apiVersion: apiregistration.k8s.io/v1 kind: APIService metadata: name: v1beta1.metrics.k8s.io spec: group: metrics.k8s.io version: v1beta1 service: name: metrics-server namespace: kube-system port: 443 groupPriorityMinimum: 100 versionPriority: 100 ``` kube-apiserver contains an **aggregator** in front of its own handlers. It merges the registered group/version into discovery, and any request whose path begins `/apis/metrics.k8s.io/v1beta1/...` is reverse-proxied to that Service rather than handled locally. ## The trust chain Proxying raises an obvious question: how does the extension server know who the original caller was, and how does it authorise them? - **Front-proxy identity.** kube-apiserver authenticates itself to the extension server with a client certificate signed by the requestheader CA, and forwards the original user's identity in headers (`X-Remote-User`, `X-Remote-Group`, `X-Remote-Extra-*`). The extension server trusts those headers *only* from a client whose certificate chains to the CA published in the `extension-apiserver-authentication` ConfigMap in `kube-system`. Without that pinning, anyone able to reach the extension server could forge an admin identity — this is the single most important thing to get right when writing one. - **Delegated authentication and authorisation.** The extension server does not implement RBAC. It calls back into the main API server with `TokenReview` (to validate bearer tokens presented directly) and `SubjectAccessReview` (to ask "may this user do this verb on this resource?"). So cluster RBAC still governs the aggregated API, which is why `kubectl auth can-i get pods.metrics.k8s.io` behaves normally. ## Storage freedom is the real differentiator A CRD's objects always land in etcd. That is fine for configuration-shaped data — tens to thousands of objects, changing occasionally. It is disastrous for high-frequency or high-volume data, because etcd is a consensus store with a practical database-size limit and every write costs a Raft round trip. metrics-server is the canonical aggregated API precisely for this reason: resource usage samples for every Pod, refreshed constantly, are held in memory with a short rolling window and never persisted. `kubectl top` and the HorizontalPodAutoscaler read them through `metrics.k8s.io` exactly as if they were normal objects. The same logic drives `custom.metrics.k8s.io` and `external.metrics.k8s.io` adapters that front Prometheus or a cloud metrics service. Other reasons to aggregate: verbs or endpoint shapes a CRD cannot express (streaming, proxying, non-CRUD actions, an arbitrary subresource with custom semantics), a need to expose an existing system's data as Kubernetes objects without copying it, or per-request logic that admission webhooks cannot express. ## The cost, and why CRDs are the default Aggregation introduces a control-plane dependency. The aggregator health-checks each APIService, and an unavailable one is reported as `Available=False` on the APIService object. The user-visible consequence is disproportionate: discovery aggregates every registered group, so a hung extension API makes `kubectl` commands slow or produce `error: unable to retrieve the complete list of server APIs`, even for people who never use that API. `kubectl get pods` becoming slow because metrics-server is unhealthy is a well-known and initially baffling symptom. You also now own: TLS certificates and their rotation, delegated authn/authz wiring, HA for the extension server, version skew against the core API, and correct RBAC on the `extension-apiserver-authentication` ConfigMap. By contrast a CRD has essentially no availability surface of its own. The main exception is a **conversion webhook**, which reintroduces a dependency on the same footing as an aggregated server. ## How to decide Start with a CRD. Move to aggregation only when you can name a requirement a CRD structurally cannot meet: data that must not live in etcd, volume or churn etcd cannot absorb, or verbs outside the CRUD-plus-subresource model. "We want it to feel native" is not a reason — CRDs already feel native. "We produce ten thousand samples a second" is. ## Inspecting what a cluster has ``` kubectl get apiservices # Local means core; others are aggregated kubectl get apiservice v1beta1.metrics.k8s.io -o yaml kubectl get crds ``` An APIService whose `SERVICE` column reads `Local` is served by kube-apiserver itself; anything naming a namespace/service is aggregated and is a live dependency of your control plane.

  • Why can an unhealthy metrics-server make unrelated kubectl commands slow?
    kubectl performs discovery to map kinds to resources, and discovery aggregates every registered group/version — including aggregated ones. If an APIService is unreachable, the aggregator waits on it before completing the discovery document, so every client pays the timeout and may see "unable to retrieve the complete list of server APIs". Aggregated discovery (GA in 1.30) reduces this, but the dependency is real.
  • How does an extension API server know it can trust the identity headers it receives?
    kube-apiserver connects with a client certificate signed by the requestheader CA, whose bundle and allowed header names are published in the `extension-apiserver-authentication` ConfigMap in kube-system. The extension server loads that configuration and accepts `X-Remote-User`/`X-Remote-Group` only from a peer presenting a certificate chaining to that CA with an allowed common name. Skipping this check would let any pod that can reach the service impersonate cluster-admin.
  • Does an aggregated API bypass RBAC?
    No. The extension server delegates authorisation back to the main API server through SubjectAccessReview, so the same Roles and ClusterRoles apply to its resources. It must implement that delegation correctly, though — a server that skips the SubjectAccessReview call is a genuine privilege-escalation hole, since the aggregator will happily proxy any authenticated user's request to it.

saying these in an interview costs you the question

  • Claiming aggregated APIs bypass RBAC or authentication
  • Thinking CRDs can store data outside etcd
  • Reaching for aggregation because it "feels more native" when a CRD would do
  • Not knowing that a failed APIService degrades discovery cluster-wide
  • Believing metrics-server persists metrics as objects in etcd

context