skip to content

Operators and CRDs

Extending the API with your own types: CustomResourceDefinitions, controllers that reconcile them, admission webhooks that intercept writes, and operators encoding day-2 ops. Interviewers probe it to see if you treat Kubernetes as an extensible control plane.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

page 1 of 2

In a Kubebuilder project built on controller-runtime, what does the framework provide, and what do you actually write inside Reconcile(ctx, req)?

level: juniorimportance: must knowfreq 58%

answer

  1. library versus scaffolding
  2. one manager, shared caches
  3. req is only namespace/name
  4. Get, converge children, write status
  5. markers become ClusterRole and CRD

basics

~20 s

controller-runtime supplies a manager that runs shared caches, clients, work queues, leader election and metrics; Kubebuilder scaffolds the API types, markers and Makefile. You write Reconcile: fetch the object named by req, converge actual state to spec, then update status.

solid answer

~40 s

controller-runtime gives you a `Manager` (from `ctrl.NewManager`) that owns a shared informer cache, a client, the scheme, leader election, health probes and a metrics endpoint, and it starts every controller registered on it. You register a controller with the builder, `ctrl.NewControllerManagedBy(mgr).For(&GameSession{}).Owns(&appsv1.Deployment{}).Complete(r)`, and implement one method: `Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error)`. The `req` carries only a namespace and a name. Inside, you `Get` the object, stop quietly if it is gone (`client.IgnoreNotFound`), compute the child objects it should have, create or patch them, and write what you observed to the `/status` subresource. Kubebuilder scaffolds the Go types, `SetupWithManager`, `cmd/main.go`, the `config/` Kustomize tree and a Makefile. `// +kubebuilder:rbac` and `// +kubebuilder:validation` markers become the ClusterRole and the CRD schema when you run `make manifests`.

code

go · 7 lines
go
func (r *GameSessionReconciler) SetupWithManager(mgr ctrl.Manager) error {
	return ctrl.NewControllerManagedBy(mgr).
		For(&gamesv1.GameSession{}).
		Owns(&appsv1.Deployment{}).
		Owns(&corev1.Service{}).
		Complete(r)
}

go deeper

for a junior

Recall the split: controller-runtime is the library with the manager and client, and Kubebuilder is the generator. Be able to write the signature Reconcile(ctx, req) and name the get, converge, update-status steps.

for a middle

Explain what the manager starts and in what order: caches sync before controllers run. Show how For and Owns route events to one reconciler, and how markers become YAML through make manifests.

for a senior

Show you have shipped one: leader election for multiple replicas, status written through the subresource, RBAC kept minimal by pruning markers, and SetupWithManager kept declarative rather than scattered through main.

for a principal

Weigh the framework choice for a platform: Go with controller-runtime versus the Java Operator SDK for a JVM team, the cost of owning generated scaffolding, and how framework upgrades are tracked across many operators.

## Two layers: library and scaffolding A **controller** is a program that watches Kubernetes objects and keeps making the cluster match what those objects declare. Writing one from scratch with raw client-go means wiring up informers, listers, work queues, rate limiters and leader election by hand. Two projects remove most of that work: - **controller-runtime** is a Go library (`sigs.k8s.io/controller-runtime`). It provides the runtime pieces: a manager, a cache-backed client, a controller builder, event handlers and predicates. - **Kubebuilder** is a CLI that generates a project laid out around controller-runtime, plus the Makefile targets that produce manifests from code. **Operator SDK** reuses the Kubebuilder layout for Go projects and adds packaging on top. ## What the manager gives you `ctrl.NewManager(cfg, ctrl.Options{...})` returns a **Manager**. Everything else hangs off it: - `mgr.GetClient()` returns a **client** that reads from a shared informer cache and writes straight to the API server. - `mgr.GetCache()` returns that shared cache. Two controllers watching Deployments share one informer. - `mgr.GetScheme()` returns the type registry that maps Go structs to group/version/kind. - **Leader election** is controlled by `LeaderElection: true` in the options (the scaffold wires it to a `--leader-elect` flag) and uses a `Lease` object by default. With two replicas of the operator, only the leader runs the controllers. - A **metrics** endpoint and **health/readiness probes** are also served. - `mgr.Start(ctx)` starts the caches, waits for them to sync, starts every controller and blocks until the context is cancelled. ## What Kubebuilder scaffolds | Path | Purpose | |---|---| | `api/v1/gamesession_types.go` | Spec and status structs, with markers | | `internal/controller/gamesession_controller.go` | The reconciler struct, `Reconcile` and `SetupWithManager` | | `cmd/main.go` | Builds the manager, registers controllers, starts it | | `config/crd`, `config/rbac`, `config/manager` | Kustomize bases for the generated YAML | | `internal/controller/suite_test.go` | An envtest-based test harness | | `Makefile` | `manifests`, `generate`, `test`, `run`, `deploy` targets | **Markers** are Go comments that `controller-gen` reads. `// +kubebuilder:rbac:groups=games.example.com,resources=gamesessions,verbs=get;list;watch` becomes a rule in the generated ClusterRole (`manager-role`). `// +kubebuilder:subresource:status` and `// +kubebuilder:validation:Minimum=1` shape the CRD. `make manifests` runs `controller-gen rbac:roleName=manager-role crd webhook ... output:crd:artifacts:config=config/crd/bases`. `make generate` produces the `DeepCopy` methods every API type needs. ## Writing Reconcile The signature is `Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error)`. `req` holds only `NamespacedName`: no event type and no old or new object. A typical body for a multiplayer game-session backend: 1. **Fetch** the `GameSession` named by `req`. If it is not found, return `ctrl.Result{}, client.IgnoreNotFound(err)`. 2. **Compute** the desired children, such as a Deployment of match servers and a Service. 3. **Converge** each child with `controllerutil.CreateOrUpdate`, setting the owner with `controllerutil.SetControllerReference` so that `Owns()` routes the child's events back here. 4. **Report** what you observed through `r.Status().Update` or `Patch`. 5. **Return** `ctrl.Result{}` when done, an error to be retried, or `ctrl.Result{RequeueAfter: d}` to be called again later. ```go // +kubebuilder:rbac:groups=games.example.com,resources=gamesessions,verbs=get;list;watch // +kubebuilder:rbac:groups=games.example.com,resources=gamesessions/status,verbs=get;update;patch // +kubebuilder:rbac:groups=apps,resources=deployments,verbs=get;list;watch;create;update;patch;delete func (r *GameSessionReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) { var gs gamesv1.GameSession if err := r.Get(ctx, req.NamespacedName, &gs); err != nil { return ctrl.Result{}, client.IgnoreNotFound(err) } dep := &appsv1.Deployment{ObjectMeta: metav1.ObjectMeta{Name: gs.Name + "-servers", Namespace: gs.Namespace}} if _, err := controllerutil.CreateOrUpdate(ctx, r.Client, dep, func() error { mutateServers(dep, &gs) return controllerutil.SetControllerReference(&gs, dep, r.Scheme) }); err != nil { return ctrl.Result{}, err } gs.Status.ReadyServers = dep.Status.ReadyReplicas return ctrl.Result{}, r.Status().Update(ctx, &gs) } ``` Idempotency, backoff and finalizers are the reconcile contract itself, and they belong to the reconcile-loop topic. This topic covers the framework that calls `Reconcile`. ## Common first-controller mistakes - **Building private clients or informers** inside the reconciler instead of using the manager's, which duplicates caches and bypasses leader election. - **Writing status with a plain `Update`** when the CRD enables the status subresource. The API server ignores status changes sent that way. - **Forgetting `SetControllerReference`**, so `Owns()` never routes the child's events back and drift on the Deployment goes unnoticed. - **Editing generated files** such as `zz_generated.deepcopy.go` or `config/rbac/role.yaml` by hand. The next `make generate` or `make manifests` overwrites them. - **Blocking in `Reconcile`** with sleeps while waiting for a child to become ready, instead of returning and letting the child's event or a `RequeueAfter` bring the reconciler back. ## The Java Operator SDK equivalent | controller-runtime | Java Operator SDK | |---|---| | `Reconcile(ctx, req)` | `UpdateControl<P> reconcile(P resource, Context<P> context)` | | `ctrl.Result{RequeueAfter: d}` | `UpdateControl...rescheduleAfter(d)` | | `r.Status().Update` | `UpdateControl.patchStatus(resource)` | | builder `For(...)` | `Reconciler<P>` plus `@ControllerConfiguration` | | finalizer handled by hand | implement `Cleaner<P>`; the SDK manages the finalizer | The Java SDK receives the **resource itself** rather than a key, but the model is the same: converge, then report status.

  • You added a `+kubebuilder:rbac` marker for Secrets but the operator logs forbidden errors in the cluster. What did you miss?
    Markers are only comments until `controller-gen` runs. `make manifests` regenerates `config/rbac/role.yaml` (the `manager-role` ClusterRole) and the CRDs under `config/crd/bases`, and the regenerated YAML then has to be deployed. Skip either step and the running ServiceAccount still has the old rules. The same applies to validation markers: the stale CRD in the cluster keeps the old schema.
  • How does the Java Operator SDK map onto this model?
    You implement `Reconciler<P>` and annotate the class with `@ControllerConfiguration`. `reconcile(P resource, Context<P> context)` returns an `UpdateControl` such as `patchStatus(resource)` or `noUpdate()`, optionally with `rescheduleAfter`. Implementing `Cleaner<P>` turns on automatic finalizer handling, and `cleanup` returns a `DeleteControl`. The SDK passes the resource itself instead of a namespace/name key, and it provides its own informer-backed event sources.
  • What happens when you run two replicas of a controller-runtime operator?
    With leader election enabled, both replicas start and compete for a `Lease`. Only the leader starts the controllers, so the other replica stays on standby, ready to take over. Runnables that do not need leader election, such as a webhook server, run on both replicas. If leader election is off, both replicas reconcile the same objects and fight each other.

controller-runtime is a building's plumbing and wiring, and Kubebuilder is the architect's standard floor plan. You still furnish one room, Reconcile, but the water and power already reach it.

saying these in an interview costs you the question

  • Reconcile receives the changed object and the event type as arguments
  • Kubebuilder is a runtime library the operator links against instead of a scaffolding tool
  • RBAC markers grant permissions directly without regenerating and applying manifests
  • Each controller should build its own informers and clients rather than use the manager's
  • controller-runtime works only for custom resources, never for built-in kinds
open as a page

What is a Kubernetes CustomResourceDefinition, and what exactly do you get once you apply one to a cluster?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A CustomResourceDefinition registers a new resource type with the Kubernetes API server. You immediately get REST endpoints for it, storage in etcd, kubectl support, watches, RBAC and labels — but nothing acts on the objects until you run a controller for them.

open as a page

A team wants automated backups, version upgrades and failover for a stateful database running on Kubernetes, all driven through the Kubernetes API. Explain what the Kubernetes Operator pattern is and how it delivers that.

level: juniorimportance: must knowfreq 68%

basics

~20 s

An Operator is a custom Kubernetes resource type plus a controller running in the cluster. Users declare the desired database in a custom object; the controller continuously drives the real system to match that spec, automating install, backup, upgrade and failover.

open as a page

In Kubernetes, a request that creates an object can be intercepted by webhooks registered through both a MutatingWebhookConfiguration and a ValidatingWebhookConfiguration. In what order does the API server run them relative to each other and to the rest of its request handling, and why does that order matter?

level: middleimportance: must knowfreq 60%

basics

~20 s

The API server authenticates and authorizes, then runs mutating admission (webhooks may patch the object), then schema validation, then validating admission (accept or reject only), then writes to etcd. Mutation runs first so validators judge the final stored object.

open as a page

A CustomResourceDefinition in `apiextensions.k8s.io/v1` requires an OpenAPI v3 validation schema. What does that schema do to the objects users submit, what happens to fields the schema does not mention, and how would you deliberately allow arbitrary content in one part of the object?

level: middleimportance: must knowfreq 50%

basics

~20 s

The schema validates types, required fields and constraints at admission, applies declared defaults, and prunes — silently strips — any field it does not describe. To keep arbitrary content, mark that subtree with x-kubernetes-preserve-unknown-fields: true.

open as a page

What changes about a Kubernetes custom resource when you enable `subresources: {status: {}}` on its CustomResourceDefinition, and why do controllers want it?

level: middleimportance: must knowfreq 40%

basics

~20 s

It splits the object into two write paths: writes to the main endpoint ignore status, and writes to /status change only status. It also makes metadata.generation increment only on spec changes, and lets RBAC grant status updates separately.

open as a page

When a Kubernetes CRD switches its storage version to v1, how are objects still stored as v1alpha1 served, and when is conversion strategy None enough?

level: middleimportance: must knowfreq 58%

basics

~20 s

A CRD stores exactly one version, but old objects stay in etcd as last written; the API server converts them in memory whenever a request needs another version. Strategy None only rewrites apiVersion, so it suits identical schemas only.

open as a page

Kubernetes controllers are described as level-triggered rather than edge-triggered, and their reconcile function must be idempotent. Explain what that means and why a controller written as a set of event handlers for create, update and delete tends to break.

level: middleimportance: must knowfreq 56%

basics

~20 s

Level-triggered means the controller reads current desired and actual state on every pass and closes the gap, instead of reacting to individual change events. Reconcile may run many times for one change, in any order, after restarts, so it must be safe to repeat and must not depend on having seen previous events.

open as a page

Your controller creates a Deployment, a Service and a Secret for each custom resource it manages. How do you ensure those objects are cleaned up when the custom resource is deleted, and how does Kubernetes garbage collection decide what to remove?

level: middleimportance: must knowfreq 47%

basics

~20 s

Set an ownerReference on each created object pointing at the custom resource. The garbage collector then deletes dependents automatically when the owner is deleted. Owner and dependent must be in the same namespace, and a cluster-scoped object cannot be owned by a namespaced one.

open as a page

In a Kubernetes custom resource, how do `metadata.generation` and `status.observedGeneration` differ, and why should its controller write the latter?

level: middleimportance: must knowfreq 62%

basics

~20 s

The API server sets metadata.generation and increments it when desired state changes. The controller copies the generation it acted on into status.observedGeneration, so readers know the status is current only when the two numbers match.

open as a page

Each webhook entry in a Kubernetes MutatingWebhookConfiguration or ValidatingWebhookConfiguration has a `failurePolicy` field set to either `Fail` or `Ignore`. Explain what each value does when the webhook backend is unreachable, times out or returns an error, and how you decide which one to use.

level: seniorimportance: must knowfreq 55%

basics

~20 s

Fail means a webhook error, timeout or unreachable backend causes the API request to be rejected; Ignore means the API server logs it and admits the request as if the webhook had approved it. Fail is safe for correctness, Ignore is safe for availability.

open as a page

When would you build a Kubernetes Operator for an application instead of shipping a Helm chart with a StatefulSet, and what does choosing the operator cost you?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Build an operator when the application needs ongoing, application-specific decisions — ordered upgrades, failover, backup and restore, resharding — that no template can make. If install-and-forget is enough, a chart plus StatefulSet is cheaper. The operator costs you a privileged, always-running component to write, secure, upgrade and monitor.

open as a page

Explain what a Kubernetes finalizer is, how a controller should use one to release external resources before an object goes away, and how you would debug a namespace or custom resource stuck in a Terminating state.

level: seniorimportance: must knowfreq 50%

basics

~20 s

A finalizer is a string in metadata.finalizers that blocks final deletion. A delete request only sets deletionTimestamp; the object persists until every finalizer is removed. The controller does its cleanup, then removes its own finalizer. Stuck Terminating usually means the responsible controller is gone or failing.

open as a page

Walk through the machinery that sits between a change in the Kubernetes API and your controller's reconcile function being called: watches, informers, caches and work queues. Why is a work queue used rather than calling reconcile directly from the event handler?

level: seniorimportance: must knowfreq 48%

basics

~20 s

A shared informer watches the API server, keeps an in-memory cache of objects, and on each event pushes only the object's key onto a rate-limited work queue. Workers pop keys and reconcile, reading from the cache. The queue deduplicates keys, bounds concurrency, and provides retry with backoff.

open as a page

In Kubernetes, what are the `status.conditions` entries a custom resource's controller reports, and how do you check them with `kubectl`?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Conditions are status entries a custom resource's controller writes, each with a type such as Ready, a status of True, False or Unknown, a reason and a message. Read them with kubectl describe or jsonpath; block on them with kubectl wait --for=condition=.

open as a page

The `sideEffects` field is mandatory on every webhook entry in a Kubernetes MutatingWebhookConfiguration or ValidatingWebhookConfiguration. What does it declare, which values are allowed, and how does the API server use it when a client sends a request with `--dry-run=server`?

level: middleimportance: should knowfreq 30%

basics

~20 s

sideEffects declares whether the webhook changes state outside the object under review. In admissionregistration.k8s.io/v1 it must be None or NoneOnDryRun. For a server dry-run request, the API server skips None-declared-unsafe webhooks and calls the others with dryRun: true set in the AdmissionReview.

open as a page

Why must a Kubernetes admission webhook be served over TLS, how does the API server decide to trust that server's certificate, and what exactly breaks when the certificate expires?

level: middleimportance: should knowfreq 40%

basics

~20 s

The API server only calls webhooks over HTTPS, and it verifies the server certificate against the caBundle in the webhook configuration, requiring a SAN matching <service>.<namespace>.svc. On expiry the handshake fails, so requests are rejected under failurePolicy: Fail or silently unchecked under Ignore.

open as a page

Why can controller-runtime's mgr.GetClient() return stale or unexpectedly expensive reads, and when should a controller use mgr.GetAPIReader() instead?

level: middleimportance: should knowfreq 46%

basics

~20 s

mgr.GetClient() serves Get and List from informer caches but sends writes to the API server, so a read can lag your own write, and the first read of a new kind starts a cluster-wide informer. mgr.GetAPIReader() reads live and caches nothing.

open as a page

When defining a CustomResourceDefinition you must set `scope` to either `Namespaced` or `Cluster`. What actually differs between the two, and what should drive the choice?

level: middleimportance: should knowfreq 35%

basics

~20 s

Namespaced objects live in a namespace, are named uniquely per namespace, and are covered by namespaced Roles and namespace deletion. Cluster objects are global, unique cluster-wide, and require ClusterRoles. Choose namespaced for tenant-owned resources, cluster for shared cluster-level configuration.

open as a page

Kubernetes operator projects are commonly rated against a five-level capability or maturity model. What do those levels describe, and how would you use them when evaluating a third-party operator before running it in production?

level: middleimportance: should knowfreq 42%

basics

~20 s

The levels run: 1 basic install, 2 seamless upgrades of the managed application, 3 full lifecycle (backup, restore, failover), 4 deep insights (metrics, alerts, logs), 5 autopilot (auto-scaling, auto-tuning, auto-remediation). Use them to check the operator actually covers the operations you would otherwise do by hand.

open as a page

A controller's reconcile fails for one object because a dependency is temporarily unavailable, and the controller then hammers the Kubernetes API with retries. Explain how requeue and rate-limited backoff are supposed to work in a controller, and what the different reconcile return values mean.

level: middleimportance: should knowfreq 36%

basics

~20 s

Returning an error requeues the key through a rate limiter with exponential backoff, so retries slow from milliseconds to minutes. Returning RequeueAfter schedules a re-check at a fixed delay for polling. Returning an empty result with no error drops the item and resets its backoff. Never retry in a loop inside reconcile.

open as a page

A Kubernetes GameSession custom resource names a shared ConfigMap in spec.profileRef; with controller-runtime, how do you reconcile every referencing GameSession when that ConfigMap changes?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Owns cannot help because the shared ConfigMap has no single controlling owner. Add Watches(&corev1.ConfigMap{}, handler.EnqueueRequestsFromMapFunc(fn)), where fn uses a field index on spec.profileRef to list referencing GameSessions and return their requests. Scope the ConfigMap cache with a selector.

open as a page

Your CustomResourceDefinition currently offers version v1alpha1 and you want to introduce v1beta1 without breaking existing objects or existing clients. Explain what the `served` and `storage` flags mean on each version entry, and when you need a conversion webhook.

level: seniorimportance: should knowfreq 35%

basics

~20 s

served: true means clients may read and write that version; exactly one version has storage: true and is the form written to etcd. Objects are converted between versions on the fly — trivially if the schemas are compatible (strategy: None), otherwise you must run a conversion webhook.

open as a page

Why does the Kubernetes API server refuse to remove v1alpha1 from a CRD's spec.versions, and how do you migrate stored objects so the removal succeeds?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The CRD's status.storedVersions still lists v1alpha1, so etcd may hold objects in it. Re-encode every object in the storage version with a StorageVersionMigration or no-op writes, trim storedVersions to v1, then delete the version entry.

open as a page

A Kubernetes CRD's conversion webhook becomes slow or unreachable; what breaks across the cluster, and how would you run that webhook so it cannot stall the API?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Every request needing conversion fails or slows: lists, watches, controllers' informers, writes, namespace deletion and migrations. There is no failurePolicy to skip it, so run the webhook redundantly with a PodDisruptionBudget, pure fast code and rotated TLS.

open as a page

You are about to roll out a cluster-wide admission webhook that intercepts every Pod creation in a production Kubernetes cluster. What latency and availability risks does that introduce, and how would you scope, size and roll it out so that a problem with the webhook cannot take the cluster down?

level: principalimportance: should knowfreq 45%

basics

~20 s

Every matching write now waits on a network call, so webhook latency becomes API latency and webhook downtime can block Pod creation. Scope with rules and selectors, exclude kube-system and the webhook's own namespace, keep timeouts at a few seconds, run it HA, and ship with failurePolicy Ignore before flipping to Fail.

open as a page

By default `kubectl get` on a custom resource shows only NAME and AGE. How do you make it display fields from the object itself, and what is the mechanism behind it?

level: juniorimportance: nice to knowfreq 25%

basics

~20 s

Add additionalPrinterColumns to the CRD version: each entry has a name, a type, and a jsonPath into the object. The API server renders the table server-side, so every client that asks for a table sees the same columns.

open as a page

In a Kubernetes CRD, what does setting deprecated: true on a version entry do for clients, and how is it different from served: false?

level: juniorimportance: nice to knowfreq 27%

basics

~20 s

Setting deprecated: true keeps the CRD version fully working but adds a warning to every response at that version, optionally replaced by deprecationWarning. Setting served: false removes the version's endpoints, so requests get 404 while stored objects stay readable.

open as a page

How does controller-runtime's envtest let you test a Kubernetes controller, and what does an envtest environment deliberately not run?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

envtest starts real etcd and kube-apiserver binaries locally, installs your CRDs and returns a rest.Config for your manager. It runs no kube-controller-manager, scheduler or kubelet, so there is no garbage collection, no Pods running and no finished namespace deletion.

open as a page

Your team wants `kubectl scale` and a Kubernetes HorizontalPodAutoscaler to drive a custom resource. What must its CustomResourceDefinition's scale subresource declare, and what must the controller keep populated?

level: seniorimportance: nice to knowfreq 30%

basics

~10 s

The CRD declares subresources.scale with specReplicasPath, statusReplicasPath and, for an HPA, labelSelectorPath. The controller applies the desired replicas and writes the actual replica count and a string label selector for the Pods it owns.

open as a page

showing 1–30 of 31