skip to content

CustomResourceDefinitions

A CRD registers a new resource type whose structural OpenAPI schema prunes unknown fields, applies defaults and enforces CEL rules, so kubectl, RBAC and watch work for it at once. Interviewers ask what a CRD gives you free and what still needs a controller.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

6

What is a Kubernetes CustomResourceDefinition, and what exactly do you get once you apply one to a cluster?

level: juniorimportance: must knowfreq 65%

answer

  1. CRD = new kind registered with the API server
  2. metadata.name must be <plural>.<group>
  3. free: REST, etcd, watch, RBAC, kubectl, ownerReferences
  4. no controller = objects just sit there
  5. deleting the CRD deletes all its objects

basics

~20 s

A CustomResourceDefinition registers a new resource type with the Kubernetes API server. You immediately get REST endpoints for it, storage in etcd, kubectl support, watches, RBAC and labels — but nothing acts on the objects until you run a controller for them.

solid answer

~50 s

A CRD is an object in `apiextensions.k8s.io/v1` that tells the API server about a new **kind**. You declare a group, one or more versions, a plural/singular name, a kind, a scope (`Namespaced` or `Cluster`) and an OpenAPI v3 schema. The API server then serves `/apis/<group>/<version>/namespaces/<ns>/<plural>` like any built-in resource. What you get for free is the whole API machinery: create/read/update/delete/patch, list and **watch**, persistence in etcd, schema validation, labels, annotations, `ownerReferences` and garbage collection, RBAC on the new resource name, `kubectl get/describe/edit`, and inclusion in `kubectl api-resources`. What you do **not** get is behavior. A `Database` custom resource sits in etcd doing nothing until a controller watches it and reconciles reality toward it. The CRD is the schema and the endpoint; the controller is the meaning. CRDs are the declarative extension point behind operators, and unlike aggregated API servers they require no extra process to serve the API.

code

yaml · 30 lines
yaml
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
  name: databases.example.com   # must be <plural>.<group>
spec:
  group: example.com
  scope: Namespaced
  names:
    plural: databases
    singular: database
    kind: Database
    shortNames: ["db"]
  versions:
    - name: v1
      served: true
      storage: true
      schema:
        openAPIV3Schema:
          type: object
          properties:
            spec:
              type: object
              required: ["engine", "sizeGb"]
              properties:
                engine:
                  type: string
                  enum: ["postgres", "mysql"]
                sizeGb:
                  type: integer
                  minimum: 1

go deeper

for a junior

Define a CRD as registering a new resource kind, and list what you get: kubectl access, storage, RBAC and watches — but no behavior.

for a middle

Add the anatomy — group, names, scope, versions, required schema — and the fact that deleting the CRD deletes all its objects.

for a senior

Discuss the CRD-versus-aggregated-API-server trade, etcd pressure from high-churn custom resources, and why CRD installation is a privileged platform-team operation.

for a principal

Treat the CRD as the published API contract for a platform capability: naming and group ownership, who may install and version them, how they interact with GitOps pruning and cluster lifecycle, and when a domain deserves an API at all.

## The idea Kubernetes ships with resources like Pod, Deployment and Service. A **CustomResourceDefinition** (CRD) lets you add your own — `Database`, `Certificate`, `KafkaTopic` — so that your domain object is stored and served by the same API server, with the same tooling, as the built-ins. A CRD is itself a Kubernetes object, in the group `apiextensions.k8s.io`, version `v1`. Applying it causes the API server to start serving a new endpoint within seconds; deleting it removes the endpoint **and deletes every object of that kind**, which is why CRD deletion is a genuinely destructive operation. ## Anatomy of a CRD The required pieces are: - **`spec.group`** — the API group, e.g. `example.com`. Use a domain you control to avoid collisions. - **`spec.names`** — `plural` (used in the URL), `singular`, `kind` (the CamelCase type name used in YAML), optional `shortNames` for kubectl, and `categories` (e.g. adding your kind to `kubectl get all`). - **`metadata.name`** — must be exactly `<plural>.<group>`; the API server rejects anything else. - **`spec.scope`** — `Namespaced` or `Cluster`. - **`spec.versions[]`** — each with a name, `served`, `storage`, and a `schema.openAPIV3Schema` (required in `v1`). ## What the API server gives you Once registered, the custom resource is a first-class API citizen: - **Full REST verbs**: create, get, list, watch, update, patch, delete, deletecollection — served at `/apis/<group>/<version>/…`, namespaced or cluster-scoped as declared. - **Persistence**: objects are stored in etcd, with resource versions and optimistic concurrency, so `kubectl apply` conflict semantics work exactly as for built-ins. - **Schema validation and pruning** against the declared OpenAPI schema, so bad YAML is rejected at submit time rather than by a controller later. - **Watch and informers**: controllers can watch for changes efficiently — this is what makes the controller pattern possible at all. - **Metadata machinery**: labels, annotations, finalizers, `ownerReferences` (so your objects participate in cascading deletion), `creationTimestamp`, and the standard `metadata.generation`. - **RBAC**: the resource plural becomes a name you can grant in a Role or ClusterRole, so access control is the same mechanism as everything else. - **Tooling**: `kubectl get`, `describe`, `edit`, `explain` (from the schema), server-side apply, `kubectl api-resources`, and anything built on the discovery API — dashboards, GitOps tools, policy engines. ## What it does not give you A CRD is data, not behavior. Creating a `Database` object does not create a database. Something must **watch** those objects and act: that is a controller, and a CRD plus its controller packaged together is what people call an operator. Those are separate topics; the point here is the separation itself. It is entirely legitimate — and common — to use a CRD with no controller at all, purely as structured, RBAC-controlled, kubectl-accessible configuration storage that other tools read. The API server also does not give you cross-object logic: no foreign keys, no uniqueness constraints beyond the object name, no transactions across objects. Validation is per object (schema plus, since 1.29, CEL rules within the object). Anything relational belongs in the controller. ## CRDs versus aggregated API servers The other extension mechanism is an **aggregated API server**: you run your own API server process and register it with an `APIService` so the main API server proxies to it. That gives you full control — custom storage backends, arbitrary subresources, custom protobuf — at the cost of running and operating another highly available server. CRDs need no extra process, store in the cluster's etcd, and cover the overwhelming majority of cases. The usual reasons to reach for aggregation are: you cannot store the data in etcd (too large, too high write rate), you need non-standard subresources or verbs, or you need to serve data that is not really stored objects at all (like `metrics.k8s.io`). Start with a CRD. ## Practical notes CRD objects are cluster-scoped even when the resources they define are namespaced, so installing one requires cluster-level rights — which is why CRD installation is typically a platform-team action separate from application deployment. Also, because every custom object lives in the cluster's etcd, high-churn or high-cardinality custom resources put real pressure on the control plane; CRDs are for configuration-shaped data, not for telemetry or event streams.

  • You applied the CRD and created a custom object, but nothing happened in the outside world. Why?
    Because a CRD only defines and stores the resource; it carries no behavior. The object is now durably in etcd and visible to kubectl and RBAC, but until a controller watches that kind and reconciles real infrastructure toward the declared spec, creating the object changes nothing. The CRD is the API, the controller is the implementation.
  • When would you choose an aggregated API server over a CRD?
    When the data cannot live comfortably in etcd — very large objects or very high write rates — or when you need capabilities CRDs do not offer, such as arbitrary custom subresources and verbs, a different storage backend, or serving computed rather than stored data, as metrics.k8s.io does. Otherwise a CRD is preferred because it requires no additional server to run and keep available.
  • What happens if you delete the CustomResourceDefinition itself?
    The API endpoint disappears and every object of that kind is deleted from etcd, with finalizers running as usual. That makes CRD deletion far more destructive than it looks in a manifest diff, so uninstall flows and GitOps prunes need to treat CRDs carefully — often by excluding them from automatic pruning.

A CRD is like adding a new table plus its REST endpoints to a shared system: you get storage, queries and permissions immediately, but nothing happens to the outside world until you also write the service that reads the rows and acts on them.

saying these in an interview costs you the question

  • Believing that creating a CRD makes something happen without a controller
  • Naming the CRD object something other than <plural>.<group>
  • Thinking CRDs need a separate API server process to be served
  • Assuming deleting the CRD leaves existing custom objects intact
  • Using CRDs as a general-purpose datastore for high-churn or high-volume data

context

open as a page

A CustomResourceDefinition in `apiextensions.k8s.io/v1` requires an OpenAPI v3 validation schema. What does that schema do to the objects users submit, what happens to fields the schema does not mention, and how would you deliberately allow arbitrary content in one part of the object?

level: middleimportance: must knowfreq 50%

basics

~20 s

The schema validates types, required fields and constraints at admission, applies declared defaults, and prunes — silently strips — any field it does not describe. To keep arbitrary content, mark that subtree with x-kubernetes-preserve-unknown-fields: true.

open as a page

What changes about a Kubernetes custom resource when you enable `subresources: {status: {}}` on its CustomResourceDefinition, and why do controllers want it?

level: middleimportance: must knowfreq 40%

basics

~20 s

It splits the object into two write paths: writes to the main endpoint ignore status, and writes to /status change only status. It also makes metadata.generation increment only on spec changes, and lets RBAC grant status updates separately.

open as a page

When defining a CustomResourceDefinition you must set `scope` to either `Namespaced` or `Cluster`. What actually differs between the two, and what should drive the choice?

level: middleimportance: should knowfreq 35%

basics

~20 s

Namespaced objects live in a namespace, are named uniquely per namespace, and are covered by namespaced Roles and namespace deletion. Cluster objects are global, unique cluster-wide, and require ClusterRoles. Choose namespaced for tenant-owned resources, cluster for shared cluster-level configuration.

open as a page

Your CustomResourceDefinition currently offers version v1alpha1 and you want to introduce v1beta1 without breaking existing objects or existing clients. Explain what the `served` and `storage` flags mean on each version entry, and when you need a conversion webhook.

level: seniorimportance: should knowfreq 35%

basics

~20 s

served: true means clients may read and write that version; exactly one version has storage: true and is the form written to etcd. Objects are converted between versions on the fly — trivially if the schemas are compatible (strategy: None), otherwise you must run a conversion webhook.

open as a page

By default `kubectl get` on a custom resource shows only NAME and AGE. How do you make it display fields from the object itself, and what is the mechanism behind it?

level: juniorimportance: nice to knowfreq 25%

basics

~20 s

Add additionalPrinterColumns to the CRD version: each entry has a name, a type, and a jsonPath into the object. The API server renders the table server-side, so every client that asks for a table sees the same columns.

open as a page