What is a Kubernetes CustomResourceDefinition, and what exactly do you get once you apply one to a cluster?
answer
- CRD = new kind registered with the API server
- metadata.name must be <plural>.<group>
- free: REST, etcd, watch, RBAC, kubectl, ownerReferences
- no controller = objects just sit there
- deleting the CRD deletes all its objects
basics
~20 sA CustomResourceDefinition registers a new resource type with the Kubernetes API server. You immediately get REST endpoints for it, storage in etcd, kubectl support, watches, RBAC and labels — but nothing acts on the objects until you run a controller for them.
solid answer
~50 sA CRD is an object in `apiextensions.k8s.io/v1` that tells the API server about a new **kind**. You declare a group, one or more versions, a plural/singular name, a kind, a scope (`Namespaced` or `Cluster`) and an OpenAPI v3 schema. The API server then serves `/apis/<group>/<version>/namespaces/<ns>/<plural>` like any built-in resource. What you get for free is the whole API machinery: create/read/update/delete/patch, list and **watch**, persistence in etcd, schema validation, labels, annotations, `ownerReferences` and garbage collection, RBAC on the new resource name, `kubectl get/describe/edit`, and inclusion in `kubectl api-resources`. What you do **not** get is behavior. A `Database` custom resource sits in etcd doing nothing until a controller watches it and reconciles reality toward it. The CRD is the schema and the endpoint; the controller is the meaning. CRDs are the declarative extension point behind operators, and unlike aggregated API servers they require no extra process to serve the API.
code
yaml · 30 linesapiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: databases.example.com # must be <plural>.<group>
spec:
group: example.com
scope: Namespaced
names:
plural: databases
singular: database
kind: Database
shortNames: ["db"]
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
required: ["engine", "sizeGb"]
properties:
engine:
type: string
enum: ["postgres", "mysql"]
sizeGb:
type: integer
minimum: 1go deeper
Define a CRD as registering a new resource kind, and list what you get: kubectl access, storage, RBAC and watches — but no behavior.
Add the anatomy — group, names, scope, versions, required schema — and the fact that deleting the CRD deletes all its objects.
Discuss the CRD-versus-aggregated-API-server trade, etcd pressure from high-churn custom resources, and why CRD installation is a privileged platform-team operation.
Treat the CRD as the published API contract for a platform capability: naming and group ownership, who may install and version them, how they interact with GitOps pruning and cluster lifecycle, and when a domain deserves an API at all.
## The idea Kubernetes ships with resources like Pod, Deployment and Service. A **CustomResourceDefinition** (CRD) lets you add your own — `Database`, `Certificate`, `KafkaTopic` — so that your domain object is stored and served by the same API server, with the same tooling, as the built-ins. A CRD is itself a Kubernetes object, in the group `apiextensions.k8s.io`, version `v1`. Applying it causes the API server to start serving a new endpoint within seconds; deleting it removes the endpoint **and deletes every object of that kind**, which is why CRD deletion is a genuinely destructive operation. ## Anatomy of a CRD The required pieces are: - **`spec.group`** — the API group, e.g. `example.com`. Use a domain you control to avoid collisions. - **`spec.names`** — `plural` (used in the URL), `singular`, `kind` (the CamelCase type name used in YAML), optional `shortNames` for kubectl, and `categories` (e.g. adding your kind to `kubectl get all`). - **`metadata.name`** — must be exactly `<plural>.<group>`; the API server rejects anything else. - **`spec.scope`** — `Namespaced` or `Cluster`. - **`spec.versions[]`** — each with a name, `served`, `storage`, and a `schema.openAPIV3Schema` (required in `v1`). ## What the API server gives you Once registered, the custom resource is a first-class API citizen: - **Full REST verbs**: create, get, list, watch, update, patch, delete, deletecollection — served at `/apis/<group>/<version>/…`, namespaced or cluster-scoped as declared. - **Persistence**: objects are stored in etcd, with resource versions and optimistic concurrency, so `kubectl apply` conflict semantics work exactly as for built-ins. - **Schema validation and pruning** against the declared OpenAPI schema, so bad YAML is rejected at submit time rather than by a controller later. - **Watch and informers**: controllers can watch for changes efficiently — this is what makes the controller pattern possible at all. - **Metadata machinery**: labels, annotations, finalizers, `ownerReferences` (so your objects participate in cascading deletion), `creationTimestamp`, and the standard `metadata.generation`. - **RBAC**: the resource plural becomes a name you can grant in a Role or ClusterRole, so access control is the same mechanism as everything else. - **Tooling**: `kubectl get`, `describe`, `edit`, `explain` (from the schema), server-side apply, `kubectl api-resources`, and anything built on the discovery API — dashboards, GitOps tools, policy engines. ## What it does not give you A CRD is data, not behavior. Creating a `Database` object does not create a database. Something must **watch** those objects and act: that is a controller, and a CRD plus its controller packaged together is what people call an operator. Those are separate topics; the point here is the separation itself. It is entirely legitimate — and common — to use a CRD with no controller at all, purely as structured, RBAC-controlled, kubectl-accessible configuration storage that other tools read. The API server also does not give you cross-object logic: no foreign keys, no uniqueness constraints beyond the object name, no transactions across objects. Validation is per object (schema plus, since 1.29, CEL rules within the object). Anything relational belongs in the controller. ## CRDs versus aggregated API servers The other extension mechanism is an **aggregated API server**: you run your own API server process and register it with an `APIService` so the main API server proxies to it. That gives you full control — custom storage backends, arbitrary subresources, custom protobuf — at the cost of running and operating another highly available server. CRDs need no extra process, store in the cluster's etcd, and cover the overwhelming majority of cases. The usual reasons to reach for aggregation are: you cannot store the data in etcd (too large, too high write rate), you need non-standard subresources or verbs, or you need to serve data that is not really stored objects at all (like `metrics.k8s.io`). Start with a CRD. ## Practical notes CRD objects are cluster-scoped even when the resources they define are namespaced, so installing one requires cluster-level rights — which is why CRD installation is typically a platform-team action separate from application deployment. Also, because every custom object lives in the cluster's etcd, high-churn or high-cardinality custom resources put real pressure on the control plane; CRDs are for configuration-shaped data, not for telemetry or event streams.
- You applied the CRD and created a custom object, but nothing happened in the outside world. Why?Because a CRD only defines and stores the resource; it carries no behavior. The object is now durably in etcd and visible to kubectl and RBAC, but until a controller watches that kind and reconciles real infrastructure toward the declared spec, creating the object changes nothing. The CRD is the API, the controller is the implementation.
- When would you choose an aggregated API server over a CRD?When the data cannot live comfortably in etcd — very large objects or very high write rates — or when you need capabilities CRDs do not offer, such as arbitrary custom subresources and verbs, a different storage backend, or serving computed rather than stored data, as metrics.k8s.io does. Otherwise a CRD is preferred because it requires no additional server to run and keep available.
- What happens if you delete the CustomResourceDefinition itself?The API endpoint disappears and every object of that kind is deleted from etcd, with finalizers running as usual. That makes CRD deletion far more destructive than it looks in a manifest diff, so uninstall flows and GitOps prunes need to treat CRDs carefully — often by excluding them from automatic pruning.
A CRD is like adding a new table plus its REST endpoints to a shared system: you get storage, queries and permissions immediately, but nothing happens to the outside world until you also write the service that reads the rows and acts on them.
saying these in an interview costs you the question
- Believing that creating a CRD makes something happen without a controller
- Naming the CRD object something other than <plural>.<group>
- Thinking CRDs need a separate API server process to be served
- Assuming deleting the CRD leaves existing custom objects intact
- Using CRDs as a general-purpose datastore for high-churn or high-volume data