A CustomResourceDefinition in `apiextensions.k8s.io/v1` requires an OpenAPI v3 validation schema. What does that schema do to the objects users submit, what happens to fields the schema does not mention, and how would you deliberately allow arbitrary content in one part of the object?
answer
- v1: schema required and must be structural
- unknown fields are pruned silently, not rejected
- default applies on write and on read
- x-kubernetes-preserve-unknown-fields for opaque subtrees
- x-kubernetes-validations = CEL, self / oldSelf for immutability
basics
~20 sThe schema validates types, required fields and constraints at admission, applies declared defaults, and prunes — silently strips — any field it does not describe. To keep arbitrary content, mark that subtree with x-kubernetes-preserve-unknown-fields: true.
solid answer
~50 sIn `v1` the schema is mandatory and must be **structural**: every field is typed, `type: object` nodes list their `properties`, and constructs like `allOf`/`oneOf` cannot introduce new fields. Three behaviors follow: 1. **Validation** — types, `required`, `enum`, `minimum`, `maxLength`, `pattern` and so on are enforced at admission, so bad objects are rejected by the API server rather than by your controller later. 2. **Pruning** — anything not described by the schema is dropped silently on write. A typo like `replcias` disappears rather than erroring, which is the number-one CRD support ticket. Server-side apply and `kubectl apply --validate` help surface it. 3. **Defaulting** — a `default` on a property fills it in when absent, on write and on read from etcd. To carry opaque content (an embedded pod template fragment, arbitrary user config) set `x-kubernetes-preserve-unknown-fields: true` on that subtree only. For cross-field rules, `x-kubernetes-validations` runs CEL expressions in the API server — for example enforcing immutability with `self == oldSelf`.
code
yaml · 30 linesschema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
required: ["engine"]
x-kubernetes-validations:
- rule: "self.maxReplicas >= self.minReplicas"
message: "maxReplicas must be >= minReplicas"
properties:
engine:
type: string
description: Database engine to provision.
enum: ["postgres", "mysql"]
x-kubernetes-validations:
- rule: "self == oldSelf"
message: "engine is immutable"
minReplicas:
type: integer
minimum: 1
default: 1
maxReplicas:
type: integer
minimum: 1
default: 3
engineConfig:
type: object
description: Opaque engine settings passed through untouched.
x-kubernetes-preserve-unknown-fields: truego deeper
Say the schema types and validates the object and that unknown fields are removed rather than rejected.
Explain structural schemas, the pruning/defaulting/validation triad, and the preserve-unknown-fields escape hatch with correct scoping.
Add CEL x-kubernetes-validations for cross-field and immutability rules, strict field validation to surface pruning, and the argument for schema-level validation over admission webhooks.
Treat the schema as the versioned public contract of the platform API: what belongs in the schema versus the controller, how validation tightening interacts with existing objects and ratcheting, and the generation pipeline that keeps schema and code in sync.
## The schema is mandatory in v1 Under the removed `apiextensions.k8s.io/v1beta1` you could ship a CRD with no schema and the API server would accept any YAML. `v1` requires `schema.openAPIV3Schema` on every version and requires it to be **structural**, which means roughly: - every field specifies a `type` (or `x-kubernetes-preserve-unknown-fields`/`x-kubernetes-int-or-string`); - objects declare their `properties` (or `additionalProperties` for map-shaped data, but not both); - `allOf`, `anyOf`, `oneOf` and `not` may constrain fields but may not *introduce* fields not present in the main schema; - the root is `type: object`. Structural schemas are what make pruning, defaulting, server-side apply and CEL validation possible, so this is not bureaucratic — it is the enabling constraint. ## Behavior 1: validation Standard OpenAPI v3 keywords are enforced at admission time by the API server itself, before any controller sees the object: `required`, `type`, `enum`, `minimum`/`maximum`, `minLength`/`maxLength`, `pattern`, `minItems`/`maxItems`, `uniqueItems`, `format`. A rejected object never reaches etcd, and the user gets an immediate, precise error at `kubectl apply` time. This is the single biggest quality win of a well-authored CRD: every constraint you express in the schema is a class of bug your controller never has to defend against, and every constraint you omit is one it must. ## Behavior 2: pruning (the surprising one) On every write, the API server **removes** fields that the schema does not describe — silently, with no error. Someone writing `spec.replcias: 5` sees their object created successfully with no replicas field, and their controller uses the default. This trips up almost everyone once. Mitigations: - `kubectl apply --validate=true` (the default in modern kubectl, which uses server-side field validation) reports unknown fields as errors or warnings. - The API server's field validation levels — `Strict`, `Warn`, `Ignore` — are selectable via the `fieldValidation` query parameter; `kubectl --validate=strict` rejects unknown fields outright. - Server-side apply tracks field ownership and surfaces conflicts, making silent loss more visible. Note that pruning applies to **stored** objects too: after a CRD update that removes a property, the field is stripped the next time the object is written. ## Behavior 3: defaulting A `default` value on a property is applied when the field is absent — when the object is created or updated, and also when it is read back from etcd (so tightening a default changes existing objects' observed values without a write). Defaults must validate against their own schema. Combined with pruning, this means the object your controller reads is always fully typed and fully defaulted, which is a meaningfully nicer contract than parsing free-form YAML. Defaulting happens as part of decoding, so it occurs before mutating admission webhooks run. ## Allowing arbitrary content on purpose Sometimes a field genuinely must carry unknown structure: an embedded third-party config blob, or a passthrough of a built-in type you do not want to vendor. Two escape hatches: - **`x-kubernetes-preserve-unknown-fields: true`** on the subtree stops pruning below that point. Apply it to the narrowest node possible — putting it at the root gives up validation for the whole object. - **`x-kubernetes-embedded-resource: true`** marks a node as a full Kubernetes object (with apiVersion/kind/metadata), used for embedded templates; it implies preservation of the embedded object's fields. There is also **`x-kubernetes-int-or-string: true`** for fields that accept either form, mirroring built-in types like `IntOrString`. ## Beyond per-field checks: CEL validation rules OpenAPI cannot express "maxReplicas must be at least minReplicas" or "this field is immutable". Since Kubernetes 1.25 (beta) and 1.29 (GA), `x-kubernetes-validations` attaches **CEL** expressions to a node: ```yaml x-kubernetes-validations: - rule: "self.maxReplicas >= self.minReplicas" message: "maxReplicas must be >= minReplicas" - rule: "self == oldSelf" message: "engine is immutable" ``` `self` is the current value at that node, `oldSelf` the previous value on updates (available on rules marked for transition). Because these run inside the API server, they need no webhook, no TLS and no availability story — a strong reason to push validation into the CRD rather than into an admission webhook whenever the rule can be expressed this way. Kubernetes 1.30 added **validation ratcheting**, so tightening a CRD's rules does not immediately break updates to existing non-conforming objects as long as the offending field is unchanged. ## Practical authoring guidance Write the schema as the API contract, not as an afterthought: mark `required` honestly, use `enum` for closed sets, set sensible `default`s, add `description` on every field (it powers `kubectl explain`), and express invariants in CEL. Keep `x-kubernetes-preserve-unknown-fields` scoped to the one node that needs it. If you generate CRDs from Go types (controller-gen), the markers map directly onto these keywords, which keeps schema and controller code from drifting apart.
- A user reports that a field they set on a custom resource keeps disappearing. What is your first hypothesis?That the field is not described by the CRD's OpenAPI schema, so the API server prunes it silently on write — usually a typo or a field added to the controller but never to the schema. Confirm by re-applying with strict field validation, which turns the unknown field into an explicit error, and fix by adding the property to the schema.
- How do you make one field of a custom resource immutable after creation without writing an admission webhook?Attach a CEL rule with x-kubernetes-validations on that property using the expression self == oldSelf and a clear message. The API server evaluates it on update, rejecting any change, with no webhook, certificate or availability concern. This is generally preferable to a validating webhook whenever the rule can be expressed over the object itself.
- What is the risk of putting x-kubernetes-preserve-unknown-fields: true at the root of the schema?You disable pruning for the whole object, so typos are stored rather than caught, defaulting and server-side apply behave less predictably, and you lose the structural guarantees that make kubectl explain and CEL validation useful. Scope the marker to the single subtree that genuinely holds opaque data.
saying these in an interview costs you the question
- Expecting unknown fields to cause an error rather than being silently pruned
- Thinking the schema is optional in apiextensions.k8s.io/v1
- Setting x-kubernetes-preserve-unknown-fields at the root to make errors go away
- Believing OpenAPI keywords can express cross-field rules or immutability without CEL
- Assuming defaults are applied only on create, not on reads from etcd