In a Kubernetes custom resource, how do `metadata.generation` and `status.observedGeneration` differ, and why should its controller write the latter?
answer
- one counter per owner
- server bumps, controller echoes
- metadata edits don't count
- status writes without subresource do
- equal numbers mean current status
basics
~20 sThe API server sets metadata.generation and increments it when desired state changes. The controller copies the generation it acted on into status.observedGeneration, so readers know the status is current only when the two numbers match.
solid answer
~40 s`metadata.generation` belongs to the API server. It starts at 1 and increments on any change outside `metadata`. With the status subresource enabled, status writes go through `/status` and leave it alone, so in practice it counts spec changes. Labels, annotations and finalizers never bump it. `status.observedGeneration` belongs to the controller: after a reconcile it writes the generation it read at the start. A reader then has a simple test. If `observedGeneration == generation`, the status and Conditions describe the current spec. If it is lower, the controller has not caught up and `Ready=True` may describe the old spec. `resourceVersion` cannot do this job, because it changes on every write, including status writes. The controller writes status with `Status().Update` or a patch on `/status`, so its Role needs `update`/`patch` on the `<plural>/status` resource.
code
go · 11 linesrun.Status.ObservedGeneration = run.Generation
meta.SetStatusCondition(&run.Status.Conditions, metav1.Condition{
Type: "Ready",
Status: metav1.ConditionTrue,
Reason: "LedgerBalanced",
Message: "7 of 7 shards reconciled",
ObservedGeneration: run.Generation,
})
if err := r.Status().Update(ctx, &run); err != nil {
return ctrl.Result{}, err
}go deeper
Remember that generation belongs to the API server and observedGeneration to the controller, and that matching numbers mean the status is current.
Explain exactly what bumps generation for a custom resource, why status writes do not once the subresource is on, and how the controller writes through /status.
Diagnose stale-status and self-triggering reconcile loops by reading these counters. Enforce separate RBAC on the status resource and wait on observedGeneration in pipelines.
Make observedGeneration plus Conditions a required status contract for every in-house operator, so delivery tooling can judge convergence the same way everywhere.
## Two counters, two owners Every Kubernetes object carries `metadata.generation`, and most well-behaved kinds also carry `status.observedGeneration`. They look alike, but different parties write them and they answer different questions. | Field | Written by | Changes when | Answers | |---|---|---|---| | `metadata.generation` | kube-apiserver | desired state changes | "which version of the request is this?" | | `status.observedGeneration` | the object's controller | the controller finishes acting on a generation | "which request does this status describe?" | | `metadata.resourceVersion` | kube-apiserver | **any** write, including status and labels | "which stored revision is this?" (used for optimistic concurrency) | ## How the API server maintains generation for custom resources For custom resources the rule in the apiextensions strategy is short: - On create, `generation` is set to **1**. - On update, the server compares everything **except `metadata`** between old and new. If anything differs, it increments `generation`. - If the CRD enables the **status subresource**, the main endpoint throws away changes to `status`, and the `/status` endpoint keeps everything from the stored object except `status`, including the old `generation`. As a result, only spec changes (and any other top-level non-status field) bump the counter. - Changes to **labels, annotations, finalizers or owner references** are metadata and never bump it. The last point of the status subresource matters most. **Without** the status subresource, a controller writes status through the main endpoint, which counts as a non-metadata change, so every status write bumps `generation`. The controller then sees a "new" generation, reconciles, writes status again and bumps it again. Generation-based change detection becomes useless, and the controller can loop forever on its own writes. That is one of the strongest reasons to enable the subresource. ## What the controller does Inside `Reconcile`, the controller should: 1. Read the object and remember `obj.Generation` as the generation it is acting on. 2. Drive the world toward that spec, for example by scaling a 7-replica Deployment of ledger workers. 3. Set `status.observedGeneration` to the generation from step 1, **not** a value re-read later. If a user edited the spec mid-reconcile, the status must not claim to cover the newer spec. 4. Set each Condition's own `observedGeneration` to the same value. 5. Write through the status subresource: `r.Status().Update(ctx, obj)` or `r.Status().Patch(...)` in controller-runtime, or `kubectl patch --subresource=status` by hand when debugging. ```go run.Status.ObservedGeneration = run.Generation meta.SetStatusCondition(&run.Status.Conditions, metav1.Condition{ Type: "Ready", Status: metav1.ConditionTrue, Reason: "LedgerBalanced", Message: "7 of 7 shards reconciled", ObservedGeneration: run.Generation, }) if err := r.Status().Update(ctx, &run); err != nil { return ctrl.Result{}, err } ``` **RBAC is separate.** `/status` is its own RBAC resource, so a rule that lists only `ledgerruns` does not cover `ledgerruns/status`. The controller's Role needs `get`, `update` and `patch` on `ledgerruns/status`. Users can get `update` on `ledgerruns` without it, so they cannot forge status. ## What readers do with it Consider a `LedgerRun` for the nightly ledger-reconciliation batch on a 3-control-plane, 27-worker cluster. At 01:58 an engineer changes `spec.shards` from 7 to 9. The object now reads: - `metadata.generation: 5` - `status.observedGeneration: 4` - `Ready=True`, reason `LedgerBalanced` A naive pipeline sees `Ready=True` and moves on, but that `True` describes the 7-shard spec from generation 4. A careful reader waits for the generations to match first: ```bash kubectl wait ledgerrun/nightly --for=jsonpath='{.status.observedGeneration}'=5 --timeout=10m kubectl wait ledgerrun/nightly --for=condition=Ready --timeout=45m ``` The same idea sits behind rollout checks on built-in workloads and behind health checks in delivery tooling: "done" means the controller has reported on the spec you actually submitted. ## Where the counters come from on screen `kubectl get ledgerrun nightly -o jsonpath='{.metadata.generation} {.status.observedGeneration}'` prints both numbers side by side. A printer column that shows both makes stale status visible in ordinary `kubectl get` output, which helps an on-call engineer at 02:10 far more than a green Ready column alone. ## Common mistakes - Using `resourceVersion` as the change signal. The controller's own status write changes it, so the controller wakes itself up. - Writing `observedGeneration` before the work is done. The object then claims to be current while it is still converging. - Forgetting `observedGeneration` completely, which leaves readers no way to spot stale status. - Declaring `observedGeneration` in the Go type but not in the CRD's structural schema. The server prunes unknown fields, so the value silently disappears. Declaring the field is the schema's job, but the symptom shows up here.
- What RBAC does the operator need to write status, and why keep it apart from users' permissions?`/status` is a separate RBAC resource: a rule listing only `ledgerruns` does not cover `ledgerruns/status`. The operator's Role gets `get`, `update` and `patch` on `ledgerruns/status`. Users get create and update on `ledgerruns` but nothing on `/status`, so they cannot mark a run `Ready` by hand. Together with the endpoint split, this means each side can write only its own half of the object.
- Why record the generation read at the start of the reconcile instead of re-reading it before the status write?A user may change the spec while the reconcile runs. If the controller re-reads and records the newer generation, the status claims to describe a spec the controller never acted on, and readers proceed too early. Recording the generation it worked from keeps the claim honest. The spec change also triggers another reconcile, which will catch up.
- Does adding a finalizer or a label bump `metadata.generation` on a custom resource?No. Generation ignores every change under `metadata`, including labels, annotations, finalizers and owner references. That is intentional: a controller adding its own finalizer must not look like a new spec to itself. `resourceVersion` does change on those writes, which is one reason it is the wrong change signal.
saying these in an interview costs you the question
- Using resourceVersion to decide whether the spec changed
- Believing labels or finalizer edits increment metadata.generation
- Writing observedGeneration before the controller has finished converging
- Assuming the /status endpoint also accepts spec changes
- Thinking RBAC on the main resource also grants the status subresource
- Trusting Ready=True without comparing observedGeneration to generation