skip to content

Why does the Kubernetes API server refuse to remove v1alpha1 from a CRD's spec.versions, and how do you migrate stored objects so the removal succeeds?

level: seniorimportance: should knowfreq 42%

answer

  1. a ledger of past storage versions
  2. API server only appends
  3. rewrite, do not just read
  4. storagemigration.k8s.io resource object
  5. patch the status subresource last

basics

~20 s

The CRD's status.storedVersions still lists v1alpha1, so etcd may hold objects in it. Re-encode every object in the storage version with a StorageVersionMigration or no-op writes, trim storedVersions to v1, then delete the version entry.

solid answer

~40 s

`status.storedVersions` lists every version that has ever been the CRD's storage version, and validation refuses to drop a listed version from `spec.versions` because objects encoded in it may still sit in etcd and could no longer be decoded. The apiextensions API server only ever appends to that list; clearing an entry is the migrator's job. Migration means writing every object again so it is re-encoded in the storage version. In v1.37 you create a `StorageVersionMigration` in `storagemigration.k8s.io/v1` naming the group and resource; kube-controller-manager sends a no-op patch to each object and, on success, sets `storedVersions` to the storage version alone. Without that controller, read and `kubectl replace` every object, then patch the CRD's `status` subresource yourself. Only then remove the old entry, ideally after a release with `served: false`.

code

yaml · 8 lines
yaml
apiVersion: storagemigration.k8s.io/v1
kind: StorageVersionMigration
metadata:
  name: suggestindexes-to-v1
spec:
  resource:
    group: autocomplete.example.com
    resource: suggestindexes

go deeper

for a junior

Recall that a CRD remembers every version it has ever stored in status.storedVersions, and that a version on that list cannot be deleted.

for a middle

Explain why reads do not migrate data and how a no-op write re-encodes an object in the storage version.

for a senior

Walk through a full retirement: deprecate, stop serving, migrate with StorageVersionMigration or by hand, verify storedVersions, then delete, and name the failure if the order is broken.

for a principal

Treat version retirement as a release process owned by the platform: who runs migrations, how long old versions stay served, and how operators ship them.

## What `status.storedVersions` records Every CustomResourceDefinition (CRD) has exactly one **storage version** — the `spec.versions` entry with `storage: true` — and the API server encodes each write in it. Changing the storage version does not re-encode the objects already in etcd, so over a CRD's life etcd can hold objects in several versions at once. The CRD tracks that in **`status.storedVersions`**: a list of every version that has ever been the storage version. The apiextensions API server **only adds** to it — whenever a CRD is created or updated and its storage version is not yet listed, that version is appended. Nothing in the API server checks etcd to remove an entry; removing one is the job of whoever migrated the data. For the search-autocomplete team's `suggestindexes.autocomplete.example.com`, a CRD that started at `v1alpha1` and now stores `v1` shows `storedVersions: ["v1alpha1", "v1"]`. ## Why removal is refused CRD validation enforces two rules on that list: - It must include the current storage version ("must have the storage version v1"). - Every listed version must still appear in `spec.versions`. Dropping `v1alpha1` fails with: `v1alpha1 was previously a storage version, and must remain in spec.versions until a storage migration ensures no data remains persisted in v1alpha1 and removes v1alpha1 from status.storedVersions`. The reason is decoding. The converter accepts only source versions that are listed in `spec.versions`. If an object stored as `v1alpha1` is read after that entry is gone, the request fails with "request to convert CR from an invalid group/version", and a list that contains even one such object fails as a whole — every controller that lists the resource stalls. ## Migrating with `StorageVersionMigration` In v1.37 the **StorageVersionMigrator** feature is GA and kube-controller-manager runs the `storage-version-migrator-controller`. You ask it to migrate a resource with an object in `storagemigration.k8s.io/v1`: ```yaml apiVersion: storagemigration.k8s.io/v1 kind: StorageVersionMigration metadata: name: suggestindexes-to-v1 spec: resource: group: autocomplete.example.com resource: suggestindexes ``` What happens next: 1. The controller sets a `StorageMigrating` condition on the CRD, recording the CRD's current `metadata.generation`. 2. It sends a **no-op patch** to every object of the resource. The API server re-encodes each one in the storage version, and because the new bytes differ from the old ones, the write reaches etcd. 3. When all patches succeed, the `StorageVersionMigration` gets a `Succeeded` condition and the controller sets `storedVersions` to just the storage version — but only if the CRD's generation still matches the one it recorded. If someone changed the CRD spec mid-run, it reports that the stored versions were not updated "due to a generation mismatch", and you run a new migration. ## Migrating by hand On a cluster without that controller, do the same thing yourself: read every object and write it back unchanged, then patch the CRD status. - `kubectl get <plural>.<group> -A -o json | kubectl replace -f -` rewrites every object; a `409 Conflict` on a busy object only means it was written meanwhile, which already re-encoded it. - `kubectl patch customresourcedefinition <name> --subresource=status --type=merge -p '{"status":{"storedVersions":["v1"]}}'` records the result. A JSON merge patch replaces the whole list. - Merely reading objects migrates nothing — reads convert in memory only. ## The safe retirement sequence 1. Ship `v1` with `storage: true` and, if schemas differ, a working conversion webhook. 2. Mark `v1alpha1` `deprecated: true` so clients see warnings. 3. Set `served: false` for at least one release so stray clients fail loudly while the data is still readable. 4. Migrate stored objects and confirm `storedVersions` is `["v1"]`. 5. Remove the `v1alpha1` entry, and its code path in the conversion webhook. ## Pitfalls | Mistake | Consequence | |---|---| | Trimming `storedVersions` without migrating | Validation lets you delete the version; old objects become unreadable and lists fail | | Migrating while the conversion webhook is down | Every patch needs conversion, so the migration fails | | Editing the CRD spec during migration | The migrator refuses to update `storedVersions` | | Removing the version in the same release that stops serving it | No window to catch clients still using it | If an old version was removed too early, re-adding its entry to `spec.versions` (with `served: false` and `storage: false`, plus its conversion code if you use a webhook) makes the stored objects decodable again, after which you can migrate properly.

  • What happens if you trim a Kubernetes CRD's storedVersions to v1 without migrating and then delete v1alpha1?
    Validation passes, but objects still encoded as `v1alpha1` can no longer be converted: reads fail with "request to convert CR from an invalid group/version", and any list containing one fails entirely, so controllers stall. Recover by re-adding the `v1alpha1` entry with `served: false` and `storage: false`, then run a real migration.
  • Why might a Kubernetes StorageVersionMigration succeed yet leave the CRD's storedVersions unchanged?
    The migrator records the CRD's `metadata.generation` when it starts. If the CRD spec changed during the run — for instance the storage version moved again — the generations no longer match, so it cannot be sure every object is in the current storage version. It marks the migration succeeded but skips the `storedVersions` update; create a new migration to finish.
  • Why keep a Kubernetes CRD version at served: false for a release before deleting it?
    Clients still using the old version then fail with 404 while the data remains decodable, so you find them without risking stored objects. It also separates the client-facing change from the storage change, which makes each easy to roll back: flipping `served` back is instant, while a deleted version entry with data left in etcd is an outage.

saying these in an interview costs you the question

  • The API server re-encodes old objects on its own after storage moves to v1.
  • Editing storedVersions by hand is itself the storage migration.
  • Listing every object with kubectl get migrates them.
  • Setting served: false on v1alpha1 converts its stored objects.
  • A storage migration can run safely while the conversion webhook is down.