skip to content

How does controller-runtime's envtest let you test a Kubernetes controller, and what does an envtest environment deliberately not run?

level: middleimportance: nice to knowfreq 33%

answer

  1. real API server, nothing else
  2. etcd plus kube-apiserver binaries
  3. KUBEBUILDER_ASSETS and CRD paths
  4. no garbage collector, no kubelet
  5. admin credentials hide RBAC gaps

basics

~20 s

envtest starts real etcd and kube-apiserver binaries locally, installs your CRDs and returns a rest.Config for your manager. It runs no kube-controller-manager, scheduler or kubelet, so there is no garbage collection, no Pods running and no finished namespace deletion.

solid answer

~50 s

`envtest.Environment` downloads or locates `etcd` and `kube-apiserver` binaries, found through `KUBEBUILDER_ASSETS` (which Kubebuilder's `make test` fills from `setup-envtest use ... -p path`) or `BinaryAssetsDirectory`. It starts them, installs the CRDs from `CRDDirectoryPaths`, and returns a `*rest.Config`. The test builds a real manager against that config, registers the reconciler, starts it in a goroutine, creates a `GameSession`, and asserts with `Eventually` that the expected Deployment and status appear. Because the API server is real, you get real schema validation, defaulting, `resourceVersion` conflicts and status-subresource behaviour, which the fake client does not reproduce. What is missing is everything else in a cluster: no built-in controllers, so no ReplicaSets, no garbage collection from ownerReferences, and namespaces stuck in Terminating; and no scheduler or kubelet, so Pods are stored but never run. Assert on ownerReferences rather than on deletion, and set dependents' status yourself.

code

bash · 2 lines
bash
export KUBEBUILDER_ASSETS="$(setup-envtest use -p path)"
go test ./internal/controller/... -v

go deeper

for a junior

Recall that envtest runs a real etcd and API server so controller tests talk to real Kubernetes APIs, and that Kubebuilder's make test sets it up.

for a middle

Explain the setup steps (binaries, CRD install, manager in a goroutine, Eventually) and list what is missing: controllers, garbage collection, scheduling and running Pods.

for a senior

Design a test pyramid for an operator: fake client for branch logic, envtest for API semantics, and a kind suite for garbage collection and RBAC. Know the flaky-test traps, such as leftover namespaces.

for a principal

Decide how much of an organisation's operator testing budget goes to envtest versus full-cluster suites, and require RBAC-realistic end-to-end runs before release.

## Where envtest fits A controller can be tested at three levels, and each catches different mistakes: | Level | What talks to your code | What it catches | |---|---|---| | Unit test with the **fake client** | An in-memory object tracker | Branch logic in `Reconcile` | | **envtest** | A real kube-apiserver backed by a real etcd | Schema validation, status subresource, conflicts, watches | | End-to-end on **kind** or a real cluster | A full cluster | Scheduling, garbage collection, networking, the deployed manifests | `envtest` (package `sigs.k8s.io/controller-runtime/pkg/envtest`) is the middle level. General integration-testing strategy belongs to the testing topics; this answer covers what envtest provides for a Kubernetes controller specifically. ## How it starts 1. **Binaries.** envtest needs `etcd` and `kube-apiserver` (and ships `kubectl` alongside them). The `setup-envtest` tool downloads a chosen Kubernetes version, and `setup-envtest use <version> -p path` prints the directory. Kubebuilder's `make test` exports that as `KUBEBUILDER_ASSETS`. Alternatively, `Environment.BinaryAssetsDirectory` points at them. 2. **Start.** `testEnv.Start()` launches both processes on local ports and returns a `*rest.Config` with admin credentials. 3. **CRDs.** Directories listed in `CRDDirectoryPaths` (usually `config/crd/bases`) are installed, and envtest waits until they are served. With `ErrorIfCRDPathMissing: true`, a wrong path fails the suite instead of silently installing nothing. 4. **Manager.** The test builds `ctrl.NewManager(cfg, ...)`, registers the real reconciler with `SetupWithManager`, and runs `mgr.Start(ctx)` in a goroutine. 5. **Assertions.** The test creates objects through a client and polls with `Eventually` (Gomega) until the reconciler has produced the expected state. 6. **Teardown.** Cancel the context, then call `testEnv.Stop()`. Setting the environment variable `USE_EXISTING_CLUSTER=true` points the same suite at a real cluster from your kubeconfig instead of starting binaries. ## What you get that the fake client lacks - **OpenAPI validation and pruning** from your generated CRD, so a bad `+kubebuilder:validation` marker fails in the test. - **Defaulting** and **CEL validation rules** as the API server applies them. - **Status subresource semantics** (when the CRD enables `/status`): a plain `Update` ignores status changes, and `Status().Update` ignores spec changes. - **Optimistic concurrency**: stale writes fail with a conflict, just as in production. - **Real watches**, so `Owns` and `Watches` wiring is exercised end to end. - **Admission webhooks**, if you configure `WebhookInstallOptions`. ## What it deliberately does not run envtest is a control-plane API with nothing acting on it: - **No kube-controller-manager.** Creating a Deployment produces no ReplicaSet and no Pods. The **garbage collector** does not run, so deleting a `GameSession` leaves its owned Deployment in place even with a correct `ownerReference`. **Namespace deletion never completes**: the namespace stays `Terminating`. - **No kube-scheduler and no kubelet.** Pods are stored but never bound to a node or started, so their status stays empty unless the test writes it. - **No cloud or network components.** Services get no load balancer, and nothing routes traffic. Practical consequences: - Assert that the child has the right `ownerReferences` instead of asserting that it disappears. - Simulate readiness by writing `status.readyReplicas` on the Deployment yourself, then check that the reconciler copies it into `GameSession.status`. - Create a fresh namespace per test instead of reusing a deleted one, and make the reconciler tolerate leftovers from earlier tests. ```go testEnv := &envtest.Environment{ CRDDirectoryPaths: []string{filepath.Join("..", "..", "config", "crd", "bases")}, ErrorIfCRDPathMissing: true, } cfg, err := testEnv.Start() Expect(err).NotTo(HaveOccurred()) mgr, err := ctrl.NewManager(cfg, ctrl.Options{ Scheme: scheme.Scheme, Metrics: metricsserver.Options{BindAddress: "0"}, }) Expect(err).NotTo(HaveOccurred()) Expect((&GameSessionReconciler{Client: mgr.GetClient(), Scheme: mgr.GetScheme()}).SetupWithManager(mgr)).To(Succeed()) ctx, cancel = context.WithCancel(context.Background()) go func() { defer GinkgoRecover() Expect(mgr.Start(ctx)).To(Succeed()) }() ``` ## Keeping envtest suites reliable - **Always poll.** The reconciler runs in another goroutine, so every assertion about its effects needs `Eventually` with a timeout. A single immediate `Get` is a race. - **Read with a live client in assertions** when the test needs to see its own write immediately, or accept the cache lag and poll. - **Pin the Kubernetes version** of the envtest binaries to the version you deploy against, so CRD validation features behave the same in CI and in the cluster. - **Share one environment per package.** Starting etcd and kube-apiserver takes seconds, so start them once in the suite setup rather than per test. - **Stop cleanly.** Cancel the manager's context before `testEnv.Stop()`, or the suite can hang waiting for the processes to exit. ## When to reach past envtest Use a kind-based end-to-end suite when the behaviour you care about depends on the missing pieces: cascading deletion, Pods becoming Ready, or the generated RBAC being sufficient. envtest runs with admin credentials, so **it does not test your ClusterRole**. That gap is the most common reason an operator passes CI and then fails with forbidden errors after deployment.

  • Your envtest passes, but the deployed operator fails with forbidden errors. Why didn't the test catch it?
    envtest hands the test an admin `rest.Config`, so the manager runs with full permissions and never exercises the generated `manager-role` ClusterRole. A missing `+kubebuilder:rbac` marker or a stale `make manifests` only shows up when the operator runs under its real ServiceAccount, for example in a kind-based end-to-end suite that deploys `config/`.
  • When is the fake client still the better choice over envtest?
    For fast, table-driven tests of branch logic inside `Reconcile`: how the code maps a spec to a child object, or how it reacts to a specific error you inject. The fake client starts instantly and needs no binaries, but it does not apply CRD schema validation, defaulting or admission, so anything that depends on API server behaviour belongs in envtest.

saying these in an interview costs you the question

  • envtest runs a complete miniature cluster including the scheduler and kubelet
  • Deleting the owner in envtest proves the ownerReference garbage collection works
  • envtest verifies that the operator's generated ClusterRole is sufficient
  • The fake client applies the same CRD validation as envtest
  • Pods created in envtest start running once they are scheduled