skip to content

Your production Kubernetes workloads run distroless images as non-root with read-only root filesystems, and on-call engineers complain they cannot debug incidents. How would you give them a debugging capability without undoing that hardening?

level: principalimportance: should knowfreq 26%

answer

  1. tools out of the image, into the debug container
  2. approved toolbox image, pre-pulled
  3. break-glass exec / ephemeralcontainers RBAC
  4. audit the subresources, recycle debugged pods
  5. reduce demand: diagnostics endpoints, dumps to object storage

basics

~20 s

Keep images minimal and move debugging out of them: an approved toolbox image attached as an ephemeral container with kubectl debug --target, granted through a time-boxed break-glass role, audited, and paired with in-app diagnostics so most incidents need no shell at all.

solid answer

~60 s

I separate three questions: what capability is needed, who may use it, and what proves it happened. **Capability.** Debugging moves from the workload image to a curated debug image in the internal registry (shell, curl, ss, tcpdump, language-specific tools), attached at incident time with `kubectl debug -it pod --image=<toolbox> --target=app`. That keeps the production image small and unprivileged while making tools reachable in seconds. For pods that will not stay up, the documented path is `kubectl debug --copy-to` with an overridden command; for machine-level failures, `kubectl debug node/` from a namespace labelled privileged. **Access.** The default on-call identity is read-only plus `pods/log` and `pods/portforward`. `pods/exec`, `pods/ephemeralcontainers` and node debug come from a break-glass role bound for a bounded window and announced in the incident channel. **Evidence.** Audit policy records those subresources at request level; pods that were debugged are recycled afterwards, since ephemeral containers cannot be removed. **Demand reduction.** Most shells are requested because something else is missing — structured logs, metrics, an admin endpoint, on-demand profiles written to object storage. I would track how often a shell was truly required and close those gaps first.

go deeper

for a junior

Know that hardened images are debugged by attaching a separate tool container rather than by adding tools to the image.

for a middle

Describe the concrete mechanics — debug image, kubectl debug with --target, --copy-to for dead pods — and why each fits.

for a senior

Add the operational and access dimensions: pre-pulled images, Pod Security constraints, audit, and recycling debugged pods.

for a principal

Own the whole trade-off: reduce the need for shells, make the remaining access break-glass and reviewable, and name the costs and the untested-fallback risk.

## Frame the tension honestly Hardening removes tools from images on purpose: smaller attack surface, fewer CVEs to patch, no shell for an attacker to pivot with. Debuggability is not free either — the answer is not to re-add a package manager to the production image, and not to tell on-call to cope. The capability should be **attachable on demand, privileged deliberately, and observable afterwards**. ## Layer 1 — make most incidents shell-free Every request for a shell is a signal that some evidence was not exported. Before granting broad interactive access, close the common gaps: structured logs on stdout at the right level with correlation IDs; per-endpoint request/error/duration metrics; a diagnostics endpoint exposing configuration, build info, dependency status and thread or goroutine dumps; and on-demand profiles or heap dumps that the process writes to object storage rather than to a local disk someone must copy off. Keep this scoped — long-term dashboards and alerting are the observability platform's job — but the last-mile diagnostics endpoints belong to the service. A team that measures "incidents that required a shell" usually finds the number falls once these exist. ## Layer 2 — a sanctioned debug path Standardize on ephemeral containers rather than mutable images: - One or two **approved debug images**, built and scanned like any other artifact, hosted in the internal registry so they pull fast and are available on air-gapped clusters. Pre-pulling them onto nodes removes the failure mode where a node with a broken runtime or full disk cannot fetch tools. - A documented decision tree: live pod misbehaving goes to `kubectl debug --target`; container not staying up goes to `--copy-to` with a changed command; node-level symptoms go to `kubectl debug node/`; anything reachable by port goes to `kubectl port-forward` instead of a shell. - Security-context profiles chosen consciously: `--profile=restricted` by default and elevated profiles only where the namespace's Pod Security admission level permits them. Expect that a `restricted` namespace rejects privileged debug containers — that is the control working, and you need a decided answer (usually a dedicated namespace for node debugging) rather than a surprise mid-incident. - Known limits written down: ephemeral containers cannot be removed, carry no resource requests and can push a tight node over allocatable, and copies are not the failing instance. ## Layer 3 — access as break-glass, not a standing grant `pods/exec` and `pods/ephemeralcontainers` mean code execution inside a workload with its service-account token and mounted secrets; node debug means root on the machine and a path to cluster-admin. Standing grants for these in production are hard to defend. A workable model: read-only plus `pods/log` and `pods/portforward` by default; an escalation role bound for a fixed window through an approved, logged mechanism; node debug held by the platform team only. Pair with an audit policy that records those subresources in full, ship those events somewhere durable, and review them periodically — the review is what makes the grant sustainable, not the grant itself. ## Layer 4 — hygiene after the incident A debugged pod is a modified pod: the ephemeral container is in its spec permanently, and anything the responder changed inside the container is invisible to Git. The rule should be that pods touched interactively are recycled once the incident is over, and that any fix applied by hand is immediately reproduced as a manifest or image change. Forgotten node-debugger pods and hand-patched replicas are two of the most common findings in cluster reviews, and both are process failures rather than tooling ones. ## Trade-offs a principal should name The curated-toolbox model costs image maintenance and scanning, and adds friction during an incident. Break-glass access adds an approval step that can be resented at 3am, so the mechanism must be fast or people will route around it. Pre-pulling debug images consumes node disk. And no debug tooling helps if the node itself is unreachable — the fallback of cloud-console or SSH access must exist and be tested, or the whole scheme has an untested single point of failure.

  • How would you measure whether this debugging strategy is working?
    Track the share of incidents where a responder needed interactive access at all, the time from wanting a debug container to having one, and the number of break-glass escalations plus their review outcomes. A falling need-a-shell rate with stable resolution time means the diagnostics gaps are closing; rising time-to-tools means the approval or image-pull path needs work.
  • An engineer argues it is simpler to just ship images with a shell and busybox. What is your counter-argument, and when might they be right?
    A shell in every production image widens the attack surface for every replica permanently to save minutes during rare incidents, and it adds packages that must be patched for CVEs. That said, for a low-risk internal batch workload, or a very small team without the capacity to maintain a toolbox image and break-glass tooling, the simpler option can be the honest engineering choice — the decision should follow the workload's exposure and the team's operational maturity, not dogma.

saying these in an interview costs you the question

  • Proposing to add a package manager or shell back into hardened production images as the default fix
  • Granting standing pods/exec in production to everyone on call
  • Ignoring that ephemeral containers cannot be removed and pods must be recycled
  • Assuming a privileged debug container will run in a namespace enforcing the restricted Pod Security Standard
  • Treating debugging purely as a tooling problem while leaving no diagnostics in the application

context