skip to content

Your platform team wants every service on a shell-less image — what must on-call engineers be given before that standard is safe?

level: principalimportance: nice to knowfreq 28%

answer

  1. a capability was removed, not added
  2. replace what the shell was for
  3. mechanism without permission is nothing
  4. the rebuild copies the image, not the world
  5. measure it on real pages

basics

~20 s

A replacement for every use the removed shell had: a permitted joined-debug-container path with a maintained tool image, diagnostics on the output stream, the terminated instance's output reaching the pager, a file-extraction route, and a one-step local rebuild with tools.

solid answer

~50 s

The standard removes exactly one capability — running a program inside the boundary — and it is safe only once each thing that capability was used for has a replacement that works under pressure at 02:00. Five of them: the **joined debug container**, enabled on the platform *and permitted for whoever carries the pager*, with a maintained tool image; **diagnostics on the output stream** by default, plus a durable path for anything too large; the **terminated instance's retained output** reaching the on-caller without a live process; a supported way to **copy a file out** of a running or stopped instance; and a **one-step local rebuild** of the same image with tools layered on. Then say out loud what you accept: every one of these is slower than typing into a live container, and an environment-only failure is still only reproducible in that environment.

go deeper

for a junior

The takeaway is that removing the shell is an operational decision, not only a build one: someone has to be able to answer an unplanned question about a running workload afterwards.

for a middle

Be able to name what the shell was used for and which replacement covers each use — the joined debug container, the output stream, the retained output of an ended instance, extraction, and the local rebuild.

for a senior

Argue the sequencing: drill the replacement paths with someone outside the platform team before the mandate, and be concrete that permission, not mechanism, is the usual gap.

for a principal

Own the trade across every future incident — a slower first touch on all of them against a property of the artifact — and set the measurement that tells you whether the replacements are real.

## What the standard actually removes It is easy to argue about this in the abstract and miss what changes operationally. A shell-less image removes one thing: **the ability to run an arbitrary program in the workload's own context**. It does not remove evidence, it does not remove access to the host, and it does not make the workload harder to observe. It removes the improvised move — the one an on-call engineer reaches for when the prepared signals did not anticipate the question they now have. That move is load-bearing in most teams precisely because it is unplanned. So the honest framing of the decision is not "is a stripped image better" — it is "can we answer unplanned questions without it, quickly, at the worst hour, with the person who happens to be paged". ## The five replacements, and what each one really requires 1. **The joined debug container.** The mechanism has to exist on your platform *and* the person paged has to be authorised to use it. This is where estates fail: the capability is demonstrated once by the platform team, who hold broad permissions, and is then discovered to be unavailable to on-call during the first incident. It also needs a **maintained tool image** — someone owns what is in it and keeps it current — and a known answer to whether adding a container disturbs the running workload on your platform, because designs differ. 2. **Diagnostics on the output stream by default.** Anything the workload can say about itself needs no program inside the boundary to be read. This is a service-authoring standard, not a platform feature, so it has to be written down and reviewed. 3. **The terminated instance's output reaching the pager.** Once a workload has ended there is nothing to join at all, so the retained output is the only evidence — and it is bounded and host-local. If the on-caller's path to it is "ask someone with host access", the replacement is not in place. 4. **A file-extraction route.** Copying a file out runs nothing inside the boundary, so it works on stripped images and, on most designs, on stopped instances. It needs to be a documented, permitted step rather than folklore. 5. **A one-step local rebuild with tools added.** The engineer needs to build the same image with an interpreter and tools layered on and run it interactively, without reverse-engineering the build first. ## What you are accepting | Replacement | Compared with a shell in the image | |---|---| | Joined debug container | slower to first look; needs permission; unavailable once the instance ends | | Output stream | only answers questions someone anticipated | | Retained output of the ended instance | bounded retention, host-local, no follow-up questions possible | | File extraction | one file at a time, no execution, no live state | | Local rebuild with tools | reproduces the image, never the environment | That last row is the one to say out loud, because it is the case the standard genuinely cannot cover. A failure that only happens in one environment depends on that environment's data, configuration, neighbours and permissions, and none of those come along with a rebuild. ## How to decide, and how to know it worked - **Sequence it.** Ship the replacements, exercise them in a drill with someone who does not own the platform, *then* make the standard mandatory. A mandate with tooling promised for later converts every incident in the gap into an argument. - **Scope it honestly.** The workloads that get paged at 02:00 and the workloads nobody has ever debugged are not the same population, and a phased standard is not a failure of nerve. - **Measure it on real pages.** After each incident, ask whether anyone needed a program inside the boundary, what they did instead, and how long it took. If the recurring answer is "asked the platform team to build a special image", the replacement is missing and the standard is borrowing time from the on-call rotation. - **Do not buy the pushback off with privilege.** Granting broad debug rights to everyone to end the argument trades the thing the standard was for against the convenience it removed, and usually keeps neither. The judgment a lead owns here is a trade across every future incident, not one: a slower first touch on all of them, in exchange for a property of the shipped artifact. That calculus changes with how often you are paged and how mature the five paths above are, and the right answer for a team that is paged twice a year is not the right answer for one that is paged twice a week.

  • Which of those five replacements is most often missing in practice?
    The permission on the debug-container path. The mechanism almost always exists and was demonstrated by people holding broad rights, so nobody noticed that the on-call rotation does not hold them. It is then discovered mid-incident, at the point where the cost of discovering it is highest.
  • How would you phase the rollout rather than mandating it everywhere at once?
    Start with the workloads whose failure modes are already well covered by their own output, keep the ones that get paged most until the debug path has been drilled by an on-caller, and set a date rather than an exception list. A standing exemption list becomes permanent; a date forces the tooling work.
  • What would make you decide the standard is not worth it for a given team?
    A team that is paged often, whose failures are environment-specific, and whose platform does not offer a joined debug container their on-callers may use. There the standard adds real minutes to every incident and the replacement for the improvised move does not exist — so fix the platform path first and revisit.

saying these in an interview costs you the question

  • Mandates the standard first and promises the tooling afterwards.
  • Assumes on-call already holds permission for the debug-container path.
  • Treats a local rebuild as equivalent to the failing environment.
  • Counts the standard as done once images build and deploy.
  • Grants broad debug privileges to everyone to stop the pushback.
  • Ignores that a terminated instance cannot be joined at all.