You own the shared base images for dozens of services. How would you standardise the ENTRYPOINT/CMD contract, PID 1 behaviour, and shutdown-signal handling across all of them, and what trade-offs would you weigh?
answer
- contract = platform API, not a Dockerfile detail
- exec form + exec "$@" = PID 1 invariant
- tini for forking services only
- grace period from measured drain, not the default
- lint + signal test + base image beats a wiki page
basics
~20 sMandate exec-form instructions, decide one convention for whether ENTRYPOINT is the app or a thin wrapper, guarantee the real process ends up as PID 1 with a SIGTERM handler, standardise grace periods against measured drain time, and enforce it with lint plus an automated termination test.
solid answer
~60 sI would treat the start-up contract as a platform API and pin four things. **Form.** Exec form only for CMD/ENTRYPOINT; shell form fails lint. Where expansion is needed, `["sh","-c","exec …"]`. **Shape.** One convention: ENTRYPOINT is a thin wrapper that `exec "$@"`s, CMD is the real command. That keeps `command`/`args` overrides in orchestrator manifests predictable, and gives one place to add cross-cutting start-up behaviour later. **PID 1.** The application process must be PID 1 and must handle SIGTERM. Add tini/`--init` only for services that fork; it is a reaping fix, not a substitute for a handler. **Shutdown.** Standard drain sequence — fail readiness, stop accepting, finish in-flight — with grace periods set from measured p99 request duration rather than the 10-second default. Enforcement matters more than the doc: a build-time lint rule plus a CI test that starts the container, signals it, and asserts a clean exit inside the budget. Trade-offs: a mandatory wrapper costs debuggability and hides behaviour from the Dockerfile; strictness pays off only if migration is cheap for existing images.
go deeper
Not expected at this level; if asked, focus on the concrete rules — exec form, exec "$@", handle SIGTERM.
Give the mechanics correctly and one enforcement idea, such as a linter rule, without needing the full fleet-migration story.
Cover the contract, PID 1 invariant, measured grace periods, and how you would verify behaviour in CI rather than trusting review.
Lead with the contract as a platform API, name the trade-offs (uniformity vs autonomy, wrapper cost, universal init), and describe a staged migration with exemptions that have owners and dates.
## Framing At fleet scale the question is not 'what does ENTRYPOINT do' but 'what promise does every image in this organisation make about how it starts and stops'. That promise is consumed by orchestrators, deploy tooling, incident responders and anyone writing a manifest. Inconsistency there shows up as ten extra seconds per pod on every deploy, dropped requests during rollouts, and engineers who cannot get a shell into an unfamiliar image during an incident. ## The decisions to pin **1. Form.** Exec (JSON) form for CMD and ENTRYPOINT, without exception. Shell form silently changes PID 1 and, for ENTRYPOINT, discards all arguments. This is cheap to enforce — BuildKit already warns, and a Dockerfile linter (hadolint or a house rule) can promote the warning to a build failure. **2. Shape of the contract.** Two defensible conventions: - *App as entrypoint*: `ENTRYPOINT ["/app/server"]`, `CMD` holds default flags. Simplest, most honest; overriding requires `--entrypoint`. - *Thin wrapper as entrypoint*: `ENTRYPOINT ["/entrypoint.sh"]` that ends in `exec "$@"`, `CMD ["/app/server", …]`. Slightly more machinery, but gives one hook for cross-cutting start-up concerns (config rendering, secret materialisation, trace-context bootstrap) and keeps orchestrator `args` overrides working naturally. Pick one and apply it everywhere. Mixed conventions are the actual cost: a manifest author must read each image's Dockerfile to know whether to set `command` or `args`. **3. PID 1 and reaping.** The invariant is: the process that knows how to shut down cleanly is the one that receives SIGTERM. Practically that means exec form, `exec` in any wrapper, and no `su`/`sudo` in the handoff (use `gosu`/`su-exec`). Add an init (tini, or the runtime's `--init` equivalent) only for images whose process genuinely forks children and would otherwise leak zombies; making it universal adds a process and a failure mode to services that do not need it, and it can mask a missing SIGTERM handler rather than fix it. **4. Shutdown semantics.** Standardise the sequence — on SIGTERM: mark unready so traffic stops arriving, stop accepting new work, drain in-flight work, close pools, exit 0 — and set the grace period from measured request duration, not the 10-second default. Long-poll, streaming and batch services need explicit, larger budgets; a service that cannot drain within its budget is a design issue to surface, not a number to keep raising. **5. Debuggability.** Every mandatory entrypoint makes `docker run img sh` stop working. Compensate deliberately: publish a debug tag for shell-less bases, document the `--entrypoint sh` invocation in the standard runbook, and keep the wrapper short enough to read in an incident. ## Enforcement A convention that lives only in a document decays. Three layers work: - **Lint at build.** Reject shell-form CMD/ENTRYPOINT, missing CMD when ENTRYPOINT is set, entrypoint scripts without `exec "$@"`, and non-executable entrypoints. - **Test the behaviour.** A shared CI job that starts the image, sends SIGTERM, and asserts the container exits 0 within the declared budget catches the failures lint cannot see — a handler that exists but never fires, a wrapper that forgot `exec`. - **Bake it into the base image.** The strongest form of a convention is one the derived image gets for free: the base declares the wrapper, the tini decision and a sane default CMD, so a service Dockerfile only overrides CMD. ## Trade-offs to name out loud - *Uniformity vs autonomy.* A rigid contract makes tooling simple and images boring, but every legitimate exception (a batch job with a 20-minute drain, a third-party image you do not control) needs an escape hatch that does not require a platform-team ticket. - *Wrapper vs no wrapper.* A wrapper is a future-proofing option with an ongoing readability and debuggability cost. If nothing cross-cutting exists yet, the honest answer may be 'no wrapper until there is a second reason for one'. - *Universal init vs targeted.* Universal `--init` removes a class of zombie bugs and adds a universal dependency; targeted use is leaner but relies on someone noticing the forking case. - *Migration cost.* A new standard applied to dozens of existing images is a migration programme. Sequence it: lint in warn mode, fix base images, flip to blocking, and accept that the last few stragglers will need per-image exemptions with an owner and a date. ## What good looks like Any engineer can read any service's manifest and know what will run; `docker stop` and rolling deploys finish in well under a second for typical services; no image in the fleet ends a deploy with SIGKILL; and the wrapper, if there is one, fits on one screen.
- Would you make tini (or `--init`) universal across the fleet?Usually not. It solves zombie reaping and signal forwarding for processes that fork children, which is a minority. Applying it everywhere adds a process and a dependency to every image and can hide the real defect — an application with no SIGTERM handler. I would default it off, enable it per-image where the workload forks, and catch the need through the termination test rather than blanket policy.
- A team needs a 20-minute drain for a batch consumer. How does that fit a standard contract?The standard should define a mechanism, not a single number: a declared shutdown budget per service, with the orchestrator's grace period derived from it. The 20-minute case then becomes a documented, reviewed value rather than an exception to the rule. It is also a prompt to ask whether the work should be checkpointable so shutdown is not held hostage by one long unit of work.
- How do you enforce this on third-party images you cannot rebuild?Wrap rather than fight: either build a thin derived image that sets the conventional entrypoint and CMD, or set `command`/`args` and the grace period in the deployment manifest. Record them as known exceptions with owners so the fleet inventory stays honest about which images meet the contract.
saying these in an interview costs you the question
- Treating the standard as a document rather than something enforced by lint, tests and the base image.
- Mandating tini everywhere as a substitute for application SIGTERM handlers.
- Keeping the default 10-second grace period without measuring how long services actually take to drain.
- Ignoring the debuggability cost of mandatory entrypoints and shell-less base images.
- Proposing a fleet-wide change with no migration path for existing images.