Which `docker run` flags does a scheduler take over, and what stays in the image?
answer
- Two documents, not one
- Host-shaped settings versus artefact properties
- Would this value differ between two hosts?
- restart, publish, mount, limits, discovery, health
- ENTRYPOINT, USER, layers and SIGTERM stay
basics
~20 sA scheduler takes over the host-shaped docker run flags — --restart, -p, -v, --cpus, --memory and health checking — restating them as fields in its own spec. The image keeps what it runs and as whom: ENTRYPOINT/CMD, USER, its filesystem, and logging to stdout.
solid answer
~50 sEverything on a `docker run` command line that describes *this host* moves into the scheduler's spec, and the image itself is untouched. That means the restart policy (`--restart`), port publishing (`-p`), volumes and bind mounts (`-v`/`--mount`), resource limits (`--cpus`, `--memory`), name-based discovery you got from `--name` on a user-defined bridge, and health checking — a scheduler declares its own probes and acts on the result. What stays with the image is the process contract: the default command in exec form (`ENTRYPOINT`/`CMD`), the `USER` it runs as, the filesystem and dependencies, reading configuration from the environment instead of baked-in files, writing logs to stdout/stderr, and exiting cleanly on `SIGTERM`. A useful test: if a value would differ between staging and production, or between two hosts, it belongs in the runtime spec and not in a layer. `EXPOSE` is only documentation; it publishes nothing.
code
dockerfile · 6 linesFROM scratch
COPY --chmod=0555 checkout /checkout
USER 10001:10001
ENV CHECKOUT_PORT=8143
EXPOSE 8143
ENTRYPOINT ["/checkout"]go deeper
Be ready to say which parts of a container come from the Dockerfile and which come from the run command. Knowing that ENTRYPOINT and USER live in the image while -p, -v and --restart are typed at run time is the whole foundation here.
Explain each flag's mechanism: what --cpus writes into a cgroup, what -p does to host NAT, how --name gets resolved on a user-defined bridge. An interviewer expects you to map each one onto its equivalent field in a runtime spec rather than just listing them.
Show the judgement call: given an existing run command, say which flags translate, which ones are a warning sign (host bind mounts, hard-coded ports) and what you would change in the image before the migration rather than after.
Own the boundary as a rule the whole fleet follows. Be ready to argue why environment-specific values must never enter a layer, and how you enforce that on images your teams build rather than trusting review to catch it.
## The image is one document, the run is another An OCI image records **what to run and the filesystem to run it in**: a default `ENTRYPOINT`/`CMD`, a default `USER`, `WORKDIR`, `ENV` defaults, `EXPOSE` metadata, labels, an architecture, and the layers themselves. Everything else you type on a `docker run` line is **runtime configuration**. It lives in the container object, not the image; it disappears with `docker rm`; and it is exactly the part a scheduler replaces. Handing a workload over is therefore not a rewrite — the image crosses the boundary unchanged, and the run command is restated in the scheduler's own vocabulary. Take an order-checkout API: a single static Go binary in a `scratch` image (it was 1.7 GB back when it shipped on a full distro with a build toolchain inside). On one host it runs as: ``` docker run -d --name checkout --restart=on-failure \ -p 8143:8143 --cpus=1.35 --memory=768m \ --network edge -v checkout-tmp:/tmp \ registry.example.com/checkout:1.4.2 ``` Flag by flag, here is who owns what after the handoff. **`--restart` — gone.** The daemon's restart policy is a single-host supervisor: it reacts to the container exiting, on this machine, for as long as this daemon is up. A scheduler *is* the supervisor. It decides whether a dead instance is restarted in place or replaced elsewhere, and it counts failures for its own backoff. Leaving both in charge gives you two supervisors racing to bring the same process back. **`-p 8143:8143` — gone.** Publishing rewrites the host's NAT rules so one host port reaches one container. It cannot express "eleven instances, reachable under one name". Addressing becomes the scheduler's concern; all the image must do is listen on `0.0.0.0` at a port it can be *told*, rather than binding `127.0.0.1` or hard-coding one. **`--name` and bridge DNS — gone.** On a user-defined bridge, the engine's embedded resolver at 127.0.0.11 turns `--name checkout` into a name other containers on that same bridge, on that same host, can resolve. That is a one-host convenience. Cross-host discovery is supplied by the scheduler, so the image must take peer addresses from configuration and never assume a container name resolves. **`--cpus` and `--memory` — restated, not dropped.** `--cpus=1.35` writes a CPU quota-per-period pair into the container's cgroup; it does not pin cores (`--cpuset-cpus` does that). `--memory=768m` sets a hard ceiling, and the kernel OOM-kills the process when it is exceeded — you see exit code 137. Every scheduler has fields for both, usually with a *requested* amount for placement in addition to the ceiling. The behaviour underneath is the same cgroup, so numbers you measured under `docker run` transfer; what changes is that something now uses them to decide *where* the container fits. **`-v` / `--mount` — restated, and often refused.** A named volume or bind mount is host-local storage. A scheduler will happily give you scratch space, but a bind mount to `/opt/checkout/data` is a promise that this container always lands on this machine — the one thing you gave up by handing over placement. **Health checking — moves.** Under plain `docker run` the engine runs the image's `HEALTHCHECK` and records a state; it does not act on it. A scheduler declares its own probes and *does* act: restarting, replacing, or pulling an instance out of rotation. **Secrets and configuration — restated.** `-e DB_PASSWORD=...` becomes a secret reference in the spec, usually surfaced as a file rather than an environment variable. ## What stays with the image - **The default command, in exec form.** `ENTRYPOINT ["/checkout"]` makes your process PID 1, so it receives `SIGTERM` directly instead of a shell swallowing it. - **The `USER`.** Declare a non-root uid in the image so every runtime inherits it, rather than relying on each caller passing `--user`. - **The filesystem and dependencies.** Nothing a scheduler sets can add a CA bundle, a timezone database or a shared library that is not in a layer. - **Behaviour under configuration.** Read settings from the environment with sane defaults; bind on all interfaces; write logs to stdout and stderr as an unbuffered stream; keep no state on local disk. - **Termination.** Handle `SIGTERM` by draining and exiting, because rolling replacement means the scheduler will send it routinely, not only during incidents. - **The architecture.** An `amd64`-only image will not run on an `arm64` node however the spec is written. ## The test to apply Ask of every value: *would this differ between two environments, or two machines?* If yes, it is runtime configuration and belongs in the spec. If no — it is genuinely a property of the artefact — it belongs in the image. Baking the production port, hostname or credential into a layer produces an image that runs in exactly one place, and turns a rollback into a change of more than the code.
- Does `EXPOSE 8143` in the Dockerfile publish anything?No. `EXPOSE` is metadata: it documents the port and gives `docker run -P` a list of ports to map to random host ports. Actual publishing is `-p` at run time, and a scheduler ignores `EXPOSE` except as a hint in tooling. An image with no `EXPOSE` at all still serves traffic perfectly well.
- Which settings can a scheduler *not* override, whatever its spec says?Anything already baked into layers or compiled in: installed packages and libraries, CA certificates, the binary itself and its compiled-in defaults, and the image's architecture. A runtime spec can change the command, the environment, the user and the resources, but it cannot add a file to the image or make an amd64 image run on arm64 hardware.
- If the runtime can pass `--user`, why still put `USER` in the Dockerfile?Because the image's default is what every caller inherits — a developer's `docker run`, a CI job, and the production spec. Declaring `USER 10001` means root is an explicit opt-in rather than the fallback. It also forces you to fix ownership of writable paths at build time instead of discovering it when one caller happens to drop privileges.
The image is the appliance; the run flags are where you plugged it in, what fuse you gave it and which shelf it sits on. Moving house replaces all of the second list and none of the first.
saying these in an interview costs you the question
- Thinks the restart policy is part of the image
- Believes EXPOSE publishes the port to the host
- Bakes the production hostname or port into a layer
- Assumes container names still resolve after the handoff
- Keeps --restart=always alongside an external supervisor
- Expects a bind mount to follow the container to another host