A process-isolated Windows container stops starting after the host is upgraded to a newer Windows build. How do you diagnose it?
answer
- The container never reached its entrypoint
- Something changed on the host, not the image
- Compare two version numbers, not one
- Process isolation binds image build to host build
- Hyper-V isolation is the bridge, not the fix
basics
~20 sProcess isolation requires the image's Windows build to match the host's, so an upgraded host rejects an image built on the old base with "The container operating system does not match the host operating system." Compare the image's OsVersion with the host build, then rebuild on the matching base tag or run it under Hyper-V isolation meanwhile.
solid answer
~50 sThe symptom is a container that never reaches its entrypoint, failing with `The container operating system does not match the host operating system.` Under process isolation the container's binaries run on the host's kernel, so the base image's build number must match the host's build; a host rebuilt from Windows Server 2019 (build `10.0.17763.x`) to Server 2022 (build `10.0.20348.x`) invalidates every `ltsc2019` image on it. Confirm both sides: `docker image inspect --format '{{.OsVersion}}' recon-batch:4.2` for the image, and the host's build number from the OS itself. The proper fix is to rebuild the image `FROM` the base tag matching the new host generation and redeploy. The immediate mitigation, while a payments reconciliation batch is waiting on its window, is `--isolation=hyperv`, which supplies the image's own kernel and ignores the host build entirely, at the cost of startup time and memory.
code
bash · 3 linesdocker image inspect --format '{{.Os}} {{.OsVersion}}' recon-batch:4.2
docker info --format '{{.OSType}} {{.OperatingSystem}}'
docker run --rm --isolation=hyperv recon-batch:4.2go deeper
Recall the rule behind the message: a process-isolated Windows container needs a base image whose Windows build matches the host's, so an upgraded host rejects images built on the previous generation.
Explain how you would prove it: read the image's OsVersion and the host's build number, note that only the revision component may differ, and describe why an empty docker logs points at creation rather than a crash.
Show the incident handling: bridge tonight's run with --isolation=hyperv, rebuild on the matching base tag, re-validate the runtime components inside the image, and keep host pools homogeneous per generation.
Own the process defect. Host-generation upgrades and image rebuilds must be one coupled change with a CI smoke run under production isolation, and the cost of supporting two Windows generations at once should be an explicit decision.
## The failure mode A nightly payments reconciliation batch runs in a Windows container built `FROM mcr.microsoft.com/windows/servercore:ltsc2019`. It has run for months. The host is rebuilt onto Windows Server 2022, and the very next run dies before producing a line of application output: ``` docker: Error response from daemon: hcsshim::CreateComputeSystem recon-batch: The container operating system does not match the host operating system. ``` The distinguishing feature is that there are no application logs at all, and `docker logs` is empty: the container never started, so nothing inside it ran. That separates this from a crash loop, from an entrypoint that exits, and from a mount or permissions problem, all of which produce at least some output or a nonzero exit code from your process. ## Why the rule exists Under process isolation the container's user-mode binaries execute directly against the host's Windows kernel. Windows does not offer a stable syscall ABI across builds the way Linux does, so a `servercore` or `nanoserver` image is built against a specific kernel build and is only supported on that build. The engine enforces this at container creation rather than letting the binaries fail unpredictably later. Concretely, the base image tags map to builds: - `ltsc2019` images carry `OsVersion` `10.0.17763.x` and process-isolate on Windows Server 2019 hosts. - `ltsc2022` images carry `OsVersion` `10.0.20348.x` and process-isolate on Windows Server 2022 hosts. Only the fourth component, the revision, is allowed to differ; on Windows Server 2019 and later a host patched to a higher revision than the image was built against still runs it. A different build number is a hard stop. ## Diagnosing it in order First establish the two numbers rather than guessing: ``` docker image inspect --format '{{.Os}} {{.OsVersion}}' recon-batch:4.2 docker info --format '{{.OSType}} {{.OperatingSystem}}' ``` The first prints something like `windows 10.0.17763.4737`; the second identifies the host. Read the host's own build number from the operating system as well, so you are not relying on a marketing name. If the two build numbers differ in any of the first three components, you have your answer and no further investigation is warranted. Second, confirm the mode actually in use. `docker inspect --format '{{.HostConfig.Isolation}}'` on a previous, successful container tells you whether the workload was ever relying on process isolation. If a developer's machine ran it under Hyper-V isolation, the mismatch was invisible there, which is often why the problem surfaces only in production. Third, check whether anything else changed with the host rebuild before declaring victory: a rebuilt host is also a new data root, a fresh image cache and possibly a different engine version. The distinctive error message keeps you honest here; if it is the OS mismatch text, the build numbers are the story. ## Fixing it, immediately and properly The immediate mitigation is `docker run --isolation=hyperv`. The container gets its own kernel from the image, so the host's build stops mattering and the batch can run tonight. Budget for a slower start and for the memory the extra kernel takes; for a single nightly batch that is irrelevant, for a fleet of long-running services it is a capacity change you should not make casually or permanently. The proper fix is to rebuild the image on the base tag matching the new host generation, `FROM mcr.microsoft.com/windows/servercore:ltsc2022`, and re-test. This is rarely a one-line change in practice: the newer base carries a newer set of OS components, and anything the application depends on inside the image, a runtime, an ODBC driver, a scheduled-task shim, has to be validated on the new base too. ## Preventing the recurrence The deeper defect is that a host upgrade and an image rebuild were treated as independent changes. Three habits stop it repeating. Pin explicitly. Never build a Windows workload `FROM` a floating tag that can move across generations; name the generation in the tag, and record the expected host build alongside it in the deployment manifest. Couple the two changes. Treat a host-generation upgrade as a fleet event that requires the images to be rebuilt and re-tested first, then rolled with the hosts, rather than a patch that operations can apply independently. Keep the host pool homogeneous per generation so a workload cannot land on a host of the wrong build by scheduling luck. Make the environments agree. If developers run under Hyper-V isolation by default and production runs process isolation, add an explicit `--isolation=process` smoke run in CI on a host of the production generation, so a mismatch is caught by a pipeline rather than by a payments batch missing its window. ## Summary The error is a version-compatibility check, not a bug. Process isolation binds the image's build to the host's build; a host generation upgrade therefore invalidates images built on the previous generation. Diagnose by comparing `OsVersion` with the host build, bridge with Hyper-V isolation if you need tonight's run, and fix by rebuilding on the matching base and coupling image rebuilds to host upgrades from then on.
- How do you tell this apart from a container that starts and then exits immediately?By where the failure appears. A build mismatch fails at container creation with `The container operating system does not match the host operating system.`, `docker logs` is empty and no exit code from your process exists. A container that started and died has an exit code in `docker inspect`, usually some log output, and a `died` event carrying that code.
- Would running the image under Hyper-V isolation permanently be an acceptable answer?As a bridge yes, as a policy rarely. Each container then carries its own kernel, so start-up latency and per-container memory rise and host density falls. For one nightly batch that is a fine trade; for a fleet of services it is a capacity decision, and it also lets the image drift further behind the host generation, which makes the eventual rebuild harder.
- The host was only patched, not upgraded, and the container still runs. Why?Patching moves the revision, the fourth component of the version, and process isolation on Windows Server 2019 and later tolerates a revision difference between image and host. Only a change in the build number, the third component, breaks the match. That is exactly why a monthly patch is safe and a generation upgrade is not.
saying these in an interview costs you the question
- Blames the entrypoint or a missing dependency
- Thinks any Windows image runs on any Windows host
- Expects the daemon to fall back to Hyper-V automatically
- Confuses the revision number with the build number
- Treats Hyper-V isolation as the permanent fix
- Upgrades hosts without rebuilding the images