A memory ceiling ends a service seconds after every start, though it is stable in steady state — why?
answer
- the wall applies to the peak
- start-up is hungrier than serving
- parsing, warming, pools, decompression
- sized from the plateau, tested at the spike
- truncated log, no exception, repeatable point
basics
~20 sMost services demand more memory while initialising than while serving: configuration and data are loaded, caches are warmed, pools are opened, often concurrently. A ceiling sized from steady-state observation is below that transient peak, so the wall is hit before serving ever begins.
solid answer
~50 sThe ceiling applies to the **peak**, not the average, and for many workloads the peak happens at start-up rather than under load. Initialisation commonly reads and parses configuration and reference data, builds indexes or caches, opens connection pools, and decompresses assets — frequently several of those at once, because start-up is where concurrency is cheapest to spend. Steady-state serving then settles well below that high-water mark. If the ceiling was chosen by watching the service run for a week, it sits under a peak nobody measured, and the kernel ends the process the moment accounted memory crosses it. What you observe is a container that dies at the same point in every start and is restarted into the identical failure, so it looks like a code fault rather than a sizing one — the tell is that it never reaches the point of serving a request.
go deeper
Remember that a memory ceiling is judged against the highest point a workload reaches, and that a service is often hungriest while starting rather than while serving. A number chosen from normal running can be below the start-up spike.
Name what makes start-up expensive — parsing configuration and reference data, warming caches, allocating pools, decompressing, often in parallel — and explain why the ceiling applies to that peak with no throttle available to soften it.
Show the diagnosis: a repeatable early death, a log truncated mid-line with no exception, an abnormal termination recorded by the platform, and a ceiling that was sized from the plateau. Then measure the real peak from the container's first instant rather than guessing.
Make it a standard. Decide whether teams size ceilings from measured peaks, whether heavy preparation is allowed to run inside the serving process, and how much headroom above the peak is expected — otherwise every team rediscovers this with a production restart loop.
## Why a start-up peak exists at all A long-running service is usually at its hungriest before it has served anything. The work of coming up is different in kind from the work of serving: - Configuration and reference data are read and parsed into in-memory structures, and the parsed form is often several times the size of the file it came from. - Caches, lookup tables and indexes are built eagerly so the first requests are not slow. - Connection pools, buffers and arenas are allocated to their configured sizes whether or not traffic needs them yet. - Archives and assets are decompressed, which needs the compressed and uncompressed forms resident at the same time. - Much of this is done **concurrently**, because start-up is exactly where parallelism looks free — so several transient peaks coincide instead of queueing. Once serving begins, most of that transient material is released and the workload settles onto a plateau that can be a fraction of the high-water mark. ## The sizing mistake that follows A memory ceiling is a wall applied continuously, to whatever the container's accounted memory happens to be at that instant. It has no notion of "this is only start-up" and no throttle to fall back on. So a ceiling derived from a week of steady-state observation — the most natural thing to do — is a ceiling chosen from the plateau while the wall will actually be tested at the spike. The failure sequence is then completely deterministic: 1. The container starts and initialisation begins. 2. Accounted memory climbs past the plateau towards the initialisation peak. 3. Reclaim frees what it can — mostly cached file pages — and buys a little room. 4. Nothing reclaimable is left; the kernel ends a process inside the boundary. 5. The platform records an abnormal exit and starts the container again, which repeats the identical sequence. ## Why it is so often misdiagnosed The symptom is a service that dies within seconds of each start and never serves a request. That reads as a startup crash, and the investigation goes to the application's own logs — which are truncated mid-initialisation, because an out-of-memory kill gives nothing a chance to flush, and which contain no exception because the process was ended rather than having failed. The signal that separates the two is outside the application: an abnormal termination recorded by the platform, at a repeatable point, with no error written by the code. A genuine startup bug normally leaves a message; a ceiling does not. The CPU analogue is worth stating so the two are not confused. A CPU ceiling too small for initialisation does not end the process — it just makes the start slow, sometimes slow enough that a health check loses patience. Only the memory ceiling ends it. ## Fixing it | Approach | What it changes | When it is right | |---|---|---| | Raise the ceiling to cover the measured peak | The wall moves above the spike | Almost always the first move; it is the honest number | | Lower initialisation concurrency | The spike flattens into a longer, lower climb | When the peak is several parallel jobs overlapping | | Load lazily instead of eagerly | Transient structures are never all resident at once | When the eager work is cache warming rather than correctness | | Move one-off preparation out of the serving process | The peak is paid by something that then exits | When the heavy step is genuinely one-time setup | Whichever you pick, measure the peak rather than inferring it: watch accounted memory from the first instant of the container's life, not from when it starts serving, and size the ceiling above the highest point that sequence reaches with margin for a slower or larger start. ## What this leaves to other decisions Two neighbouring choices interact with this one and should not be conflated with it. The **reservation** does not change any of it — nothing about a reservation prevents a kill, because it is a placement input and not an enforcement. And a ceiling raised to cover a rare spike is capacity that the workload holds the right to use continuously, which is a density argument rather than a correctness one. Get the container to survive its own start first; argue about the cost of the headroom second. ## Where platforms differ Some platforms let a workload declare a separate, larger allowance for an initialisation phase, or run heavy preparation as a distinct short-lived step before the serving process begins, so the peak and the plateau do not have to share one number. Others offer only the single ceiling. Where the separation exists it is the cleanest fix; where it does not, the ceiling has to cover the peak.
- How would you tell this apart from an ordinary crash during initialisation?A code fault usually writes something — an exception, a failed dependency, a validation message — and exits through the application's own path. An out-of-memory kill leaves the log cut off mid-line with no error at all, and the platform records an abnormal termination rather than an application exit. Repeatability at the same point with a silent log is the tell.
- Does raising the reservation instead of the ceiling help a workload that dies during start-up?No. The reservation only influences which node the replica is placed on and how much capacity is claimed there; nothing enforces it on the process and nothing consults it when accounted memory crosses the wall. The ceiling is the enforced number, so the ceiling is the one that has to cover the start-up peak.
- If the peak is genuinely transient, is raising the ceiling wasteful?Less than it looks. A ceiling is a cap, not a claim — it is the reservation that holds capacity whether it is used or not. Raising the ceiling alone costs nothing until the workload actually uses the headroom, though it does mean a single replica can consume more of a node during its start. That trade belongs to a density discussion, not to getting the service to survive booting.
saying these in an interview costs you the question
- Assumes steady-state observation is enough to size a memory ceiling
- Reads a silent truncated log as an application crash
- Thinks raising the reservation will stop the kill
- Expects a graceful shutdown to be attempted before the kill
- Believes the peak must be below the ceiling only on average