Two workloads on one host want different values for a host-wide kernel tunable — what can the container boundary do about that?
answer
- two classes of setting, not one
- does the kernel scope it per view?
- network-scoped values can differ per container
- one machine-wide number for every tenant
- conflict settled by placement, not configuration
basics
~20 sNothing, for a value the kernel keeps per machine. Some settings follow a container's own view of a resource and can differ per workload; a single machine-wide value is one number for everyone, so the conflict is settled by placement, not by configuration.
solid answer
~50 sTunables fall into two classes. Some are scoped to the same per-resource view a container already has — network-related ones such as a listen backlog ceiling or the ephemeral port range typically follow the container's own network view, so each workload can carry its own value. The rest are a single value for the whole machine: the memory overcommit policy, a ceiling on filesystem watch registrations, and similar. Which class a given setting is in is a property of the kernel, not something the platform can decide, so verify rather than assume. For the machine-wide class there is no per-container answer at all. You either agree one value every workload on that host can live with, or you stop co-locating the two and give the demanding one its own host population, where the value becomes part of that host's definition.
code
pseudocode · 14 linesfor each setting the vendor asks for:
scope = how the kernel scopes this setting
if scope is per-container-view-of-the-resource:
apply it to this workload only
neighbours keep their own value
else if every workload already on this host tolerates the new value:
apply it once to the host
record it as part of the host definition
else:
place this workload on a host population built with that value
keep the requirement as a placement constraintgo deeper
Know that not every setting a container depends on is its own: some knobs belong to the machine, and changing one of those changes it for every workload there.
Explain the split — settings the kernel scopes to a container's own view can differ per workload, a single machine-wide value cannot — and say that which class a knob is in is the kernel's decision.
Show how you would resolve the conflict in practice: verify the scope, agree one tolerable value or separate the workloads onto distinct host populations, and record the value so a rebuilt host comes up with it.
Own the policy. Every extra host population is more configuration to certify, patch and keep placed, so decide when one tenant's tuning is worth fragmenting the fleet and when the workload is the thing that should change.
## Two classes of setting A container gets separate views of several resources — its own filesystem, process table, network, ids and hostname — and an accounting group that caps what it consumes. Kernel tunables interact with that split in exactly one way: **a tunable can differ between two containers only if the kernel scopes it to a view the containers already have separately.** That gives two classes: | Class | Typical examples | Can two tenants differ? | Where the value belongs | |---|---|---|---| | Scoped to the container's own view of a resource | listen backlog ceiling, ephemeral port range | yes, per workload | the workload's declaration, where the platform permits it | | One value for the whole machine | memory overcommit policy, ceiling on filesystem watch registrations | no | the host's provisioning | The crucial point is that **membership of those classes is the kernel's decision, not the platform's**. No amount of boundary work promotes a machine-wide value into a per-container one, and a platform that lets you write such a setting into a workload's declaration is either applying it to the host or refusing it, never giving each tenant a private copy. Check the scope of the specific setting before designing around it; the two classes are not distinguishable by intuition, which is why platforms usually gate which settings a workload may touch at all. ## Why the machine-wide class is awkward A machine-wide value has three uncomfortable properties: - **It outlives the workload that wanted it.** If it were carried by whichever container happened to ask, the host's behaviour would drift as workloads came and went, and a replacement host would come up different. - **It applies to the neighbours.** Raising a ceiling for one tenant raises it for every other tenant on that machine, including the ones whose capacity planning assumed the old number. - **It is invisible from inside.** A workload's own declaration says nothing about it, so the same workload can behave differently on two machines for reasons its owners cannot see in their own configuration. There is a fourth, subtler one: a shared ceiling is also **exhaustible**. A global allowance — registrations, ports, table entries — is drawn down by all the workloads on the machine together, so one greedy tenant can consume most of it and leave a quiet neighbour failing against a limit nobody changed. The symptom then appears in the wrong place entirely, and the only thing the failures have in common is the host. ## Resolving the conflict 1. **Confirm the scope** of the exact setting being argued about. Half of these disputes evaporate here, because the setting turns out to follow the container's own view and both tenants can simply have their own value. 2. **If it is machine-wide, look for one tolerable value.** Test the other workloads against the proposed number rather than assuming indifference; the tenant who did not ask for the change is the one who discovers it. 3. **If no shared value exists, separate the workloads.** The demanding one moves to a host population where the value is part of how those hosts are built, and the requirement becomes a placement constraint that somebody maintains. 4. **Record it where hosts are built**, so a rebuilt or replaced machine comes up with the same value. A tunable applied by hand once is a machine that silently regresses on its next restart. 5. **Re-review it like a host change**, because that is what it is — not a workload setting that happens to live somewhere unusual. ## What this tells you about the boundary The container boundary is a set of partitioned views plus an accounting group. It is not a machine, and it does not virtualise the kernel's own configuration. Everything the kernel holds as one value for itself — its version, its loaded modules, its machine-wide tunables — is shared, whatever the workload declarations look like. Density and multi-tenancy arguments that ignore this class of setting tend to fail late, when a tenant's tuning request turns out to be a request to retune everyone. The practical discipline is simply to treat these values as **host properties with tenants attached**: scope verified, value reviewed, change announced, and placement used as the tool of last resort when two tenants genuinely cannot share one number.
- Why does a machine-wide value belong in the host's definition rather than in the workload's spec?Because it outlives the workload and applies to its neighbours. Carried by whichever container asked for it, the host's behaviour would change as workloads come and go, and a replacement machine would come up different. Building it into how the host is provisioned makes it reproducible, reviewable, and true of the host even when the requesting workload has moved on.
- The shared machine-wide ceiling is high enough for everyone. Where is the remaining risk?It is shared, so it is also exhaustible. One workload can draw down most of a global allowance and leave a neighbour failing against a limit that neighbour never changed. The symptom surfaces in the wrong place, and the only thing the failures have in common is the machine they are running on.
saying these in an interview costs you the question
- Assumes any kernel setting can be made per container by the platform.
- Thinks the container boundary partitions machine-wide kernel values.
- Puts a host-wide value in a workload spec and expects isolation.
- Forgets that raising a global ceiling for one tenant raises it for all.
- Treats a shared global allowance as impossible for one workload to exhaust.