A Kubernetes golden-path template deploys a fraud-rules engine as a 7-replica Deployment; on a developer's single-node laptop cluster six pods stay Pending. What leaked, and how should the platform respond?
answer
- Pending: read the scheduler event
- a rule the developer never wrote
- one node, one pod per node
- spread degrades, required anti-affinity does not
- test the template on the dev shape
basics
~20 sThe template's production scheduling default leaked: required pod anti-affinity on kubernetes.io/hostname allows one replica per node, so one node runs one pod. The platform should make such defaults environment-aware, explain them, and keep failures debuggable.
solid answer
~40 sThe template almost certainly sets required pod anti-affinity on `kubernetes.io/hostname` so replicas spread across nodes in production. On a single-node laptop cluster only one replica fits; the other six stay `Pending` with a `FailedScheduling` event saying the node `didn't match pod anti-affinity rules`. That is an **abstraction leak**: a setting the developer never wrote, and may not understand, decides the outcome. The platform response has three parts. Fix the default: use `topologySpreadConstraints` with `maxSkew: 1`, or preferred anti-affinity, which degrade gracefully when only one node exists, or give the template a local profile. Make it visible: document each default, show the rendered YAML, and point developers to `kubectl describe pod`. And test it: run the template in CI against the same single-node cluster shape developers use.
code
yaml · 24 linesapiVersion: apps/v1
kind: Deployment
metadata:
name: fraud-rules-engine
spec:
replicas: 7
selector:
matchLabels:
app.kubernetes.io/name: fraud-rules-engine
template:
metadata:
labels:
app.kubernetes.io/name: fraud-rules-engine
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app.kubernetes.io/name: fraud-rules-engine
containers:
- name: engine
image: registry.example.com/fraud-rules-engine:4.12.3go deeper
Recall that a Pending pod's reason is in its events, and that a template can set scheduling rules you never wrote.
Explain why required anti-affinity on kubernetes.io/hostname caps a node at one matching pod, and how topology spread computes skew.
Diagnose from events and rendered output, fix the default once in the template, and add a CI test on the cluster shape developers really use.
Treat leaks as product defects: decide how defaults vary by environment, how they are documented, and how you measure where developers get stuck.
## The incident A team adopts the platform's golden-path template for its **fraud-rules engine**. The values file asks for 7 replicas, the same as production. A developer starts a **single-node development cluster on a laptop** (kind or minikube), applies the rendered output, and sees one pod `Running` and six `Pending`. The developer never wrote any scheduling rules, so there is nothing in *their* file to fix. ## Diagnosing it 1. `kubectl get pods -l app.kubernetes.io/name=fraud-rules-engine` shows 1 `Running`, 6 `Pending`. 2. `kubectl describe pod <a pending pod>` shows a `FailedScheduling` event. The scheduler builds its message from a `0/1 nodes are available` prefix and the plugin reasons, here `node(s) didn't match pod anti-affinity rules`. 3. Rendering the template and reading the pod spec reveals the cause: `affinity.podAntiAffinity.requiredDuringSchedulingIgnoredDuringExecution` with `topologyKey: kubernetes.io/hostname`. That rule means *no two pods with this label on the same node*. With one node, exactly one pod can ever be placed: 7 requested, 1 running, 6 pending. ## Why this is an abstraction leak A golden path promises that developers can describe *what* they want without knowing *how* Kubernetes realises it. A **leak** is any case where the hidden *how* decides the result and the developer must understand it to proceed. Leaks are unavoidable; the platform's job is to make them rare, loud and explainable. Typical leaks on this kind of platform: - **scheduling defaults** tuned for multi-node production, as here - **resource defaults** that fit a typical service but not a memory-heavy rules engine - **probe defaults** whose timings suit a fast-starting service but not one that loads a large rule set - **admission rejections** whose error names a field the developer never saw ## Fixing the default | Option | Single-node laptop | Multi-node production | |---|---|---| | Required pod anti-affinity on hostname | 1 of 7 scheduled | Strict one-per-node spread | | Preferred pod anti-affinity on hostname | All 7 scheduled | Best-effort spread | | `topologySpreadConstraints`, `maxSkew: 1`, `whenUnsatisfiable: DoNotSchedule`, no `minDomains` | All 7 scheduled, because one eligible node means skew is always 0 | Even spread, enforced | | Environment profile that sets replicas to 1 locally | 1 of 1 scheduled | Unchanged | Topology spread is often the better production default: it enforces an even distribution but, with a single eligible node, has nothing to be uneven against. Required anti-affinity is a hard cap of one pod per node and is right only when co-location is genuinely unsafe. An environment profile is also reasonable, but it means local runs no longer exercise the production shape. ## Making the leak survivable Changing one default fixes this incident; the platform should also reduce the cost of the *next* leak: - **Show the rendered output.** Developers should be able to render the template locally and read the real `Deployment`, or see it in the pull request. - **Document every non-obvious default** with the reason for it, next to the value that overrides it. - **Label rendered objects** so a developer can trace a field back to the template version that produced it. - **Teach the three commands** that explain most failures: `kubectl describe pod`, `kubectl get events -n <namespace>`, and `kubectl logs`. - **Test the template in CI** against a single-node cluster, the shape developers actually run, and assert that all replicas become ready. - **Offer an escape hatch**: a documented parameter to relax the spread rule, rather than forcing teams to fork the template. ## What not to do - Do not tell developers to edit the rendered YAML by hand: the next template upgrade overwrites it and the fork drifts. - Do not silently lower replicas everywhere; production still needs the spread. - Do not treat this as developer error. The developer followed the golden path; the path had a hole. ## The senior signal Interviewers want to hear that you **diagnose from events, not guesses**, that you know **why required anti-affinity and topology spread differ with one node**, and that you treat a leak as a **platform product defect** — fixed once in the template, tested, and explained — rather than a support ticket answered one team at a time.
- Why does a topologySpreadConstraints entry with maxSkew 1 still place all seven replicas on one node?Skew is the difference between the pod count in a topology domain and the minimum count across eligible domains. With one eligible node there is only one domain, so the minimum equals that node's count and skew stays 0 however many pods land there. Setting `minDomains` above 1 would change that and leave pods Pending.
- The same template later fails an admission check naming a field the developer never set. How should the platform handle that class of leak?Make the rejection explainable: the policy message should name the template value that controls the field and link to its documentation. Then fix the root cause in the template so the default passes policy, and add a CI check that renders the template and runs it through the same admission rules with a server-side dry run.
saying these in an interview costs you the question
- The developer misconfigured the Deployment, so it is their problem to fix.
- Six Pending pods on a laptop mean the node is simply out of CPU.
- Edit the rendered YAML by hand to remove the affinity rule.
- Topology spread and required anti-affinity behave the same on one node.
- Golden paths are only worth building if they never leak.