A latency-sensitive service and its in-cluster cache should run close together, while the service's own replicas must stay spread for availability. How would you design that placement with Kubernetes affinity rules, and what would you refuse to guarantee?
answer
- podAffinity pulls, podAntiAffinity pushes — same term shape
- zone for co-location, hostname for spreading
- hard affinity = cross-Deployment placement dependency + bootstrap deadlock
- autoscaler cannot fix hard podAffinity with an empty node
- need a guarantee? DaemonSet or sidecar, not affinity
basics
~20 sUse soft podAffinity from the service to the cache at zone topology to cut cross-zone latency and cost, plus podAntiAffinity on the service's own label for spreading. Keep both soft — hard co-location plus hard spreading over-constrains the scheduler and produces Pending Pods.
solid answer
~50 sTwo rules on the same Pod spec, pulling in opposite directions: - `podAffinity` selecting `app=cache` with `topologyKey: topology.kubernetes.io/zone`, **preferred** with a high weight. Zone-level co-location removes cross-zone network hops and inter-AZ data charges without pinning both workloads to one machine. - `podAntiAffinity` selecting the service's own label with `topologyKey: kubernetes.io/hostname`, also **preferred**, so replicas spread across machines. Both soft is the key decision. Hard co-location makes the service's schedulability depend on another Deployment's placement and creation order — a cold cluster where the cache has not yet scheduled leaves the service Pending. Hard spreading caps replicas at the domain count. Together, hard-and-hard can be unsatisfiable at any replica count. What I would not guarantee: that co-location survives. Affinity is `IgnoredDuringExecution`, so a cache Pod rescheduling into another zone silently breaks the pairing with no signal. If the locality is load-bearing, run the cache as a DaemonSet or sidecar instead of hoping the scheduler keeps them together.
code
yaml · 18 linesspec:
affinity:
podAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 80
podAffinityTerm:
labelSelector:
matchLabels:
app: cache
topologyKey: topology.kubernetes.io/zone
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: api
topologyKey: kubernetes.io/hostnamego deeper
Know the two rule types exist and that podAffinity pulls Pods together while podAntiAffinity pushes them apart.
Write both rules correctly, pick zone for co-location and hostname for spreading, and explain why required plus required is likely unschedulable.
Cover the bootstrap dependency of hard podAffinity, autoscaler behaviour, and diagnosing Pending Pods when two rules conflict.
Frame it as best-effort optimisation versus architectural invariant: soft rules for cost and latency, DaemonSet or sidecar when a guarantee is needed, and a policy view on letting one team's manifest couple to another's placement.
## The two forces This is the canonical affinity design question because the two rules genuinely conflict. Co-location pulls Pods together; spreading pushes them apart. The scheduler resolves both in the same cycle, and every hard constraint you add shrinks the feasible set multiplicatively. ## Rule one: co-location with podAffinity `podAffinity` attracts a Pod towards topology domains that already contain matching Pods. The term is the same shape as anti-affinity: `labelSelector` (which Pods to be near), namespace scoping, and `topologyKey` (how close "near" means). Choosing the topologyKey is the whole design: - `kubernetes.io/hostname` — same machine. Lowest latency, loopback-ish networking, but it packs both workloads into one failure domain and makes the pair a single unit of loss. Almost never the right choice for a shared cache. - `topology.kubernetes.io/zone` — same availability zone. Removes the cross-AZ round trip (typically the difference between sub-millisecond and low single-digit milliseconds) and, on public clouds, removes per-GB inter-AZ transfer charges, which on a chatty cache path is frequently the larger win. Keeps the two workloads on different machines. - `topology.kubernetes.io/region` — usually meaningless inside one cluster. Zone is the answer for a cache. State it as a cost-and-latency decision, not a latency decision alone. ## Rule two: spreading with podAntiAffinity Select the service's own Pod labels with `kubernetes.io/hostname` so replicas do not share a machine. Note the asymmetry in topologyKey between the two rules: zone-level attraction, hostname-level repulsion. That combination is satisfiable — several replicas in one zone, on different nodes, all near a cache instance in that zone — whereas zone-level repulsion plus zone-level attraction to a single-replica cache is contradictory. ## Why both should be soft A hard podAffinity creates a **placement dependency between two independent Deployments**. Consequences: - **Ordering.** On a fresh cluster the cache may not be scheduled yet; the service has nothing to be near and stays Pending. Anti-affinity has no such bootstrap problem, affinity does. There is no dependency ordering in Kubernetes to fix this. - **Cascade.** If the cache is evicted and rescheduled elsewhere, new service Pods follow it, but existing ones do not move — you get a split population with no alarm. - **Autoscaling interference.** Node autoscaling simulates whether adding a node fixes a Pending Pod. A hard podAffinity is often *not* fixed by an empty new node, because the new node's domain contains no cache Pod, so the autoscaler correctly concludes scale-up will not help and leaves the Pod Pending. Soft rules degrade instead: worse locality under pressure, still running. ## What you refuse to guarantee The honest boundary is that affinity is a **scheduling-time hint with no runtime enforcement**. Every form ends in `IgnoredDuringExecution`. Nothing re-evaluates after placement, nothing emits an event when locality is lost, and no controller repairs it. So: - Do not promise that co-located Pods stay co-located. - Do not build a correctness property on locality — for example, do not let the service assume the cache in its zone holds its data. - Do not treat it as a cost control you can report on; if inter-AZ spend matters, measure it, do not assume the rule held. If locality must be guaranteed, change the topology instead of the constraint: run the cache as a **DaemonSet** so every node has one, or as a **sidecar** in the same Pod so the pairing is structural. Both remove the scheduler from the guarantee entirely. That is usually the better principal-level answer — replace a soft scheduling wish with an architectural invariant. ## Second-order concerns - **Scheduling cost.** Two inter-Pod terms per Pod, each a scan over matching Pods, on every scheduling attempt including every rollout. On a large cluster this is a real budget item; keep both selectors narrow. - **Blast radius.** Zone-level co-location deliberately concentrates a request path inside one zone. That is fine when zone loss means losing that zone's replicas anyway, and bad if the cache is a single replica whose zone loss takes out every dependent. - **Restore path.** After a zone outage, rules that were soft let everything reschedule into the surviving zones; hard rules can leave you unable to recover capacity at exactly the wrong moment. Soft rules are also a disaster-recovery decision. - **Ownership.** Cross-Deployment affinity couples two teams' release schedules through the scheduler. Whoever owns the platform should decide whether that coupling is allowed to be expressed in a manifest at all. ## How to present it Lead with the tradeoff, not the YAML: soft-and-soft for a best-effort optimisation, structural change (DaemonSet or sidecar) when it is a requirement, and hard rules only for correctness properties such as quorum separation where being Pending genuinely beats being wrong.
- Why can node autoscaling rescue a Pod blocked by hard anti-affinity but usually not one blocked by hard podAffinity?The autoscaler simulates adding a template node and re-runs scheduling. A brand-new empty node satisfies anti-affinity by definition, because it hosts none of the repelling Pods. For podAffinity the new node's topology domain contains no attracting Pod either, so the simulation still fails and the autoscaler correctly declines to add a node it knows would not help.
- If zone-level co-location with a cache is a hard requirement rather than an optimisation, what would you change?Stop expressing it as a scheduling preference. Run the cache as a DaemonSet so every node has a local instance, or as a sidecar container in the same Pod so co-location is structural and cannot drift. Alternatively route through a zone-aware Service so the client picks a same-zone endpoint at request time instead of relying on where Pods were placed.
Affinity is asking colleagues to sit near each other on day one. Nobody re-seats them when someone moves desks. If two people must always sit together, you do not repeat the request — you give them one shared desk.
saying these in an interview costs you the question
- Making both the co-location and the spreading rules required, then being surprised by unschedulable Pods.
- Believing the scheduler will move Pods later to restore co-location — nothing re-evaluates affinity after placement.
- Using hostname-level podAffinity for a cache, which concentrates the whole path onto one machine.
- Ignoring that hard podAffinity depends on another Deployment already being scheduled, creating a bootstrap deadlock.
- Claiming zone co-location is purely a latency optimisation and missing the inter-AZ data transfer cost, which is often the bigger number.