skip to content

EC2 offers three placement group strategies — cluster, spread and partition. What does each one do to where instances physically land, and which kind of workload does each suit?

level: middleimportance: should knowfreq 40%

answer

  1. adjacency: asset or liability
  2. one AZ, lowest latency
  3. distinct hardware, small count
  4. seven is the recurring limit
  5. replica sets across partitions

basics

~20 s

Cluster packs instances close together in one Availability Zone for the lowest network latency. Spread puts each instance on distinct underlying hardware to avoid correlated failure. Partition groups instances into blocks that share no hardware between blocks, for replicated distributed systems.

solid answer

~50 s

A placement group tells EC2 **how to distribute instances across the physical fleet**, which is otherwise opaque to you. **Cluster** packs them tightly within a single Availability Zone so that instances sit close on the network — lowest latency and highest per-flow throughput, at the cost of a shared failure domain and a real chance of insufficient-capacity errors if you add instances later. **Spread** guarantees each instance sits on distinct underlying hardware, with a limit of seven running instances per Availability Zone per group; it suits a small number of critical peers such as domain controllers or a quorum of coordinators. **Partition** is the middle ground: up to seven partitions per Availability Zone, each on its own racks, with many instances per partition — this is what you want for HDFS, Cassandra or Kafka-style systems, where you can align replicas to partitions so no single rack failure takes out every copy. `aws ec2 describe-instances` reports the partition an instance landed in.

code

bash · 6 lines
bash
aws ec2 create-placement-group --group-name kafka-pg \
  --strategy partition --partition-count 5

aws ec2 describe-instances \
  --filters Name=placement-group-name,Values=kafka-pg \
  --query 'Reservations[].Instances[].{Id:InstanceId,Partition:Placement.PartitionNumber,AZ:Placement.AvailabilityZone}'

go deeper

for a junior

Know the three strategy names and the one-line purpose of each: cluster for low latency, spread for keeping instances off shared hardware, partition for large replicated systems.

for a middle

Explain the constraints that come with each — a cluster group's single Availability Zone and capacity risk, the seven-per-AZ limits on spread and partition — and say which workload shape each fits.

for a senior

Judge when a placement group earns its constraints at all, and design around the capacity risk: launch the full set in one request, keep types homogeneous, and align partitions to the data system's own replica placement.

for a principal

Own the availability argument. Decide when latency between peers is worth concentrating a workload in one Availability Zone, and make sure that choice is a deliberate, documented trade rather than something a template inherited.

## The problem placement groups solve Normally you have no say in which physical host, rack or network segment your instances land on — EC2 spreads them however it likes within the Availability Zone you chose. That is fine until physical adjacency starts to matter, and it matters in exactly two opposite directions: sometimes you want instances **as close together as possible** (latency), and sometimes you want them **as far apart as possible** (correlated failure). Placement groups are the control for both. A placement group is a free, zone-scoped construct you create and then reference at launch. ## Cluster: pack them together A cluster placement group asks EC2 to place instances close to each other on the network, inside a **single Availability Zone**. The result is the lowest inter-instance latency and the highest per-flow throughput EC2 offers, which is what HPC codes, tightly coupled simulations, and low-latency trading or distributed-training workloads need. The costs are real: - **One AZ means one failure domain.** A cluster group is architecturally the opposite of high availability. If the workload must survive an AZ event, a cluster group is the wrong tool. - **Capacity risk.** Packing many instances close together is a demanding request. The standard mitigation is to launch all the instances you need **in a single request**, so EC2 plans the placement once, and to use a homogeneous instance type — adding instances to an existing cluster group later is where `InsufficientInstanceCapacity` typically appears. To get the network benefit you also need instances whose types support enhanced networking, which effectively means current-generation types. ## Spread: keep them apart A spread placement group does the reverse: every instance is placed on **distinct underlying hardware** — different racks, with their own power and network — so that no single hardware failure can take out two of them. The trade is scale: a spread group permits at most **seven running instances per Availability Zone**. Spread it across three zones and you have twenty-one, which tells you what it is for. Use it for a small set of instances whose simultaneous loss would be catastrophic and which are not trivially replaceable: a quorum of consensus nodes, a pair of domain controllers, licence servers, a handful of stateful singletons. It is not a way to place a large web fleet. ## Partition: the middle ground A partition placement group divides the group into **up to seven partitions per Availability Zone**. Partitions share no underlying racks with each other, but a partition can hold as many instances as you like. So you get failure isolation *between* partitions with no cap on total fleet size. This is the shape that large replicated data systems want — HDFS, HBase, Cassandra, Kafka — because those systems already have a notion of replicas that must not share a failure domain. You place replica sets in different partitions and a rack-level failure loses one replica rather than all of them. EC2 tells you which partition an instance landed in, so topology-aware software can be configured to match: ```bash aws ec2 describe-instances \ --filters Name=placement-group-name,Values=kafka-pg \ --query 'Reservations[].Instances[].{Id:InstanceId,Partition:Placement.PartitionNumber,AZ:Placement.AvailabilityZone}' ``` You can let EC2 assign partitions or request a specific one at launch, which is how a broker or data node is pinned to the partition its rack-awareness configuration expects. ## Choosing between them The decision is a single question: **is adjacency an asset or a liability for this workload?** - Latency between peers dominates, and the workload can be rebuilt if the zone fails → **cluster**. - A handful of instances must never fail together → **spread**. - A large replicated system with its own replica placement rules → **partition**. - None of the above → **no placement group at all**, which is the right answer for most services. Ordinary stateless fleets behind a load balancer get their availability from spanning Availability Zones, and adding a placement group only introduces capacity constraints for no benefit. ## Practical notes Placement groups are per-region constructs referenced at launch; you cannot merge two groups, and moving a running instance into or out of a group requires it to be stopped. Mixing instance types inside a cluster group weakens the guarantee, because EC2 cannot always place dissimilar hardware close together. And none of this replaces multi-AZ design: a placement group shapes placement *within* the zones you already chose.

  • Why would launching extra instances into an existing cluster placement group fail with an insufficient-capacity error?
    A cluster group asks EC2 to place instances physically close together, so a later addition must find free capacity in that same tight area — which may no longer exist. The mitigations are to launch the full set in one request so placement is planned once, to use a single instance type, and to be willing to stop and relaunch the whole group if you must grow it substantially.
  • Can a placement group span Availability Zones?
    Cluster groups cannot — they are single-AZ by definition. Spread and partition groups can span zones within the region, and their limits are expressed per zone: seven running instances per AZ for spread, seven partitions per AZ for partition. That per-zone framing is what makes them usable in a multi-AZ design, whereas a cluster group is explicitly trading availability for latency.
  • How does a partition placement group help software that already has rack awareness?
    The instance's partition number is exposed through the EC2 API, so you can map partitions to whatever the software calls a rack or failure domain and configure replica placement accordingly. The data system then guarantees that copies of a shard live in different partitions, and because partitions share no underlying racks, a rack-level failure removes at most one copy.
  • When should you use no placement group at all?
    For most workloads. A stateless fleet behind a load balancer spread across Availability Zones already has the failure isolation it needs, and a placement group would only add capacity constraints and launch complexity. Reach for one when you have a measured latency requirement between peers, or a specific correlated-failure risk you can name.

saying these in an interview costs you the question

  • Thinks a cluster placement group improves availability
  • Believes a spread group can hold a large fleet
  • Says placement groups replace multi-AZ deployment
  • Assumes any instance type can join a cluster group
  • Uses a placement group by default for ordinary web servers

context