skip to content

Your team wants to move an EC2 fleet from x86 instances to AWS Graviton types such as m7g. What actually has to change for the workload to run, and how would you validate the move before committing?

level: seniorimportance: should knowfreq 48%

answer

  1. arm64, not a resize
  2. architecture-specific images
  3. native extensions are the blocker
  4. one tag, two architectures
  5. measure cost per request

basics

~20 s

Graviton instances are arm64, so everything compiled has to change: an arm64 AMI, arm64 runtimes, recompiled native code and multi-architecture container images. Interpreted code usually moves untouched; native dependencies and third-party agents are where migrations stall.

solid answer

~50 s

Graviton is AWS's own arm64 processor, so the migration is an **instruction-set** change, not a resize. Concretely: boot an arm64 AMI, use arm64 builds of the runtime, rebuild anything with native code, and publish **multi-architecture container images** so the same tag runs on both. Managed runtimes — the JVM, Python, Node, Go — all have arm64 builds, so pure application code usually moves untouched; the friction is native extensions, JNI libraries, prebuilt wheels, vendored binaries and third-party monitoring or security agents that ship x86-only. Validate by benchmarking your own workload rather than trusting a headline number: run a canary alongside the x86 fleet on identical traffic and compare latency percentiles, throughput per instance and cost per request. Migrate the stateless tier first, keep the build pipeline emitting both architectures, and keep the option to roll back to the x86 type.

code

bash · 4 lines
bash
docker buildx build \
  --platform linux/amd64,linux/arm64 \
  -t 111122223333.dkr.ecr.eu-west-1.amazonaws.com/api:1.4.2 \
  --push .

go deeper

for a junior

Know that Graviton instances use arm64 rather than x86, and that images and compiled binaries must match the architecture of the instance you launch on.

for a middle

Explain which layers are architecture-specific — image, runtime, native extensions, agents — and how a multi-architecture container image lets one tag serve both fleets.

for a senior

Lay out a migration you would actually run: inventory the compiled dependencies, canary on live traffic behind the same load balancer, compare cost per request and p99, and keep rollback cheap.

for a principal

Decide whether the programme is worth funding at all: weigh fleet-wide savings against build-pipeline and dependency-audit cost, set the policy for new services defaulting to arm64, and own the vendor conversations for agents that lack arm64 builds.

## What Graviton is Graviton is a family of processors AWS designs itself, used by instance types carrying a `g` after the generation digit: `m7g`, `c7g`, `r8g`, `t4g`. They implement the **arm64** (also written aarch64) instruction set rather than x86_64. AWS positions them as offering better price-performance than comparable x86 types, and for many workloads they do — but the number that matters is the one you measure on your own traffic, not the one in a launch blog post. Because the change is architectural, the migration question is really a **software supply-chain** question: what in your stack is compiled, and does an arm64 build of it exist? ## The layers that must change **1. The machine image.** AMIs are architecture-specific. An x86_64 AMI id will simply fail to launch on an m7g type. Every place that names an image — launch templates, image lookups, base images in a Dockerfile — needs an arm64 equivalent. **2. The runtime.** The JVM, CPython, Node.js, Go and .NET all ship arm64 builds, and code written in those languages is normally architecture-neutral. Go and Rust cross-compile with a target flag. This layer is rarely the problem. **3. Native code.** This is where migrations actually stall: - Python packages that ship prebuilt wheels only for `manylinux_x86_64` fall back to compiling from source, which needs a toolchain and may fail. - Node modules built through `node-gyp`, and any JNI library loaded by a JVM service. - Vendored static binaries: CLI tools baked into an image, database clients, media codecs. - Anything with hand-written assembly or CPU-specific intrinsics. **4. Third-party agents.** APM agents, security and EDR products, log shippers and licence daemons all have to publish arm64 builds. Vendor coverage is now broad but not universal, and this is the single most common blocker outside your own code — check it before you plan anything else. ## Containers: build multi-arch If the fleet runs containers, the right answer is a multi-architecture image published under one tag, so the platform picks the correct variant at pull time: ```bash docker buildx build --platform linux/amd64,linux/arm64 \ -t 111122223333.dkr.ecr.eu-west-1.amazonaws.com/api:1.4.2 --push . ``` This keeps deployment manifests architecture-agnostic and makes rollback trivial: the same tag still runs on x86 hosts. The build pipeline needs to produce both, which usually means either arm64 build machines or emulated builds — emulation is correct but slow, so native arm64 builders are the practical choice for anything with a compile step. ## Validating the move A credible plan has four parts: 1. **Inventory** — enumerate every compiled artefact and agent, and confirm an arm64 build exists. This is desk work and it prevents most surprises. 2. **Canary** — run a small arm64 fleet behind the same load balancer as the x86 fleet, taking real traffic. Do not benchmark synthetically first; production traffic exposes the pathological cases. 3. **Compare the right numbers** — p50/p95/p99 latency, throughput per instance, error rate, and **cost per request** rather than cost per instance. A type that is cheaper per hour but slower per request may be a loss. 4. **Watch for the non-obvious** — different cache sizes and memory bandwidth can shift GC behaviour or thread-pool tuning, and workloads that lean on specific x86 vector instructions can regress. Crypto and compression paths are worth checking explicitly. ## Sequencing Move the **stateless** tier first: it is easy to roll back and it carries most of the fleet cost. Keep both architectures buildable for the whole transition rather than cutting over the pipeline. If the fleet is managed as a group, mixing architectures within one group is awkward because the launch template pins the image — running two groups behind the same load-balancer target group is the cleaner shape while you compare. ## The judgment an interviewer wants They are listening for you to say that this is not free. The savings are real and often substantial for large stateless fleets, but you pay in build-system complexity and dependency auditing. For a small fleet, the engineering time can exceed the saving. Naming that threshold — and naming the agent-compatibility check as the first thing you would do — is what distinguishes a considered answer from an enthusiastic one.

  • How do you keep container deployments architecture-agnostic across the transition?
    Publish a multi-architecture image under a single tag, built with `docker buildx build --platform linux/amd64,linux/arm64`. The registry stores a manifest list and each host pulls the variant it can run, so deployment manifests never mention an architecture and rollback to x86 hosts needs no rebuild.
  • Which measurement decides whether the migration was worth it?
    Cost per unit of work — cost per request or per job — not cost per instance-hour. A Graviton type with a lower hourly price that serves fewer requests per second at your latency target is not a saving. Compare against the x86 fleet under identical live traffic and include p99 latency, since a regression there can cost more than the instance savings.
  • What in a JVM service is most likely to break on Graviton?
    Not the bytecode — the JVM itself has solid arm64 builds. The risks are JNI libraries and any dependency bundling a native `.so`, plus agents attached at startup such as APM or profiling agents that ship x86-only. Memory-model differences also mean concurrency bugs that happened to be masked on x86's stronger ordering can surface on arm64.

saying these in an interview costs you the question

  • Calls it just a cheaper instance type, no changes needed
  • Thinks the same AMI id boots on both architectures
  • Assumes a Docker image runs on any CPU architecture
  • Trusts the advertised price-performance number without benchmarking
  • Forgets third-party monitoring and security agents

context