skip to content

In an AWS Glue Spark job, what does moving from G.1X to G.2X workers change?

level: seniorimportance: should knowfreq 45%

answer

  1. bigger workers, not more of them
  2. the unit is bundled CPU and memory
  3. doubling one, halving the other, is a wash
  4. OOM and spill point one way, slow stages another
  5. the count can change during the run

basics

~20 s

G.2X gives each worker twice the DPU of G.1X — double the vCPUs, memory and disk per worker — so it raises per-executor capacity rather than total parallelism. Halving the worker count when you double the worker type keeps total DPUs, and cost, roughly flat.

solid answer

~50 s

AWS Glue sizes a Spark job as a **worker type** times a **number of workers**. A DPU (Data Processing Unit) is Glue's bundled unit of vCPU and memory; a **G.1X** worker is 1 DPU and a **G.2X** worker is 2 DPUs, so G.2X doubles the vCPUs, memory and attached disk that each worker gets. Larger types (G.4X, G.8X) exist for memory-hungry jobs, and G.025X is a small type intended for low-volume streaming. Moving up the worker type is the lever for per-executor pressure — heap OOMs, heavy spill to disk, a large broadcast, one enormous partition. It is *not* the lever for "the job is slow because there isn't enough parallelism"; that is more workers. Because you are billed on DPU-hours, doubling worker type and halving worker count leaves capacity roughly unchanged, which is exactly how you trade parallelism for per-task headroom. Enabling `--enable-auto-scaling` lets Glue vary the executor count during the run up to your configured maximum.

code

bash · 11 lines
bash
aws glue update-job --job-name orders-etl --job-update '{
  "Role": "GlueETLRole",
  "Command": {"Name": "glueetl", "ScriptLocation": "s3://scripts/orders.py"},
  "WorkerType": "G.2X",
  "NumberOfWorkers": 10,
  "DefaultArguments": {
    "--enable-auto-scaling": "true",
    "--enable-metrics": "true",
    "--enable-spark-ui": "true"
  }
}'

go deeper

for a junior

Recall that a Glue Spark job is sized as a worker type plus a worker count, and that G.2X workers are twice the size of G.1X ones.

for a middle

Explain the DPU as the bundled unit, why doubling the type and halving the count is capacity-neutral, and what --enable-auto-scaling varies during a run.

for a senior

Match the lever to the symptom from real evidence — OOM and spill versus insufficient parallelism versus skew — and argue in DPU-hours rather than wall-clock time.

for a principal

Own how sizing decisions are made and revisited across a fleet of jobs: what gets measured, who reviews it as data grows, and where FLEX and auto scaling belong in the platform's cost model.

## The sizing model A Glue Spark job's capacity is two numbers: which **worker type** and how many **workers**. Everything else — executors, cores per executor, memory — follows from that; you do not hand-tune `--num-executors` the way you would with `spark-submit`. The unit underneath is the **DPU**, a Data Processing Unit, which bundles vCPU and memory together: - **G.1X** — 1 DPU per worker. - **G.2X** — 2 DPUs per worker: double the vCPUs, double the memory, more attached disk. - **G.4X / G.8X** — larger still, for jobs whose per-task working set genuinely needs it. - **G.025X** — a fractional worker intended for low-volume streaming jobs. Billing is by DPU-hours consumed, so `worker_type_dpus × number_of_workers × duration` is the shape of the cost. That is why the substitution matters: 10 × G.2X and 20 × G.1X are the same capacity purchase with very different runtime characteristics. Note also that the capacity you request includes the Spark driver, not only executors. ## What going up a worker type actually fixes More DPU per worker means a bigger executor: more heap for each task's working set, more cores sharing that heap, and more local disk for shuffle and spill. It is the right move when the symptom is **per-executor pressure**: - Executors dying with out-of-memory errors. - Heavy spill to disk during shuffles or sorts. - A broadcast side that is too big to fit comfortably. - Genuine skew where one partition is far larger than the rest and simply needs somewhere to live. It is the wrong move when the symptom is **not enough parallelism** — the job is CPU-bound across many well-balanced tasks and the cluster is simply too small. There, more workers is the answer, and moving to G.2X while halving the count makes it worse. It is also the wrong move for the most common Glue slowness of all: a source split into tens of thousands of tiny objects, where the job spends its time listing and opening files rather than computing. No worker size fixes that; compaction upstream does. ## Auto scaling Setting the job parameter `--enable-auto-scaling` to `true` (available from Glue 3.0 onward) lets Glue add and remove executors during the run based on demand, up to the maximum number of workers you configured. Since you are billed for what is actually consumed, this suits jobs whose shape varies run to run — a backfill that reads a month one day and an hour the next, or a pipeline with a heavy stage followed by a light one. It does not remove the need to pick a sensible worker type: auto scaling changes *how many* executors exist, not *how big* each one is. For jobs with no latency requirement, Glue also offers a **FLEX** execution class, which runs on spare capacity in exchange for less predictable start and run times. It is appropriate for backfills and non-urgent nightly work, and inappropriate for anything with an SLA attached. ## How to size in practice Start from evidence, not from a default: 1. Turn on the job's metrics and the Spark UI (`--enable-metrics`, `--enable-spark-ui`) so you can see executor memory, task duration distribution and shuffle spill instead of guessing. 2. Diagnose the symptom class. OOM and spill → bigger worker type. Uniformly long stages with all executors busy → more workers. Long tail of one or two tasks → skew, which needs a data fix (salting, a different partition key, a broadcast) more than a hardware fix. 3. Change one variable at a time and compare DPU-hours, not wall clock — a job that finishes in half the time on twice the capacity has not saved anything. 4. Re-measure after the data grows. Glue sizing set once during a proof of concept and never revisited is where a large share of a data platform's bill quietly accumulates. ## What interviewers are checking That you understand a worker type is a per-executor size, not a throughput dial; that you can express the cost identity (type × count × duration) and therefore know the neutral substitution; and that you reach for measurement before capacity. The weak answer is "the job was slow so we moved to G.2X" with no account of what was actually constrained.

  • What does enabling auto scaling on a Glue job actually vary?
    The `--enable-auto-scaling` parameter (Glue 3.0 and later) lets Glue add and remove executors during the run according to demand, up to the maximum number of workers configured on the job. It varies how many executors exist, never how large each one is — the worker type is still a fixed decision you make before the run starts.
  • A Glue job on 20 G.1X workers is slow because one task runs for an hour while the rest finish in minutes. Does G.2X help?
    Mostly no. That profile is skew: one partition holds far more data than the others, so a bigger executor may keep it from spilling but the stage still waits on a single task. The fix is in the data — repartition on a better key, salt the hot key, or broadcast the small side so the skewed join disappears.
  • When would you use Glue's FLEX execution class?
    For runs where completion time does not matter: backfills, reprocessing, exploratory or nightly non-urgent jobs. FLEX uses spare capacity, so start and run times are less predictable in exchange for a cheaper execution class. Anything with an SLA, a downstream consumer waiting on it, or a tight schedule window should stay on the standard class.

saying these in an interview costs you the question

  • Treating worker type as a general speed dial
  • Believing G.2X doubles the number of executors
  • Ignoring that requested capacity includes the driver
  • Expecting a bigger worker type to fix data skew
  • Sizing once during a POC and never re-measuring

context