In AWS Glue, when would you run a Python shell job instead of a Spark ETL job?
answer
- not everything needs a cluster
- one process versus many
- Glue has more than one job type
- a fraction of a DPU, no Spark
- DynamicFrames and bookmarks need Spark
basics
~20 sA Glue Python shell job runs one ordinary Python process with no Spark cluster behind it, so it fits small single-node work: calling an API, issuing SQL to a warehouse, moving a few megabytes. Spark ETL jobs are for distributed volumes.
solid answer
~50 sAWS Glue can run three kinds of job. A **Spark ETL job** starts a managed Spark cluster and gives your script a `GlueContext`, `DynamicFrame`s and job bookmarks. A **Python shell job** runs a single plain-Python script on a small slice of capacity — you can request 1 DPU or 0.0625 DPU — with no Spark at all, which makes it the right home for lightweight glue work: firing a `COPY`/`MERGE` statement at Redshift, calling a vendor API, checking a file landed, emitting a metric. A **Ray job** runs a Ray cluster for Python-native distributed work. The rule of thumb is that if the data itself never has to be distributed, a Python shell job starts faster and costs less capacity than spinning up Spark to run ten lines of boto3. Note that job bookmarks and the DynamicFrame API belong to Spark jobs only.
code
python · 12 linesimport sys
import boto3
from awsglue.utils import getResolvedOptions
args = getResolvedOptions(sys.argv, ["cluster_id", "sql"])
redshift = boto3.client("redshift-data")
resp = redshift.execute_statement(
ClusterIdentifier=args["cluster_id"],
Database="analytics",
Sql=args["sql"],
)
print(resp["Id"])go deeper
Be able to say that Glue can run a plain single-process Python script as well as a Spark job, and give one example of work that belongs in each.
Explain the capacity difference — 1 or 0.0625 DPU for a Python shell job versus workers of DPUs for Spark — and name what the Python shell option loses: DynamicFrames, Glue transforms and job bookmarks.
Show that you audit job types against actual data volume, and can point at Glue jobs in an estate whose Spark cluster exists only to run a statement that a sixteenth of a DPU could issue.
Own the guidance that keeps teams from defaulting to one job type: where the boundary between control work and data work sits, and what it costs the platform bill when everything is provisioned as Spark.
## The three job types AWS Glue's compute side is not one thing. When you create a job you pick its type, and that choice decides what runtime your script gets: - **Spark ETL job** — Glue provisions a managed Apache Spark cluster and runs your PySpark or Scala script on it. The script gets a `SparkContext`, a Glue-specific `GlueContext` layered over it, and access to the Glue-only pieces: `DynamicFrame`, the `ApplyMapping`/`ResolveChoice`/`Relationalize` transforms, Data Catalog reads and writes, and job bookmarks. - **Python shell job** — Glue runs a single ordinary Python script in one process. There is no Spark, no cluster, no `GlueContext`. You get a Python interpreter with a set of preinstalled libraries (boto3, pandas, numpy and friends), and you can pull extra packages in with the `--additional-python-modules` job parameter. - **Ray job** (Glue for Ray) — Glue provisions a Ray cluster so you can distribute Python-native workloads that are awkward to express in Spark. ## Capacity shape A Spark job's capacity is expressed in workers of a given worker type, each worth one or more DPUs (a Data Processing Unit is Glue's bundled unit of vCPU and memory). A Python shell job is sized very differently: it takes either **1 DPU** or **0.0625 DPU** (one sixteenth). That sixteenth-DPU option is the entire point — it is the cheapest way in Glue to run a script that mostly waits on someone else's API. Since Glue bills by capacity consumed over the run's duration, running a ten-line boto3 script on a Spark cluster means paying for a cluster's worth of DPUs, plus the time it takes to bring Spark up, to do work a single interpreter finishes in seconds. ## What a Python shell job is actually used for In real pipelines the Python shell job is the connective tissue around the Spark jobs: - Issue a warehouse statement and wait for it — `COPY`, `MERGE`, `VACUUM`, a stored procedure call. - Call a third-party REST API for a small extract, or trigger a downstream system. - Do control-plane work: check for a file, publish a custom CloudWatch metric, write a manifest, tag an S3 prefix. - Run a small pandas transformation over a file that comfortably fits in one process's memory. ```python import sys import boto3 from awsglue.utils import getResolvedOptions args = getResolvedOptions(sys.argv, ["target_prefix"]) s3 = boto3.client("s3") # small, single-node control work — no Spark involved ``` Note that `awsglue.utils.getResolvedOptions` is available in a Python shell job, so job parameters are read the same way as in a Spark job — but nothing else from the `awsglue` Spark surface is. ## What you give up Choosing a Python shell job means giving up everything that hangs off Spark: - **No DynamicFrame and no Glue transforms.** `resolveChoice`, `relationalize`, `ApplyMapping` are DynamicFrame methods; DynamicFrames live in `GlueContext`, which lives in a Spark job. - **No job bookmarks.** Glue's built-in incremental state tracking is tied to the Glue Spark readers, so a Python shell job that wants to be incremental has to keep its own high-water mark somewhere (a DynamoDB item, an S3 marker, a control table). - **No horizontal scale.** One process, one machine's worth of memory. The moment the input outgrows that, you are rewriting the job. ## How to choose Ask what the data volume forces. If the job's whole reason for existing is to move or reshape enough data that one machine cannot hold it, it is a Spark job. If the data volume is incidental and the job is really an action — a call, a statement, a check — it is a Python shell job. If it is Python-native distributed computation that maps badly onto DataFrames, Ray is the reason that option exists. The failure mode interviewers are listening for is the team that has one hammer: every Glue job in the account is a Spark job, several of them exist only to run a single SQL statement, and nobody can explain why the bill has a floor. The opposite failure exists too — a Python shell job quietly pulling a growing table into pandas until it dies on memory.
- Can a Python shell job use job bookmarks to process only new files?No. Job bookmarks are state kept for the Glue Spark readers (`create_dynamic_frame.*` with a `transformation_ctx`), so they are a Spark ETL job feature. A Python shell job that needs to be incremental has to store its own high-water mark — a value in DynamoDB, a marker object in S3, or a control table row — and filter on it itself.
- How do you get a library that is not preinstalled into a Glue Python shell job?Pass the `--additional-python-modules` job parameter with a comma-separated list of packages, which Glue installs with pip before your script runs. For a wheel or an internal package you can point that parameter at an S3 path, or supply your own `.egg`/`.whl` through the job's Python library path setting.
saying these in an interview costs you the question
- Claiming every Glue job runs Spark under the hood
- Thinking DynamicFrames are available in a Python shell job
- Assuming Python shell jobs can use job bookmarks
- Sizing a Python shell job in workers rather than DPUs
- Running one SQL statement on a full Glue Spark cluster