skip to content

Glue Jobs and DynamicFrames

The compute half of Glue: managed Spark jobs with an AWS-specific DataFrame variant and built-in incremental bookmarks. Interviewers ask about DynamicFrame versus DataFrame and about bookmarks, because both are where Glue diverges from plain Spark.

on this pageshow

questions

6

In AWS Glue, what does a DynamicFrame give you that a Spark DataFrame does not?

level: middleimportance: must knowfreq 78%

answer

  1. one schema, or one per record
  2. what happens when a field is two types
  3. the Glue-only reader is the one with state
  4. ChoiceType, error records, bookmarks
  5. toDF crosses into Catalyst territory

basics

~20 s

A DynamicFrame carries a schema per record instead of one schema for the whole dataset, so conflicting types survive as a ChoiceType rather than failing or being coerced. It also brings Glue-only transforms, Data Catalog reads, error capture and job bookmarks.

solid answer

~50 s

A Spark `DataFrame` has one schema for the whole dataset, which Spark either infers by scanning or you supply. A Glue `DynamicFrame` is self-describing per record: each record carries its own schema, and when the same field is a string in some records and a long in others, Glue keeps both as a **ChoiceType** rather than erroring or silently coercing. That is the point — messy semi-structured input loads without a schema pass, and you resolve the ambiguity deliberately with `resolveChoice`. DynamicFrames also carry the Glue-specific surface: `create_dynamic_frame.from_catalog`/`from_options`, the `ApplyMapping`, `Relationalize` and `unnest` transforms, per-record error capture, and **job bookmarks** via `transformation_ctx`. What they do not have is the full Spark SQL and Catalyst surface, so the normal pattern is to read as a DynamicFrame, call `toDF()` for joins, windows and aggregations, then `DynamicFrame.fromDF()` to write.

code

python · 20 lines
python
from awsglue.dynamicframe import DynamicFrame

dyf = glueContext.create_dynamic_frame.from_catalog(
    database="raw",
    table_name="events",
    transformation_ctx="events_src",
)
dyf = dyf.resolveChoice(specs=[("amount", "cast:double")])

df = dyf.toDF()
df = df.filter("amount > 0").groupBy("country").sum("amount")

out = DynamicFrame.fromDF(df, glueContext, "agg")
glueContext.write_dynamic_frame.from_options(
    frame=out,
    connection_type="s3",
    connection_options={"path": "s3://curated/events/"},
    format="parquet",
    transformation_ctx="events_sink",
)

go deeper

for a junior

Recall that Glue has its own frame type built for messy input, and that toDF() and DynamicFrame.fromDF() move between it and a Spark DataFrame.

for a middle

Explain the per-record schema and ChoiceType, name what the Glue reader adds (catalog integration, bookmarks, error records), and describe the read-as-DynamicFrame, transform-as-DataFrame, write-back pattern.

for a senior

Show judgment about where the boundary sits in a real script: what you lose by reading with Spark directly, why choices are resolved before toDF(), and how you route error records rather than failing the run.

for a principal

Own the standard for a team: when DynamicFrames are mandatory (untyped ingest, incremental sources) versus when plain DataFrames are the default, and the review rule that stops each pipeline from inventing its own answer.

## One schema versus a schema per record A Spark `DataFrame` is a distributed table with **one** schema. That schema comes from a file footer (Parquet, ORC), a catalog, an explicit `StructType` you supply, or an inference pass in which Spark reads part or all of the data to guess. If a JSON field is a number in most records and a string in a few, inference picks one and the odd records become nulls, or the read fails. A Glue `DynamicFrame` inverts that. It is a collection of **self-describing records**: each record carries its own field types. Glue never needs a separate inference pass to load the data, and a field with inconsistent types across records is represented as a **ChoiceType** — a field that is, honestly, more than one type in this dataset. Nothing is dropped and nothing is guessed. That design exists because Glue's job is to ingest whatever is sitting in S3, including semi-structured data nobody validated on the way in. ## What else rides on the DynamicFrame The type model is only half the answer. The Glue-only surface hangs off DynamicFrames: - **Catalog and connection reads.** `glueContext.create_dynamic_frame.from_catalog(database=..., table_name=..., transformation_ctx=...)` and `create_dynamic_frame.from_options(...)` return DynamicFrames, and `glueContext.write_dynamic_frame.from_options(...)` writes them. - **Job bookmarks.** Incremental processing state is tracked against the Glue reader and keyed by the `transformation_ctx` you pass. Read the same S3 prefix with `spark.read.json(...)` and there is no bookmark — nothing is tracked. - **Glue transforms.** `resolveChoice`, `applyMapping`, `relationalize`, `unnest`, `unbox`, `dropFields`, `selectFields`, `splitFields`, `spigot`, and Glue's own `Join`. - **Error records.** A DynamicFrame keeps records it could not parse rather than failing the job; `errorsCount()` tells you how many there were and `errorsAsDynamicFrame()` hands them to you so you can route them to a quarantine path. ## What you give up Catalyst — Spark's query optimizer — plans over DataFrames. Predicate pushdown, join strategy selection, whole-stage codegen and adaptive query execution are DataFrame machinery. The DynamicFrame API is a thinner layer with a much smaller vocabulary: no window functions, no rich SQL expression language, no `spark.sql("...")` against it directly. So the practical pattern in almost every mature Glue script is a sandwich: ```python dyf = glueContext.create_dynamic_frame.from_catalog( database="raw", table_name="events", transformation_ctx="events_src") dyf = dyf.resolveChoice(specs=[("amount", "cast:double")]) df = dyf.toDF() # Spark DataFrame: joins, windows, SQL df = df.filter("amount > 0").groupBy("country").sum("amount") out = DynamicFrame.fromDF(df, glueContext, "out") glueContext.write_dynamic_frame.from_options( frame=out, connection_type="s3", connection_options={"path": "s3://curated/events/"}, format="parquet", transformation_ctx="events_sink") ``` Read as a DynamicFrame to get the catalog integration, the bookmark and the type tolerance; convert to a DataFrame for the actual transformation work; convert back to write. `toDF()` collapses the per-record schemas into one — which is exactly why you resolve choices **before** converting, not after. ## Cost and correctness consequences `toDF()` is not free in the abstract sense: the flexible representation has to be reconciled into a single schema, and DynamicFrame processing generally carries more per-record overhead than columnar DataFrame processing. On large, well-typed Parquet input where the schema is already stable, a team that reads straight into a DataFrame will often see a faster job — at the price of losing bookmarks and the Glue transforms. The correctness trap is the mirror image: converting a DynamicFrame with an unresolved ChoiceType to a DataFrame, or writing it to Parquet, forces a single type on a field that genuinely has two. Deciding that quietly is how a column of amounts loses its string-formatted rows. ## What interviewers are checking They want to hear that you know *why* AWS added a second frame type rather than reusing Spark's, that you can name the concrete things it buys (ChoiceType, bookmarks, catalog reads, error records), and that you know when to leave it — because the DynamicFrame API is not where you write a windowed aggregation.

  • If DataFrames are faster, why not read every Glue source with spark.read directly?
    Because reading outside `create_dynamic_frame.*` gives up the pieces the job may depend on: job bookmarks (they are tracked against the Glue reader and its `transformation_ctx`), Data Catalog and Glue connection integration, per-record error capture, and tolerance for inconsistent types. On stable, well-typed Parquet with no incremental requirement, reading straight into a DataFrame is a legitimate choice.
  • What does errorsAsDynamicFrame() give you, and what would you do with it?
    It returns the records the DynamicFrame could not parse as a DynamicFrame of their own, alongside `errorsCount()`. In production you write it to a quarantine prefix and alarm on the count, so malformed input becomes a visible, inspectable dataset instead of either a failed job or a silent gap in the output.
  • What happens to a ChoiceType column when you call toDF()?
    The flexible per-record schema has to collapse into one Spark schema, so the ambiguity is decided for you rather than by you. Resolve it first with `resolveChoice` — cast to a target type, split into per-type columns, or match the Data Catalog — so the choice is explicit and reviewable in the script.

saying these in an interview costs you the question

  • Calling a DynamicFrame just a wrapper with no real difference
  • Thinking Catalyst optimizes DynamicFrame transforms the same way
  • Converting to a DataFrame before resolving a ChoiceType
  • Believing job bookmarks work with spark.read on the same path
  • Claiming DynamicFrames require an upfront schema inference pass

context

open as a page

In an AWS Glue job, what does a job bookmark track between runs?

level: middleimportance: must knowfreq 72%

basics

~20 s

A Glue job bookmark is server-side state, kept per job, recording which input a previous run already processed — for S3 which objects were consumed, for JDBC how far the bookmark key columns advanced — so the next run reads only what is new.

open as a page

In an AWS Glue job, what does resolveChoice do to a column with a ChoiceType?

level: middleimportance: should knowfreq 48%

basics

~20 s

resolveChoice turns a DynamicFrame field that holds more than one type into a single well-defined shape. You choose the rule: cast to one type, project to one type, split into a column per type (make_cols), wrap the types in a struct (make_struct), or match the Data Catalog.

open as a page

An AWS Glue job with bookmarks enabled reprocesses every S3 object each run — what would you check?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Check that the script calls job.init and job.commit, that the source uses create_dynamic_frame with a stable transformation_ctx rather than spark.read, that the run really passed job-bookmark-enable, and that the job was not recreated or renamed since the last successful run.

open as a page

In an AWS Glue Spark job, what does moving from G.1X to G.2X workers change?

level: seniorimportance: should knowfreq 45%

basics

~20 s

G.2X gives each worker twice the DPU of G.1X — double the vCPUs, memory and disk per worker — so it raises per-executor capacity rather than total parallelism. Halving the worker count when you double the worker type keeps total DPUs, and cost, roughly flat.

open as a page

In AWS Glue, when would you run a Python shell job instead of a Spark ETL job?

level: juniorimportance: nice to knowfreq 40%

basics

~20 s

A Glue Python shell job runs one ordinary Python process with no Spark cluster behind it, so it fits small single-node work: calling an API, issuing SQL to a warehouse, moving a few megabytes. Spark ETL jobs are for distributed volumes.

open as a page