skip to content

An AWS Glue job with bookmarks enabled reprocesses every S3 object each run — what would you check?

level: seniorimportance: should knowfreq 55%

answer

  1. it fails green, which is the clue
  2. three things must all be true
  3. the state has a name, and names change
  4. reading past the Glue reader loses it
  5. nothing is remembered without the final call

basics

~20 s

Check that the script calls job.init and job.commit, that the source uses create_dynamic_frame with a stable transformation_ctx rather than spark.read, that the run really passed job-bookmark-enable, and that the job was not recreated or renamed since the last successful run.

solid answer

~50 s

Work down the chain that has to hold for a Glue bookmark to advance. **Is the option on for the run?** `--job-bookmark-option` defaults to disabled and can be overridden per run, so check the actual run's arguments, not just the job definition. **Does the script commit?** `job.init(args['JOB_NAME'], args)` at the start and `job.commit()` at the end are what persist the position — a script that ends without commit reprocesses forever and never errors. **Is the read going through Glue?** Bookmarks are tracked by the Glue readers; `spark.read.json("s3://...")` bypasses them entirely. **Is `transformation_ctx` present and unchanged?** State is keyed by that string, so a missing one is untracked and a renamed one looks like a brand-new source. Finally, a job that was deleted and recreated, or renamed, starts with empty state. Confirm by checking whether the job ran green but the input row count stayed at the full history.

code

python · 23 lines
python
# broken: bookmark can never advance
args = getResolvedOptions(sys.argv, ["JOB_NAME"])
glueContext = GlueContext(SparkContext.getOrCreate())
df = glueContext.spark_session.read.json("s3://raw/events/")
df.write.mode("append").parquet("s3://curated/events/")

# fixed: Glue reader + transformation_ctx + init/commit
job = Job(glueContext)
job.init(args["JOB_NAME"], args)
dyf = glueContext.create_dynamic_frame.from_options(
    connection_type="s3",
    connection_options={"paths": ["s3://raw/events/"]},
    format="json",
    transformation_ctx="events_src",
)
glueContext.write_dynamic_frame.from_options(
    frame=dyf,
    connection_type="s3",
    connection_options={"path": "s3://curated/events/"},
    format="parquet",
    transformation_ctx="events_sink",
)
job.commit()

go deeper

for a junior

Know that a Glue bookmark has prerequisites in the script itself, and that job.commit() missing from the end is the classic reason a job keeps rereading everything.

for a middle

Walk the chain: run arguments, init and commit, Glue reader versus plain Spark read, and a present and stable transformation_ctx.

for a senior

Diagnose from evidence rather than guesses — read volume per run, what changed at the last deploy — and know the mirror-image failure where the bookmark silently skips late or rewritten objects.

for a principal

Set the standard that makes this class of bug rare: transformation_ctx treated as a contract, deploys that update rather than replace jobs, idempotent sinks so a reprocessing run costs money rather than correctness.

## Why this is a diagnosis question Glue bookmarks fail silently in the safe direction. Nothing errors; the run is green; the only symptom is that the job reads the whole prefix every time, so the runtime and the DPU-hours climb with the size of history and the sink gets duplicates unless it happens to be an overwrite. Interviewers use this scenario because the fix requires knowing what the bookmark mechanism actually depends on. ## The checklist, in the order it pays off ### 1. Was the option actually set on the run? `--job-bookmark-option` has three values — `job-bookmark-enable`, `job-bookmark-disable`, `job-bookmark-pause` — and it can be set on the job's default arguments *and* overridden in `start-job-run`. A job defined with bookmarks on but started by an orchestrator that passes its own argument map can be running with them off. Look at the arguments on the individual run. ### 2. Does the script call init and commit? ```python job = Job(glueContext) job.init(args["JOB_NAME"], args) # ... job.commit() ``` `job.commit()` is the moment new bookmark state becomes durable. This is the single most common cause: a script that was refactored, or generated and then hand-edited, and lost its commit. Symptomatically it is indistinguishable from bookmarks being off — every run starts from the beginning — so check both. Also check that `job.init` receives the resolved arguments, not a hand-built dict: it needs `JOB_NAME` and the bookmark-related arguments to identify and load the right state. ### 3. Is the read going through a Glue reader? Bookmark state is tracked by `glueContext.create_dynamic_frame.from_catalog(...)` and `from_options(...)`. A script that reads with `spark.read.parquet("s3://raw/events/")` — often introduced by someone optimising for DataFrame performance — has no bookmark at all for that source. The rest of the script can look perfectly bookmark-shaped, commit included, and still reprocess everything. ### 4. Is `transformation_ctx` present, and has it changed? State is keyed by job plus `transformation_ctx`. Two failure shapes: - **Missing** — a reader with no `transformation_ctx` is not tracked. - **Renamed** — a refactor from `"events_src"` to `"events_source"` presents Glue with a source it has never seen, so the first run after the deploy reads all history. That one is worth recognising by its shape: bookmarks worked, a deploy happened, one giant run followed, and then things settled. ### 5. Did the job identity change? Bookmark state belongs to the job. Delete and recreate the job with the same name from CloudFormation or Terraform, or rename it, and the state does not follow. A CI pipeline that replaces rather than updates the job will do this on every deploy. ### 6. Did someone reset it? `aws glue reset-job-bookmark --job-name ...` and the console's reset action clear the position. If it happens repeatedly, look for an operational runbook or a deploy step that resets as a matter of habit. ## The mirror-image bug While you are in here, know the opposite failure, because interviewers often follow up with it: **the bookmark skipping data that should have been processed**. For S3 the bookmark is object-level and driven by path and last-modified time, so an object rewritten in place under a path already consumed, or one that lands with a last-modified timestamp behind the recorded position, can be missed. For JDBC sources the bookmark keys must be monotonically increasing — a mutable `updated_at` that can move backwards, or a reused identity value, leaves rows permanently below the high-water mark. Nothing errors there either. ## Confirming and preventing Confirm by instrumenting rather than guessing: log the input record count per run, and watch whether it equals the full history or only the increment. Glue's job metrics and the run's own logs will show the read volume; a bookmark that is working shows a small, steady read against a growing prefix. Prevent it with three habits: treat `transformation_ctx` values as part of the job's contract and never rename them casually; make the deploy pipeline update the job rather than replace it; and make sinks idempotent regardless — partition overwrite or merge on a key — so that a reprocessing run is expensive rather than corrupting.

  • What is the opposite failure — a bookmark that skips data that should have been read?
    For S3 the bookmark is object-level and driven by path and last-modified time, so a file overwritten in place under an already-consumed path, or one landing with a timestamp behind the recorded position, is skipped. For JDBC, bookmark keys must increase monotonically; a mutable timestamp that moves backwards leaves rows permanently below the high-water mark. Neither raises an error.
  • A deploy replaced the Glue job via CloudFormation and the next run took six hours. What happened?
    Bookmark state belongs to the job, so a stack change that deletes and recreates it rather than updating in place starts with empty state and the next run reads the whole history. Fix the pipeline to update the job resource, and treat a full-history read after a deploy as a signal to check job identity rather than the script.
  • How would you prove from the outside whether the bookmark is advancing?
    Compare the input volume per run against the prefix's growth: log the record or byte count read by the source and watch it in CloudWatch. A working bookmark reads a small, roughly constant increment while the prefix grows; a broken one reads a total that tracks the whole history. Run duration alone is a weaker signal because it also moves with cluster size.

saying these in an interview costs you the question

  • Blaming S3 eventual consistency for the reprocessing
  • Assuming enabling the option is sufficient on its own
  • Not knowing spark.read bypasses bookmark tracking
  • Renaming transformation_ctx and calling it cosmetic
  • Recommending reset-job-bookmark as the fix for reprocessing

context