In an AWS Glue job, what does a job bookmark track between runs?
answer
- state that survives between runs
- kept per job, not in your data
- the reader needs a name for its own position
- the string that names a source's position
- nothing is remembered until commit
basics
~20 sA Glue job bookmark is server-side state, kept per job, recording which input a previous run already processed — for S3 which objects were consumed, for JDBC how far the bookmark key columns advanced — so the next run reads only what is new.
solid answer
~50 sJob bookmarks are AWS Glue's built-in incremental-processing state. Glue stores them per job and keys each entry by the `transformation_ctx` string you pass to a reader, so one script can track several sources independently. For **S3** sources the bookmark records which objects have already been consumed, based on their path and last-modified time; for **JDBC** sources it records the high-water value of the bookmark key columns, which must increase monotonically. Three things must all be true for it to work: the job's `--job-bookmark-option` is `job-bookmark-enable`, the script creates a `Job` and calls `job.init(args['JOB_NAME'], args)` at the start and `job.commit()` at the end, and every source and sink you want tracked passes a stable `transformation_ctx`. The bookmark advances only on `job.commit()`. You can pause it (`job-bookmark-pause`) to re-run without moving state, or clear it with `aws glue reset-job-bookmark`.
code
python · 26 linesimport sys
from awsglue.utils import getResolvedOptions
from awsglue.context import GlueContext
from awsglue.job import Job
from pyspark.context import SparkContext
args = getResolvedOptions(sys.argv, ["JOB_NAME"])
glueContext = GlueContext(SparkContext.getOrCreate())
job = Job(glueContext)
job.init(args["JOB_NAME"], args)
orders = glueContext.create_dynamic_frame.from_catalog(
database="raw",
table_name="orders",
transformation_ctx="orders_src",
)
glueContext.write_dynamic_frame.from_options(
frame=orders,
connection_type="s3",
connection_options={"path": "s3://curated/orders/"},
format="parquet",
transformation_ctx="orders_sink",
)
job.commit() # without this the bookmark never advancesgo deeper
Recall that a Glue job bookmark is what makes a job pick up only new input, and that it is a setting on the job rather than something you build in your script.
Explain what is recorded for S3 versus JDBC sources, and name the three prerequisites: the enable option, job.init/job.commit, and a stable transformation_ctx per reader.
Demonstrate the operational picture: pause versus reset, the at-least-once consequence when a run dies after writing but before committing, and why the sink still has to be idempotent.
Own the policy question of when a team should rely on Glue-managed bookmark state at all versus an explicit, inspectable watermark the platform can audit, replay and reason about across tools.
## The problem bookmarks solve A Glue job reading `s3://raw/events/` every hour would, by default, read everything in that prefix every hour. Job bookmarks are AWS Glue's managed answer: Glue keeps state about what previous runs already consumed, and the Glue readers use it to skip that data on the next run. The state lives on the AWS side, attached to the job — you do not create a table for it and you cannot query it as data. You interact with it through the job's bookmark option, the console's reset action, and the `reset-job-bookmark` API. ## What is actually recorded This depends on the source: - **Amazon S3** — the bookmark records which objects have been processed, identified by path and last-modified timestamp. On the next run, the Glue reader lists the prefix and skips objects it has already seen. Note the consequence: it is object-level, and it is timestamp-driven. A file that is overwritten in place, or that lands with a last-modified time behind the bookmark's high-water mark, is where the model bites. - **JDBC sources** — the bookmark records the maximum value seen for the **bookmark key** column or columns, supplied through the `jobBookmarkKeys` connection option with `jobBookmarkKeysSortOrder`. Those columns must be monotonically increasing (an identity key, an append-only timestamp); if the key can go backwards or be updated, rows are silently skipped. There is no bookmark for a Python shell job — bookmarks are a feature of the Glue Spark readers. ## The three things that must all hold Bookmarks fail closed and quietly, and every failure comes back to one of three requirements. **1. The option is on.** `--job-bookmark-option` takes `job-bookmark-enable`, `job-bookmark-disable` or `job-bookmark-pause`, set on the job and overridable per run. `pause` is the useful middle setting: the run reads from the current bookmark position but does not advance it, which is how you re-run a failed load without moving state. **2. The job object is wired up.** The script must construct Glue's `Job` and bracket its work: ```python from awsglue.job import Job job = Job(glueContext) job.init(args["JOB_NAME"], args) # ... read, transform, write ... job.commit() ``` `job.commit()` is what makes the new bookmark state durable. A script that ends without it processes data and then forgets it did — the next run reads the same input again. **3. Every tracked reader has a `transformation_ctx`.** The bookmark is keyed by this string, which is how a script with three sources keeps three independent positions: ```python orders = glueContext.create_dynamic_frame.from_catalog( database="raw", table_name="orders", transformation_ctx="orders_src") ``` That string is part of the identity of the state. Rename `orders_src` to `orders_source` in a refactor and Glue sees a source it has never met — the next run reprocesses the entire history. Passing `transformation_ctx` on the sink as well matters for writers that need to track their own output. ## Operating them - **Reset**: `aws glue reset-job-bookmark --job-name my-job` clears the state so the next run starts from scratch; the console exposes the same action. Reset before a deliberate full reload. - **Pause**: use `job-bookmark-pause` to re-run a window without disturbing the position. - **Disable**: the job reads everything every time, which is the correct setting for a job whose sink is an idempotent full overwrite. A bookmark is not a substitute for idempotent writes. If a run fails *after* writing output but before `job.commit()`, the bookmark did not advance, so the next run reprocesses that input and writes it again. That is the right failure direction — at-least-once rather than data loss — but it means the sink must tolerate the repeat, by overwriting a partition or merging on a key rather than blindly appending. ## What good answers include Candidates who have actually run Glue in production mention `job.commit()` and `transformation_ctx` unprompted, know that the option defaults to off, and can state the at-least-once consequence: bookmarks tell you what has been *read*, and say nothing about what has been *written*.
- What is job-bookmark-pause for, and when would you use it?With `--job-bookmark-option job-bookmark-pause` the run reads from the current bookmark position but does not advance it on commit. It is the setting for re-running a load you are debugging, or for a parallel test run of the same job, without losing the position or having to reset and reprocess everything afterwards.
- If a Glue run writes its output and then fails before job.commit(), what happens on the next run?The bookmark never advanced, so the next run reads the same input again and writes it again. That is at-least-once by design — it prefers duplication to loss — which is why the sink must be idempotent: overwrite the target partition or merge on a business key rather than appending blindly.
- What do bookmark keys require of a JDBC source column?They must be monotonically increasing, supplied through the `jobBookmarkKeys` connection option with an explicit sort order. An identity column or an append-only insert timestamp works; a mutable `updated_at` that can be set backwards, or a key that is reused, causes rows to fall below the recorded high-water mark and be skipped without any error.
saying these in an interview costs you the question
- Believing bookmarks are enabled by default on a new job
- Forgetting job.commit() and expecting state to persist
- Thinking the bookmark records what was written, not what was read
- Renaming transformation_ctx without expecting a full reprocess
- Treating bookmarks as exactly-once delivery to the sink