skip to content

How do you make a Great Expectations checkpoint actually fail the pipeline task that runs it?

level: seniorimportance: must knowfreq 55%

answer

  1. the tool reports, something else must enforce
  2. one boolean, read it or ignore it
  3. errors are captured, not thrown
  4. not every expectation deserves to stop a DAG
  5. stop it before consumers see it

basics

~20 s

Read the returned result's success flag and raise from the task - Great Expectations reports, it does not enforce. Its actions store results and rebuild Data Docs whether validation passed or failed, so nothing stops on its own.

solid answer

~50 s

Running a checkpoint returns a result object with a top-level `success`; if your task never inspects it, the run is pure reporting and the pipeline continues with bad data. The enforcement is yours: `if not result.success: raise ...` in the task body, or an orchestrator operator that fails the task on a false result. Two refinements matter in production. First, check for **recorded exceptions** as well as failures - an expectation whose column no longer exists is captured as an error in the result rather than crashing the run, and a naive `success` check may not tell you the check never really ran. Second, decide deliberately which suites **block** and which only **warn**, because a gate that fails the whole DAG on a cosmetic expectation gets disabled within a month. Blocking gates belong before publication so the failure mode is a stale table rather than a wrong one.

code

python · 4 lines
python
result = checkpoint.run()

if not result.success:
    raise RuntimeError("orders_gate failed - see Data Docs for details")

go deeper

for a junior

Recall the core fact: running a checkpoint returns a result with a success flag, and your code has to check it and raise - nothing stops on its own.

for a middle

Explain that actions run regardless of outcome, so the returned success boolean is the only control signal, and that an operator or an explicit raise in the task turns it into a task failure.

for a senior

Show production judgment: distinguish failed from could-not-run, tier blocking versus warning suites, place the gate before publication, and put the failing expectation names and counts into the alert.

for a principal

Own the enforcement policy across pipelines - which classes of violation may ever stop a business-critical load, who can bypass a gate and for how long, and how bypasses are made visible instead of becoming permanent.

## The default is report-only, and that surprises people A Great Expectations Checkpoint validates a batch against one or more suites and then runs its configured actions - persisting the validation result, rebuilding Data Docs, perhaps posting a notification. All of that happens whether validation passed or failed. What it does **not** do is throw. The invocation returns a result object carrying a top-level `success` that is false if any validation inside it failed, and that boolean is inert until something reads it. So the shortest correct answer is: enforcement is a line of your code. ```python result = checkpoint.run() if not result.success: raise RuntimeError("orders_gate failed - see Data Docs") ``` Orchestrators offer the same thing packaged: community operators for the common schedulers invoke a named checkpoint and fail the task when the result is unsuccessful. Either way the mechanism is identical - somebody translates a boolean into a non-zero exit. ## Failure is not the only bad outcome There are three distinct states, and treating them as two is a classic production hole: 1. **Validation ran and passed** - proceed. 2. **Validation ran and failed** - an expectation was evaluated and rows violated it. 3. **Validation could not run properly** - an expectation raised, most often because a column it references no longer exists in the source. Great Expectations captures such an error into the result with exception information rather than letting it crash the process, which is good for producing a complete report and bad if your gate logic only looks at the aggregate boolean and assumes a returned result means everything was checked. A robust gate therefore asserts that the run was *complete* as well as successful: no captured exceptions, and the expected number of validations actually executed. A suite silently reduced to two evaluable expectations after a schema change is not a gate. A fourth state deserves its own handling: the checkpoint task itself throwing for infrastructure reasons - the results store unreachable, the warehouse connection refused. If your orchestrator retries that task and then continues on the last attempt's exit code, make sure a genuine data failure is not indistinguishable from an infrastructure hiccup, and that the gate **fails closed** rather than being skipped. ## Blocking versus warning, and where to put the gate An all-or-nothing gate does not survive contact with a real team. Ten expectations, one of which flags a cosmetic formatting drift, will one night stop a business-critical load, someone will add a bypass flag at 3 a.m., and the flag will never be removed. The durable pattern is to tier the checks: - **Blocking suite** - a small set of invariants whose violation makes downstream results *wrong*: key uniqueness, key not-null, row-count floor, referential expectations for the joins that matter. This suite raises. - **Warning suite** - the broader set that describes health rather than correctness: distributions, optional-field completeness, format tolerances. This one records, renders and notifies without stopping anything. Placement matters as much as severity. Validate *before* publication - on the landing or staging table rather than after the mart is swapped in - so that a failing gate leaves consumers looking at yesterday's correct data instead of today's wrong data. A gate placed after the publish can only tell you how long you have been serving something broken. ## What a good failure tells the on-call Raising an exception with no context turns a data problem into a debugging expedition. The result already contains what is needed: which expectation failed, `unexpected_count` and `unexpected_percent`, and a sample of the actual offending values. Put the failing expectation names and those counts into the exception message and the alert, and link to the Data Docs page for the run. "orders_gate: expect_column_values_to_be_unique(order_id) failed, 412 duplicates of 2.1M" is actionable at 3 a.m.; "data quality check failed" is not. ## Rerun behaviour Finally, think about what a fixed rerun does. Because the gate is upstream of publication and the validation itself is side-effect-free, a rerun after fixing the source is usually safe - but only if the pipeline steps around it are idempotent. A gate that fails after a partial load has already committed rows leaves you re-running a load that will double-write unless the load itself is written to be repeatable. The quality gate does not make an unsafe pipeline safe; it only tells you sooner. ## In an interview Lead with the one-liner - GX reports, your code enforces - then show operational depth: distinguish failure from a captured exception, tier blocking versus warning suites, validate before publish, and put the failing expectation and counts into the alert.

  • An expectation references a column the source dropped. Does the checkpoint crash?
    No - the error is captured into the result with exception information rather than propagating, so the run finishes and reports. That is why a gate should assert completeness as well as success: check that no exceptions were recorded and that the expected validations actually executed, otherwise a schema change can quietly shrink your suite to whatever still evaluates.
  • Why not simply make every expectation blocking?
    Because one cosmetic expectation will eventually stop a business-critical load at 3 a.m., someone adds a bypass, and the bypass becomes permanent. Split into a small blocking suite of invariants that make results wrong - key uniqueness, not-null keys, a row-count floor - and a broader warning suite that records and notifies without halting the pipeline.
  • Should the gate run before or after the table is published?
    Before. If validation runs on the landing or staging output and blocks the swap, a failure leaves consumers on yesterday's correct data. A gate placed after publication can only measure how long you have been serving wrong numbers, and remediation then involves rolling back something people have already queried.
  • What belongs in the alert when a gate fails?
    The failing expectation names with their unexpected_count and unexpected_percent, a couple of the actual offending values, and a link to the Data Docs page for that run. The result object already carries all of it. A bare "data quality check failed" forces the on-call to reconstruct from scratch what the result had already computed.

saying these in an interview costs you the question

  • Assuming a failing checkpoint halts the pipeline automatically
  • Checking success but ignoring captured expectation exceptions
  • Making every expectation blocking, which invites permanent bypasses
  • Validating after publish, when consumers already read the bad data
  • Raising a bare error with no failing expectation or counts

context