skip to content

Data Quality and Contracts

Ingestion that loads bad data faithfully is still a broken pipeline, so mature platforms validate what arrives and agree in writing on what producers owe consumers. Expect to be asked what you check, where you check it, and what happens to the records that fail.

on this pageshow

explore

questions

18

In Great Expectations, what is an Expectation Suite and what does validating one produce?

level: juniorimportance: must knowfreq 72%

answer

  1. rules as data, not as assert statements
  2. a named bundle scoped to one dataset
  3. names like expect_column_values_to_not_be_null
  4. output carries counts and offending samples
  5. rendered into a static HTML report

basics

~20 s

An Expectation Suite is a named collection of declarative assertions about one dataset - not-null, unique, between, matches-regex. Validating it against a batch of data returns a Validation Result with per-expectation pass/fail plus counts and samples of the offending values.

solid answer

~40 s

An **Expectation** is one declarative assertion about data: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`. An **Expectation Suite** is a named, serializable bundle of those assertions describing what one dataset must look like; because it serializes to a config (name plus kwargs) it can be reviewed in a pull request and versioned like code. Validation runs a suite against a *batch* of data and returns a Validation Result: an overall `success` boolean plus, per expectation, its own success flag and a `result` block carrying `element_count`, `unexpected_count`, `unexpected_percent` and a capped sample of the failing values. Those results also render into **Data Docs**, a static HTML report showing what was expected and what actually happened. Great Expectations 1.x reorganized the Python API, but this suite/validation vocabulary carried over from 0.x.

code

python · 13 lines
python
import great_expectations as gx

context = gx.get_context()
suite = context.suites.add(gx.ExpectationSuite(name="orders_raw"))
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeBetween(column="amount", min_value=0)
)

go deeper

for a junior

Be ready to name three or four real expectation types and say in one sentence what a suite is. Knowing that the output is a pass/fail result with counts of offending rows is enough at this stage.

for a middle

Explain the mechanics: how a suite serializes to config, what a batch is, and what fields the validation result carries - element_count, unexpected_count, and the sampled unexpected values controlled by result_format.

for a senior

Show judgment about which datasets deserve suites and how strict to make them. Interviewers listen for the point that a profiled suite unedited becomes alert noise, and that validation without someone acting on success is decoration.

for a principal

Own the question of where declarative quality rules live across the org - which layer validates, who reviews a change to a suite, and how you keep a growing library of suites from becoming an unowned graveyard of stale thresholds.

## What Great Expectations is Great Expectations (GX) is an open-source Python framework for expressing data-quality rules as **declarative assertions** and running them as a gate inside a data pipeline. Rather than scattering ad-hoc checks like `assert df.order_id.notnull().all()` through transformation code, you declare what the data must look like, store that declaration as a reviewable artifact, and run it against data at a pipeline step or on a schedule. ## An Expectation An Expectation is a single named assertion plus keyword arguments - for example `expect_column_values_to_be_between(column="amount", min_value=0)`. They come in families: - **Column map expectations** evaluate row by row: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`, `expect_column_values_to_be_in_set`. - **Column aggregate expectations** evaluate one statistic over the column: `expect_column_mean_to_be_between`, `expect_column_distinct_values_to_be_in_set`. - **Table expectations** describe the whole batch: `expect_table_row_count_to_be_between`, `expect_table_columns_to_match_ordered_list`. Every expectation is serializable - the expectation name plus its kwargs - which is precisely why a suite reads as a spec rather than as code. ## The Expectation Suite A suite is a named collection of expectations scoped to one dataset or asset (say, `orders_raw`). It lives in the **Data Context** - GX's entry point and configuration root, either backed by a project directory on disk or created ephemerally in memory - so suites can be stored beside the pipeline, diffed, and reviewed. You author a suite three ways: by hand, interactively against a sample batch while you inspect it, or by profiling a sample so GX proposes expectations from observed ranges and distributions. A profiled suite is a **draft, not a contract**: unedited profiler output encodes accidents of whatever sample it saw, and the usual failure is a suite full of over-tight ranges that fires every Monday. ## What validation produces Validation applies a suite to a **batch** - a defined slice of data, such as one day's partition of a table or one landed file. The output is a validation result for each expectation and an overall suite-level result: - top-level `success` - false if any expectation failed; - per expectation: its own `success`, the expectation's configuration, and a `result` dictionary with `element_count` (rows evaluated), `unexpected_count`, `unexpected_percent`, and a bounded `partial_unexpected_list` of actual offending values. The verbosity of that block is controlled by `result_format`, which ranges from `BOOLEAN_ONLY` through `BASIC` and `SUMMARY` to `COMPLETE`. `COMPLETE` returns every unexpected value, which is wonderful for debugging and dangerous on a wide failure: a column that is entirely wrong will serialize millions of values into the result store. ## Data Docs GX renders suites and validation results into **Data Docs**, a static HTML site: what we expect of this dataset, and what the last runs found. It is the artifact you point an analyst at when they ask why the load was blocked. It is documentation and evidence, not a monitoring system - it has no alerting of its own. ## What a suite does not do by itself Three things trip newcomers up. First, running a suite does not stop anything - somebody must read `success` and act on it. Second, GX validates; it never repairs data. Third, GX does not schedule itself - an orchestrator task, a job step, or a CI job invokes it. ## In an interview Be able to say the vocabulary chain out loud: Data Context holds the configuration; a data source and asset define where the data is; a batch is the slice being checked; an Expectation Suite is the set of rules; validation produces a Validation Result; a Checkpoint is the runnable unit that ties suite to batch and fires actions; Data Docs is the rendered report. Being able to place *where validation executes* in a pipeline is the point of learning the tool's nouns at all.

  • What is the difference between a column map expectation and a table expectation in Great Expectations?
    A column map expectation such as `expect_column_values_to_match_regex` evaluates every row of one column and reports how many rows were unexpected. A table expectation such as `expect_table_row_count_to_be_between` or `expect_table_columns_to_match_ordered_list` makes a single assertion about the batch as a whole, so it either passes or fails outright with no unexpected-row count.
  • Is a suite generated by profiling a sample safe to ship as-is?
    No. Profiling proposes expectations from whatever the sample happened to contain, so ranges and value sets encode accidents of that snapshot. Treat it as a first draft: keep the expectations that express a real invariant, delete the coincidences, and widen numeric bounds to what the business actually guarantees. Otherwise the suite becomes noisy and people start ignoring failures.
  • What does the result_format setting change?
    It controls how much detail a validation result carries: `BOOLEAN_ONLY` just pass/fail, `BASIC` and `SUMMARY` add counts and a capped sample of unexpected values, `COMPLETE` returns every unexpected value. `COMPLETE` is excellent for debugging one run and hazardous as a default, because a wholly broken column serializes an enormous list into the result store.

A suite is a checklist you hand to an inspector, not the inspection itself - the Validation Result is the filled-in checklist that comes back with the failures circled.

saying these in an interview costs you the question

  • Calling expectations imperative Python asserts rather than serializable config
  • Thinking a failing suite stops the pipeline automatically
  • Shipping profiler output unedited as the production suite
  • Believing Great Expectations cleans or corrects bad rows
  • Confusing Data Docs with an alerting or monitoring system

context

open as a page

Why route rejected rows to a quarantine table instead of dropping them or failing the load?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A quarantine table keeps each rejected row plus the reason it failed, so good data still lands while bad data stays visible and replayable. Dropping destroys the evidence silently; failing the whole load punishes every valid row for a few defects.

open as a page

How do you enforce a data contract in CI on the producing team's pull requests?

level: middleimportance: must knowfreq 58%

basics

~20 s

Keep the contract spec in the producer's repository and add a CI job that diffs the proposed spec against the released one, classifies each change as compatible or breaking, and fails the pull request on a breaking change unless a major version bump and consumer sign-off accompany it.

open as a page

Which ingestion validation failures should fail the whole load instead of quarantining rows?

level: middleimportance: must knowfreq 56%

basics

~20 s

Fail the load when the defect makes the whole batch untrustworthy: an unrecognisable schema, a row count or control total that disagrees with the manifest, duplicate keys that would corrupt a merge, or a reject rate far above baseline. Independent per-row defects belong in quarantine.

open as a page

How do you make a Great Expectations checkpoint actually fail the pipeline task that runs it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Read the returned result's success flag and raise from the task - Great Expectations reports, it does not enforce. Its actions store results and rebuild Data Docs whether validation passed or failed, so nothing stops on its own.

open as a page

Under a data contract, how does a producer safely drop a field consumers still read?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Never delete in place. Publish a contract version that marks the field deprecated with a removal date and a replacement, keep emitting it through the agreed window, watch lineage and query logs for remaining readers, then remove it in a major version once usage reaches zero.

open as a page

What does a data contract specify beyond the column names and types?

level: juniorimportance: should knowfreq 45%

basics

~20 s

A data contract pins down the schema plus everything a schema cannot say: field meanings and units, the key and grain, freshness and volume expectations, a named owning team, a version, and how breaking changes get announced.

open as a page

In Great Expectations, what does a Checkpoint add beyond an Expectation Suite?

level: middleimportance: should knowfreq 58%

basics

~20 s

A Checkpoint is the runnable unit: it binds an Expectation Suite to a specific batch of data, executes the validation, and fires configured actions such as storing the result and rebuilding Data Docs. The suite alone only declares rules.

open as a page

In Great Expectations, what does the mostly parameter change about an expectation's result?

level: middleimportance: should knowfreq 42%

basics

~20 s

mostly sets a tolerance: the expectation passes when at least that fraction of evaluated rows satisfy the condition, instead of requiring every row. It applies to row-wise column expectations, not to aggregate or table-level ones.

open as a page

What do freshness and volume SLOs add to a data contract that schema checks miss?

level: middleimportance: should knowfreq 50%

basics

~20 s

Schema checks prove the shape is right; freshness and volume SLOs prove the data actually arrived. A contract states a maximum staleness and an expected row-count range, so a perfectly-typed but empty or half-loaded dataset still raises an alert.

open as a page

What error context must a quarantine table carry alongside each rejected record?

level: middleimportance: should knowfreq 54%

basics

~20 s

Store the payload exactly as it arrived, the check that rejected it with a stable reason code, where it came from (file and line, or topic, partition and offset), the pipeline run id, arrival and rejection timestamps, and a status field that drives replay.

open as a page

When Great Expectations validates a billion-row warehouse table, where does the computation run?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It depends on the backend. Against a SQL data source, expectations are translated into queries the warehouse executes, and only aggregates come back. Against the pandas backend, the batch is pulled into the Python process's memory first.

open as a page

A streaming ingestion consumer crashes on the same record after every restart — how do you isolate it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Classify the failure first. If it is deterministic, bound the attempts, divert that record with its key, position and error into a dead-letter destination, advance past it, and alert. Never divert on transient failures — an outage would dead-letter the whole stream.

open as a page

How do you replay quarantined records after a fix without double-loading rows that already landed?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Replay through the normal pipeline with an idempotent keyed merge, select only the rejects whose reason matches the fix, guard the write with the record's own version or event time so stale payloads cannot overwrite newer state, and mark each record replayed under a conditional status update.

open as a page

How would you roll out data contracts across teams whose producers see no benefit?

level: principalimportance: should knowfreq 38%

basics

~20 s

Start where a repeat incident already hurts, cover one dataset, bootstrap the contract from what the producer already emits so day one is green, make the check cheap and local to their CI, then expand only after the mechanism has visibly caught a real break.

open as a page

Why can a Great Expectations regex expectation pass on a column that is entirely null?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

Most row-wise value expectations evaluate only non-null values, so a fully null column offers nothing to violate the pattern and the check passes vacuously. Assert completeness separately with expect_column_values_to_not_be_null.

open as a page

What does generating DDL, tests and docs from a data contract spec buy you?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Generation makes the contract load-bearing instead of descriptive. If the target DDL, the validation checks and the catalog entry are all produced from one spec, the documented contract cannot drift from the enforced one, because there is only one source.

open as a page

At what quarantine reject rate should an ingestion pipeline stop rather than keep loading?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

There is no universal number. Set the threshold per source and per reason code against that feed's own baseline, as a rate rather than a count, and stop when a partial load would mislead consumers more than a late load would delay them.

open as a page