skip to content

Great Expectations

The best-known Python framework for expressing data-quality rules as declarative expectations and running them as a gate inside a pipeline. Knowing its vocabulary — suites, checkpoints, validation results — lets you talk concretely about where validation actually executes.

on this pageshow

explore

questions

6

In Great Expectations, what is an Expectation Suite and what does validating one produce?

level: juniorimportance: must knowfreq 72%

answer

  1. rules as data, not as assert statements
  2. a named bundle scoped to one dataset
  3. names like expect_column_values_to_not_be_null
  4. output carries counts and offending samples
  5. rendered into a static HTML report

basics

~20 s

An Expectation Suite is a named collection of declarative assertions about one dataset - not-null, unique, between, matches-regex. Validating it against a batch of data returns a Validation Result with per-expectation pass/fail plus counts and samples of the offending values.

solid answer

~40 s

An **Expectation** is one declarative assertion about data: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`. An **Expectation Suite** is a named, serializable bundle of those assertions describing what one dataset must look like; because it serializes to a config (name plus kwargs) it can be reviewed in a pull request and versioned like code. Validation runs a suite against a *batch* of data and returns a Validation Result: an overall `success` boolean plus, per expectation, its own success flag and a `result` block carrying `element_count`, `unexpected_count`, `unexpected_percent` and a capped sample of the failing values. Those results also render into **Data Docs**, a static HTML report showing what was expected and what actually happened. Great Expectations 1.x reorganized the Python API, but this suite/validation vocabulary carried over from 0.x.

code

python · 13 lines
python
import great_expectations as gx

context = gx.get_context()
suite = context.suites.add(gx.ExpectationSuite(name="orders_raw"))
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
suite.add_expectation(
    gx.expectations.ExpectColumnValuesToBeBetween(column="amount", min_value=0)
)

go deeper

for a junior

Be ready to name three or four real expectation types and say in one sentence what a suite is. Knowing that the output is a pass/fail result with counts of offending rows is enough at this stage.

for a middle

Explain the mechanics: how a suite serializes to config, what a batch is, and what fields the validation result carries - element_count, unexpected_count, and the sampled unexpected values controlled by result_format.

for a senior

Show judgment about which datasets deserve suites and how strict to make them. Interviewers listen for the point that a profiled suite unedited becomes alert noise, and that validation without someone acting on success is decoration.

for a principal

Own the question of where declarative quality rules live across the org - which layer validates, who reviews a change to a suite, and how you keep a growing library of suites from becoming an unowned graveyard of stale thresholds.

## What Great Expectations is Great Expectations (GX) is an open-source Python framework for expressing data-quality rules as **declarative assertions** and running them as a gate inside a data pipeline. Rather than scattering ad-hoc checks like `assert df.order_id.notnull().all()` through transformation code, you declare what the data must look like, store that declaration as a reviewable artifact, and run it against data at a pipeline step or on a schedule. ## An Expectation An Expectation is a single named assertion plus keyword arguments - for example `expect_column_values_to_be_between(column="amount", min_value=0)`. They come in families: - **Column map expectations** evaluate row by row: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`, `expect_column_values_to_be_in_set`. - **Column aggregate expectations** evaluate one statistic over the column: `expect_column_mean_to_be_between`, `expect_column_distinct_values_to_be_in_set`. - **Table expectations** describe the whole batch: `expect_table_row_count_to_be_between`, `expect_table_columns_to_match_ordered_list`. Every expectation is serializable - the expectation name plus its kwargs - which is precisely why a suite reads as a spec rather than as code. ## The Expectation Suite A suite is a named collection of expectations scoped to one dataset or asset (say, `orders_raw`). It lives in the **Data Context** - GX's entry point and configuration root, either backed by a project directory on disk or created ephemerally in memory - so suites can be stored beside the pipeline, diffed, and reviewed. You author a suite three ways: by hand, interactively against a sample batch while you inspect it, or by profiling a sample so GX proposes expectations from observed ranges and distributions. A profiled suite is a **draft, not a contract**: unedited profiler output encodes accidents of whatever sample it saw, and the usual failure is a suite full of over-tight ranges that fires every Monday. ## What validation produces Validation applies a suite to a **batch** - a defined slice of data, such as one day's partition of a table or one landed file. The output is a validation result for each expectation and an overall suite-level result: - top-level `success` - false if any expectation failed; - per expectation: its own `success`, the expectation's configuration, and a `result` dictionary with `element_count` (rows evaluated), `unexpected_count`, `unexpected_percent`, and a bounded `partial_unexpected_list` of actual offending values. The verbosity of that block is controlled by `result_format`, which ranges from `BOOLEAN_ONLY` through `BASIC` and `SUMMARY` to `COMPLETE`. `COMPLETE` returns every unexpected value, which is wonderful for debugging and dangerous on a wide failure: a column that is entirely wrong will serialize millions of values into the result store. ## Data Docs GX renders suites and validation results into **Data Docs**, a static HTML site: what we expect of this dataset, and what the last runs found. It is the artifact you point an analyst at when they ask why the load was blocked. It is documentation and evidence, not a monitoring system - it has no alerting of its own. ## What a suite does not do by itself Three things trip newcomers up. First, running a suite does not stop anything - somebody must read `success` and act on it. Second, GX validates; it never repairs data. Third, GX does not schedule itself - an orchestrator task, a job step, or a CI job invokes it. ## In an interview Be able to say the vocabulary chain out loud: Data Context holds the configuration; a data source and asset define where the data is; a batch is the slice being checked; an Expectation Suite is the set of rules; validation produces a Validation Result; a Checkpoint is the runnable unit that ties suite to batch and fires actions; Data Docs is the rendered report. Being able to place *where validation executes* in a pipeline is the point of learning the tool's nouns at all.

  • What is the difference between a column map expectation and a table expectation in Great Expectations?
    A column map expectation such as `expect_column_values_to_match_regex` evaluates every row of one column and reports how many rows were unexpected. A table expectation such as `expect_table_row_count_to_be_between` or `expect_table_columns_to_match_ordered_list` makes a single assertion about the batch as a whole, so it either passes or fails outright with no unexpected-row count.
  • Is a suite generated by profiling a sample safe to ship as-is?
    No. Profiling proposes expectations from whatever the sample happened to contain, so ranges and value sets encode accidents of that snapshot. Treat it as a first draft: keep the expectations that express a real invariant, delete the coincidences, and widen numeric bounds to what the business actually guarantees. Otherwise the suite becomes noisy and people start ignoring failures.
  • What does the result_format setting change?
    It controls how much detail a validation result carries: `BOOLEAN_ONLY` just pass/fail, `BASIC` and `SUMMARY` add counts and a capped sample of unexpected values, `COMPLETE` returns every unexpected value. `COMPLETE` is excellent for debugging one run and hazardous as a default, because a wholly broken column serializes an enormous list into the result store.

A suite is a checklist you hand to an inspector, not the inspection itself - the Validation Result is the filled-in checklist that comes back with the failures circled.

saying these in an interview costs you the question

  • Calling expectations imperative Python asserts rather than serializable config
  • Thinking a failing suite stops the pipeline automatically
  • Shipping profiler output unedited as the production suite
  • Believing Great Expectations cleans or corrects bad rows
  • Confusing Data Docs with an alerting or monitoring system

context

open as a page

How do you make a Great Expectations checkpoint actually fail the pipeline task that runs it?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Read the returned result's success flag and raise from the task - Great Expectations reports, it does not enforce. Its actions store results and rebuild Data Docs whether validation passed or failed, so nothing stops on its own.

open as a page

In Great Expectations, what does a Checkpoint add beyond an Expectation Suite?

level: middleimportance: should knowfreq 58%

basics

~20 s

A Checkpoint is the runnable unit: it binds an Expectation Suite to a specific batch of data, executes the validation, and fires configured actions such as storing the result and rebuilding Data Docs. The suite alone only declares rules.

open as a page

In Great Expectations, what does the mostly parameter change about an expectation's result?

level: middleimportance: should knowfreq 42%

basics

~20 s

mostly sets a tolerance: the expectation passes when at least that fraction of evaluated rows satisfy the condition, instead of requiring every row. It applies to row-wise column expectations, not to aggregate or table-level ones.

open as a page

When Great Expectations validates a billion-row warehouse table, where does the computation run?

level: seniorimportance: should knowfreq 35%

basics

~20 s

It depends on the backend. Against a SQL data source, expectations are translated into queries the warehouse executes, and only aggregates come back. Against the pandas backend, the batch is pulled into the Python process's memory first.

open as a page

Why can a Great Expectations regex expectation pass on a column that is entirely null?

level: middleimportance: nice to knowfreq 32%

basics

~10 s

Most row-wise value expectations evaluate only non-null values, so a fully null column offers nothing to violate the pattern and the check passes vacuously. Assert completeness separately with expect_column_values_to_not_be_null.

open as a page