In Great Expectations, what is an Expectation Suite and what does validating one produce?
answer
- rules as data, not as assert statements
- a named bundle scoped to one dataset
- names like expect_column_values_to_not_be_null
- output carries counts and offending samples
- rendered into a static HTML report
basics
~20 sAn Expectation Suite is a named collection of declarative assertions about one dataset - not-null, unique, between, matches-regex. Validating it against a batch of data returns a Validation Result with per-expectation pass/fail plus counts and samples of the offending values.
solid answer
~40 sAn **Expectation** is one declarative assertion about data: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`. An **Expectation Suite** is a named, serializable bundle of those assertions describing what one dataset must look like; because it serializes to a config (name plus kwargs) it can be reviewed in a pull request and versioned like code. Validation runs a suite against a *batch* of data and returns a Validation Result: an overall `success` boolean plus, per expectation, its own success flag and a `result` block carrying `element_count`, `unexpected_count`, `unexpected_percent` and a capped sample of the failing values. Those results also render into **Data Docs**, a static HTML report showing what was expected and what actually happened. Great Expectations 1.x reorganized the Python API, but this suite/validation vocabulary carried over from 0.x.
code
python · 13 linesimport great_expectations as gx
context = gx.get_context()
suite = context.suites.add(gx.ExpectationSuite(name="orders_raw"))
suite.add_expectation(
gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeUnique(column="order_id")
)
suite.add_expectation(
gx.expectations.ExpectColumnValuesToBeBetween(column="amount", min_value=0)
)go deeper
Be ready to name three or four real expectation types and say in one sentence what a suite is. Knowing that the output is a pass/fail result with counts of offending rows is enough at this stage.
Explain the mechanics: how a suite serializes to config, what a batch is, and what fields the validation result carries - element_count, unexpected_count, and the sampled unexpected values controlled by result_format.
Show judgment about which datasets deserve suites and how strict to make them. Interviewers listen for the point that a profiled suite unedited becomes alert noise, and that validation without someone acting on success is decoration.
Own the question of where declarative quality rules live across the org - which layer validates, who reviews a change to a suite, and how you keep a growing library of suites from becoming an unowned graveyard of stale thresholds.
## What Great Expectations is Great Expectations (GX) is an open-source Python framework for expressing data-quality rules as **declarative assertions** and running them as a gate inside a data pipeline. Rather than scattering ad-hoc checks like `assert df.order_id.notnull().all()` through transformation code, you declare what the data must look like, store that declaration as a reviewable artifact, and run it against data at a pipeline step or on a schedule. ## An Expectation An Expectation is a single named assertion plus keyword arguments - for example `expect_column_values_to_be_between(column="amount", min_value=0)`. They come in families: - **Column map expectations** evaluate row by row: `expect_column_values_to_not_be_null`, `expect_column_values_to_be_unique`, `expect_column_values_to_be_between`, `expect_column_values_to_match_regex`, `expect_column_values_to_be_in_set`. - **Column aggregate expectations** evaluate one statistic over the column: `expect_column_mean_to_be_between`, `expect_column_distinct_values_to_be_in_set`. - **Table expectations** describe the whole batch: `expect_table_row_count_to_be_between`, `expect_table_columns_to_match_ordered_list`. Every expectation is serializable - the expectation name plus its kwargs - which is precisely why a suite reads as a spec rather than as code. ## The Expectation Suite A suite is a named collection of expectations scoped to one dataset or asset (say, `orders_raw`). It lives in the **Data Context** - GX's entry point and configuration root, either backed by a project directory on disk or created ephemerally in memory - so suites can be stored beside the pipeline, diffed, and reviewed. You author a suite three ways: by hand, interactively against a sample batch while you inspect it, or by profiling a sample so GX proposes expectations from observed ranges and distributions. A profiled suite is a **draft, not a contract**: unedited profiler output encodes accidents of whatever sample it saw, and the usual failure is a suite full of over-tight ranges that fires every Monday. ## What validation produces Validation applies a suite to a **batch** - a defined slice of data, such as one day's partition of a table or one landed file. The output is a validation result for each expectation and an overall suite-level result: - top-level `success` - false if any expectation failed; - per expectation: its own `success`, the expectation's configuration, and a `result` dictionary with `element_count` (rows evaluated), `unexpected_count`, `unexpected_percent`, and a bounded `partial_unexpected_list` of actual offending values. The verbosity of that block is controlled by `result_format`, which ranges from `BOOLEAN_ONLY` through `BASIC` and `SUMMARY` to `COMPLETE`. `COMPLETE` returns every unexpected value, which is wonderful for debugging and dangerous on a wide failure: a column that is entirely wrong will serialize millions of values into the result store. ## Data Docs GX renders suites and validation results into **Data Docs**, a static HTML site: what we expect of this dataset, and what the last runs found. It is the artifact you point an analyst at when they ask why the load was blocked. It is documentation and evidence, not a monitoring system - it has no alerting of its own. ## What a suite does not do by itself Three things trip newcomers up. First, running a suite does not stop anything - somebody must read `success` and act on it. Second, GX validates; it never repairs data. Third, GX does not schedule itself - an orchestrator task, a job step, or a CI job invokes it. ## In an interview Be able to say the vocabulary chain out loud: Data Context holds the configuration; a data source and asset define where the data is; a batch is the slice being checked; an Expectation Suite is the set of rules; validation produces a Validation Result; a Checkpoint is the runnable unit that ties suite to batch and fires actions; Data Docs is the rendered report. Being able to place *where validation executes* in a pipeline is the point of learning the tool's nouns at all.
- What is the difference between a column map expectation and a table expectation in Great Expectations?A column map expectation such as `expect_column_values_to_match_regex` evaluates every row of one column and reports how many rows were unexpected. A table expectation such as `expect_table_row_count_to_be_between` or `expect_table_columns_to_match_ordered_list` makes a single assertion about the batch as a whole, so it either passes or fails outright with no unexpected-row count.
- Is a suite generated by profiling a sample safe to ship as-is?No. Profiling proposes expectations from whatever the sample happened to contain, so ranges and value sets encode accidents of that snapshot. Treat it as a first draft: keep the expectations that express a real invariant, delete the coincidences, and widen numeric bounds to what the business actually guarantees. Otherwise the suite becomes noisy and people start ignoring failures.
- What does the result_format setting change?It controls how much detail a validation result carries: `BOOLEAN_ONLY` just pass/fail, `BASIC` and `SUMMARY` add counts and a capped sample of unexpected values, `COMPLETE` returns every unexpected value. `COMPLETE` is excellent for debugging one run and hazardous as a default, because a wholly broken column serializes an enormous list into the result store.
A suite is a checklist you hand to an inspector, not the inspection itself - the Validation Result is the filled-in checklist that comes back with the failures circled.
saying these in an interview costs you the question
- Calling expectations imperative Python asserts rather than serializable config
- Thinking a failing suite stops the pipeline automatically
- Shipping profiler output unedited as the production suite
- Believing Great Expectations cleans or corrects bad rows
- Confusing Data Docs with an alerting or monitoring system