skip to content

In Great Expectations, what does the mostly parameter change about an expectation's result?

level: middleimportance: should knowfreq 42%

answer

  1. a tolerance, not a mute button
  2. fraction of rows that must satisfy the rule
  3. only meaningful for row-wise column checks
  4. true counts still appear in the result
  5. hides gradual drift below the threshold

basics

~20 s

mostly sets a tolerance: the expectation passes when at least that fraction of evaluated rows satisfy the condition, instead of requiring every row. It applies to row-wise column expectations, not to aggregate or table-level ones.

solid answer

~50 s

By default a row-wise expectation such as `expect_column_values_to_match_regex` fails if a single row violates it. `mostly` accepts a fraction between 0 and 1 and turns that into a threshold: with `mostly=0.99`, the expectation succeeds as long as at least 99% of evaluated rows satisfy the condition. The result still reports the true `unexpected_count` and `unexpected_percent`, so you keep visibility into the tail - only the pass/fail verdict moves. It applies to **column map** expectations, which have a per-row notion of success; it is meaningless on `expect_table_row_count_to_be_between` or a column-mean expectation, which are single assertions. Use it where a small dirty fraction is a known reality rather than an incident - free-text emails, an optional field the app fills in late. Its danger is that it hides slow degradation: a column drifting from 0.5% bad to 0.98% bad stays green until the day it crosses.

code

python · 5 lines
python
gx.expectations.ExpectColumnValuesToMatchRegex(
    column="email",
    regex=r"^[^@\s]+@[^@\s]+$",
    mostly=0.99,
)

go deeper

for a junior

Recall that mostly is a fraction between 0 and 1 that lets an expectation pass when most - not all - rows satisfy it, and that the default requires every row to pass.

for a middle

Explain the mechanics: it converts the verdict into a threshold on the proportion of evaluated rows, applies only to row-wise column expectations, and leaves the reported counts untouched.

for a senior

Show that you derive the threshold from measured rates and that you watch unexpected_percent as a trend, because a fixed tolerance silently absorbs a threefold degradation until the day it does not.

for a principal

Frame it as declaring an acceptable defect rate on behalf of consumers - who agrees to that rate, who reviews it, and when a tolerance is really an unowned problem that should have a remediation date instead.

## The default is unforgiving, on purpose A column map expectation in Great Expectations evaluates each row of a column and counts how many fail. The default verdict is binary and strict: one unexpected row and `success` is false. That is the right default for genuine invariants - a primary key that is null once is a broken key. But plenty of real columns are not invariants. A free-text `email` field captured from a marketing form will contain garbage at some small rate forever; a `phone` column will hold a handful of unparseable entries; an enrichment field is null for the tiny fraction of rows the vendor could not match. If those columns fail your gate every night, the gate stops meaning anything, and people start rerunning past it out of habit. That is what `mostly` exists to prevent. ## What mostly actually does `mostly` takes a fraction in the range 0 to 1 and converts the verdict into a threshold test. With `mostly=0.99`, the expectation succeeds when the proportion of evaluated rows meeting the condition is at least 0.99. Crucially it changes **only the verdict**. The result block still records `element_count`, `unexpected_count`, `unexpected_percent` and a sample of the actual offending values, so the failing rows remain visible in the validation result and in Data Docs even on a green run. You are declaring an acceptable rate, not blinding the check. ## Where it applies and where it does not It is meaningful only where there is a per-row notion of success - the `expect_column_values_to_*` family: not-null, unique, between, in-set, match-regex, match-strftime-format. It has no meaning on: - **table expectations** like `expect_table_row_count_to_be_between` or `expect_table_columns_to_match_ordered_list`, which make one assertion about the batch; - **column aggregate expectations** like `expect_column_mean_to_be_between`, which reduce the column to a single statistic before comparing. If you find yourself wanting a tolerance on an aggregate, what you actually want is a wider `min_value`/`max_value` band. ## The interaction that surprises people The denominator matters. For most map expectations other than the explicit null checks, missing values are excluded from evaluation - so `mostly` is a fraction of the **non-null** rows, not of the batch. A column that is 90% null and whose remaining 10% is clean will sail past `expect_column_values_to_match_regex(..., mostly=0.99)`, because the nulls were never in the denominator. The fix is to state completeness separately: pair the format expectation with `expect_column_values_to_not_be_null`, itself given a `mostly` if some nullness is legitimate. Two expectations, two thresholds, two clear failure messages. ## The operational hazard: slow drift A threshold converts a continuous signal into a binary one, and that is exactly what makes it dangerous over time. Suppose a column normally runs 0.3% unparseable and you set `mostly=0.99`. An upstream form change pushes it to 0.9%: a threefold degradation, and your gate stays green for months until a bad day tips it over 1% and you get a page with no history to explain it. Two habits mitigate this. First, set the threshold close to observed reality rather than at a round comfortable number, so the headroom is small and deliberate. Second, treat `unexpected_percent` as a metric you look at over time, not merely as an input to a boolean - the trend is where the early warning lives. ## Choosing the number Derive it, do not guess it. Measure the current rate over a few weeks of batches, ask whether that rate is *acceptable* or merely *current*, and set the threshold just above the acceptable rate. If the honest answer is that any bad row is a defect, do not use `mostly` at all - use the strict default and route the rejects, so the pipeline can proceed while the offending rows are captured for someone to fix. A tolerance that is really a resignation should be written down as such, with an owner and a date, not buried as a float in a suite. ## In an interview Say what it does in one line, then show judgment: it is a declared acceptable defect rate, the result still reports the true counts, it applies to row-wise expectations only, its denominator excludes nulls for most expectations, and its cost is that it masks gradual degradation unless someone watches the rate.

  • Does mostly hide the failing rows from the validation result?
    No. Only the pass/fail verdict changes. The result block still carries element_count, unexpected_count, unexpected_percent and a sample of the offending values, and Data Docs renders them on a green run. That is precisely why the unexpected rate is worth trending over time rather than reading only the boolean.
  • Why does mostly not apply to expect_table_row_count_to_be_between?
    That is a table-level expectation: it makes a single assertion about the batch, so there is no per-row population over which a fraction could be computed. The same holds for aggregate expectations such as a column mean. If you want slack on those, widen the min and max bounds instead.
  • How would you pick the value rather than guessing it?
    Measure the current unexpected rate over several weeks of real batches, then decide whether that rate is acceptable or merely current. Set the threshold just above the acceptable rate so headroom is small and deliberate. If any bad row is genuinely a defect, skip mostly, keep the strict default, and route the rejects instead.

saying these in an interview costs you the question

  • Describing mostly as suppressing or hiding failing rows
  • Applying it to a table-level or aggregate expectation
  • Picking a round number instead of measuring the real rate
  • Forgetting that nulls are excluded from the denominator
  • Never trending unexpected_percent, so drift goes unnoticed

context