skip to content

A sales table is narrowed to customers whose total exceeds 1,000; how does that differ from keeping only the individual orders above 1,000?

level: middleimportance: must knowfreq 68%

answer

  1. which unit is being judged
  2. whole group, or one row
  3. the group-level form needs a summary first
  4. surviving group sizes are the tell
  5. same word, two operations

basics

~20 s

The customer total is a condition evaluated once per group: every order of a qualifying customer survives and none of a failing one. The order amount is a condition evaluated once per row, which can leave a customer partly present.

solid answer

~40 s

These are two operations that English spells almost the same way. Testing each customer's total is a condition evaluated **once per group**: the rows must be split by customer first, the group's summary computed, and the verdict then keeps or drops all of that customer's rows together. Testing each order's amount is a condition evaluated **once per row**: no split is needed at all, and a customer can come through with three of their nine orders. The tell is in the survivors — after a group-level condition every surviving customer still has every order they started with, while after a row condition the group sizes have moved. Establish which is meant before writing either, because they answer different questions: which customers count, against which orders count.

go deeper

for a junior

Know that a condition on a group's total keeps or drops all of that group's rows together, while a condition on each record keeps or drops rows one at a time.

for a middle

Explain why the group-level form needs the split first — the value it tests is computed from the rows sharing a key — and that it hands back the original rows, not summaries.

for a senior

Work backwards from the surviving group sizes to say which narrowing actually ran, and catch a report whose stated cohort rule does not match the operation underneath it.

for a principal

Get the definition written once. A cohort rule expressed as a group-level threshold in one place and a row-level one in another yields two headline numbers from one source, and each team will defend its own.

## One English word, two operations Narrowing a table means keeping some of it and dropping the rest, but *the unit being judged* can be either a row or a whole group, and the sentence said out loud rarely makes that clear. 'Drop the small customers' and 'drop the small orders' are one word apart and produce different tables from the same input. - A **condition evaluated once per row** looks at one record's own values. Each row survives or leaves on its own merits. A customer with nine orders can end up with three. - A **condition evaluated once per group** looks at all the rows sharing a key together — their count, their total, their span — and returns a single verdict for that key. Every row of the customer survives, or none does. This is the confusion the subject exists to dissolve, and it is not only a beginner's confusion: in several ecosystems the same call name does both, and which one you get depends on whether it was applied to a table or to a **grouped handle** — what a grouping call hands back before any per-group computation has run, recording which rows belong to which key and nothing else. ## What each one needs before it can run | | Condition once per row | Condition once per group | |---|---|---| | Needs the split first | no | yes | | Input to the test | one record's values | all the rows sharing a key | | Verdict granularity | one row | one whole group | | Group sizes afterwards | changed | unchanged for survivors | | Distinct keys afterwards | possibly fewer | fewer, or the same | The last two rows are the useful ones. A group-level condition cannot change the size of a group it kept, because it never looked at rows individually. A row condition can, and usually does. ## A worked pair Take a table of orders carrying a customer key. Customer A has nine orders totalling 1,800; customer B has two orders totalling 900; customer C has a single order of 1,400. 1. **Keep customers whose total exceeds 1,000.** A and C pass, B fails. The result holds A's nine rows and C's one row: ten rows, two customers, and every order of the survivors intact — including A's small orders, none of which would clear the bar on its own. 2. **Keep orders above 1,000.** Only C's single order qualifies. The result holds one row and one customer. A, whose total is the largest in the table, has gone entirely, because no individual order cleared the bar. The two results share almost nothing, and both are reasonable answers to differently worded questions. ## What a group-level condition returns A surprising number of candidates expect the group-level form to return one row per surviving group, on the grounds that the condition was about groups. It does not. It returns the original rows of the groups that passed, at full width and full length, with nothing folded. Collapsing to one row per group is a separate operation: the condition chooses which groups, the collapse chooses the shape, and wanting one row per surviving customer means doing both. ## Reading a result backwards Given only the output, you can usually tell which one ran: - Every surviving key has exactly the row count it had in the input, so a group-level condition ran, or nothing did. - Some keys are present with fewer rows than before, so a row condition ran. - Keys have disappeared entirely, and either could be responsible: a group-level condition rejected them, or a row condition removed their last surviving row. That last line matters in a diagnosis, because a vanished customer is not by itself evidence of a cohort cut. ## They compose, and the order is a decision Both can appear in one pass: narrow the rows, split, then reject groups. But a row condition applied before the split changes what each group's summary *is*, so it also changes which groups a later group-level condition accepts. Fix the order deliberately and record it; it is not a formatting choice. ## Where this bites - A report described as covering active customers only, which actually dropped small orders, so every per-customer total it shows is lower than the real one. - A cohort defined by a group-level threshold in one place and re-derived with a row-level one in another, producing two different cohort sizes from what everyone believes is one definition. - A group-level condition written where a per-row one was meant, quietly multiplying the surviving row count by roughly the average group size and overwhelming a downstream step that was sized for the smaller result.

  • Why can a condition evaluated once per group not be applied before the split?
    Because the value it tests does not exist yet. The test consumes all the rows sharing a key — their count, total or span — so the rows must be keyed and brought together first. A row condition tests values already present on each record, so it can run at any point before the split.
  • Both narrowings removed customer B. What does that tell you?
    Very little on its own. A group-level condition can reject B outright, and a row condition can remove B's last surviving row. Look instead at the customers that remain: unchanged row counts point to a group-level condition, shrunken ones to a row condition.
  • How would you get one row per qualifying customer rather than all their orders?
    Two steps. The condition evaluated once per group chooses which customers survive and returns their original rows; a collapse then folds each survivor to one row. Selecting groups and reshaping the result are separate decisions, and doing only the first leaves you with the full-length table.

saying these in an interview costs you the question

  • Treats a threshold on a group total as a row condition
  • Says a group-level condition can keep part of a group
  • Expects a group-level condition to return one row per group
  • Assumes both narrowings select the same customers
  • Thinks a vanished key proves a group-level condition ran