skip to content

Why do the extreme rows in a transformed dataset stay identifiable when ordinary rows do not?

level: seniorimportance: should knowfreq 30%

answer

  1. The middle of the data is fine
  2. Look at the ends, not the average
  3. Biggest, oldest, first, only
  4. Rank and rarity survive value replacement

basics

~20 s

Being unusual is itself an identifier. The largest account, the only customer in a small region, the row ten times the median: those are recognisable by size or rarity whatever the transformation did to the values, because rank and isolation survive it.

solid answer

~50 s

Transformations are chosen to preserve shape, because the tests need it: amounts stay in range, dates keep their order, categories keep their proportions. Shape is exactly what gives an unusual row away. The biggest balance is still the biggest, the sole row in a small region is still alone, and anyone who knows the business recognises both immediately. Small groups behave the same way -- if only two rows share a rare combination and you are one of them, you have identified the other. So the screen has to be asymmetric: look at the top and bottom of every distribution and at the smallest groups, not at a random sample of rows, which lands in the comfortable middle every time. Then handle the survivors deliberately: fold the row into a wider group, substitute a manufactured row with the same shape, or record it as accepted residual.

code

pseudocode · 11 lines
pseudocode
// screen the ends and the smallest groups, not a random sample
for each column in numericAndDateColumns(dataset):
    flag(topRows(column, 5))
    flag(bottomRows(column, 5))

for each group in groupRows(dataset, by = categoricalCombinations):
    if group.rowCount <= SMALL_GROUP:
        flag(group.rows)

review(flaggedRows)   // widen the group, substitute a manufactured row, or record as residual
// then run the screen again: every change moves the boundary

go deeper

for a junior

Know that replacing values does not make every row anonymous, and that unusual rows -- the biggest, the oldest, the only one of its kind -- deserve extra care. Recognising the idea is enough at this level; nobody expects you to run the screen.

for a middle

Explain why it happens: transformations are chosen to preserve shape, and rank, rarity and group size pass through untouched. Be able to name the three survivors -- the extreme value, the lone member of a category, and the group of two or three.

for a senior

Show the screen you would actually run: the ends of each distribution, the smallest groups, the aggregate one row dominates -- and say what you did with what it found. Interviewers want to hear that you re-ran it after changing the data.

for a principal

Own the standing rule: what group size the organisation treats as too small for this data, whether extremes are substituted or suppressed by default, and how much distribution fidelity you will trade away. Both directions are defensible; say which conditions pick between them.

## A transformation is chosen to keep the data usable Data is transformed so that tests can still run against it, which means the transformation is deliberately shape-preserving. Amounts stay in a plausible range. Dates keep their order. Categories keep their proportions. Text keeps its length. That is not a mistake -- a transformation that destroyed all of it would break the very tests the dataset exists for -- but shape is exactly what makes an unusual row recognisable. | What the transformation preserves | Why it is preserved | How an extreme row leaks through it | | --- | --- | --- | | Rank and magnitude of a figure | Totals, limits and pagination must behave | The largest balance is still the largest | | Category proportions | Branching by category must still be exercised | The only row in a rare category is still alone | | Date ordering and spacing | Ageing, expiry and reporting windows must hold | The oldest account is still the oldest | | Row counts per group | Aggregates and reconciliations must still add up | A group of one is still a group of one | ## Being unusual is itself an identifier Three shapes account for most survivors: - **The extreme value.** The largest account, the highest claim, the longest tenure, the one balance with an extra digit. Anyone who works in the business knows who that is, and in many industries it is public. - **The lone member of a category.** The single customer in a small country, the only account on a withdrawn product, the one row of a rare type. There is nothing to hide behind. - **The tiny group.** Not one row but two or three sharing a rare combination. This feels safer and often is not: if you are one of the three, the other two are no longer anonymous to you. The common thread is that the transformation replaced the *values* while the *position* of the row in the distribution -- its rank, its rarity, its isolation -- passed through untouched. Position is what a reader recognises, and no amount of value replacement disturbs it. ## Screening the ends rather than the middle A review that samples rows at random almost always samples the middle, where everything is fine. The screen that finds survivors is deliberately asymmetric: 1. For every numeric and date column, inspect the top and bottom rows rather than a random sample. 2. For every categorical column, and for every combination the tests rely on, count rows per group and list the smallest groups. 3. For every aggregate a report will display, ask whether one row dominates it. A total that is ninety per cent one customer names that customer. 4. Repeat the screen after every change to the dataset, because each change moves the boundary. ## Handling what the screen finds There is no single correct action, and the choice depends on what the tests need: - **Suppress the row.** Simple, and it distorts totals and removes the boundary condition the row represented. - **Fold it into a wider group.** Widen the region, the range or the period until the group is comfortably large. Keeps a row, loses precision. - **Substitute a manufactured row.** Keep a row with the same magnitude and the same rare category but no real subject behind it. This is usually the best answer when a test genuinely needs an extreme, because the test cares about the shape, not about whose row it is. - **Keep it, and say so.** Sometimes the extreme is unavoidable. Then it belongs in a written residual rather than in a quiet decision nobody recorded. ## Why it never converges in one pass Removing the top ten rows promotes the next ten to being the extremes, and each removal changes group sizes elsewhere, sometimes creating small groups that did not exist before. Treat the screen as a loop, not as a one-time gate: change the dataset, run the screen again, look at what moved. Teams that run it once -- at the moment the transformation was written -- discover months later that a new column, a new region or a refreshed extract quietly recreated the problem. ## The habit worth building When someone says a dataset has been protected, ask two questions: what does the biggest row look like now, and how small is the smallest group. Both are cheap to answer, and both are answered by looking at the data rather than at the process that produced it. Most datasets that fail, fail at the ends -- and the ends are where nobody looks.

  • What do you do when a test genuinely depends on the extreme row you want to remove?
    Replace it rather than delete it. Manufacture a row with the same magnitude and the same rare category, belonging to no real subject, and let the test assert on the shape. If the test depends on the actual value matching another system, that is a signal it should run against a controlled dataset built for it rather than against an extract of real records.
  • Why can suppressing the top ten rows leave the dataset no safer than before?
    Removing the extremes promotes the next rows down to being the extremes, and they are just as recognisable. Each removal also changes group sizes elsewhere and can create small groups that did not exist before. Treat the screen as a loop -- change, re-screen, look at what moved -- and prefer widening groups over trimming rows one at a time.

Blur every face in a team photograph and you still know which one is the person two heads taller than everybody else.

saying these in an interview costs you the question

  • Assumes a replaced value makes the row anonymous
  • Checks only the average row, never the ends of the distribution
  • Thinks removing the single top row settles it
  • Treats a group of two as safe because no column names anyone
  • Believes shuffling a column hides who the largest customer is
  • Samples rows at random and calls that a review