skip to content

A condition column over 1,000 rows holds 412 true, 500 false and 88 outcomes that are neither — how many rows come back?

level: middleimportance: must knowfreq 63%

answer

  1. not every outcome is true or false
  2. the third outcome must go somewhere
  3. designs disagree, and two are silent
  4. resolve it before addressing rows

basics

~20 s

412, 500, or an error — the design decides. Most treat an outcome that is neither true nor false as not-true and return 412; some refuse the operation; some emit a placeholder row for each, returning 500 of which 88 hold no values.

solid answer

~50 s

The question has no single answer, and that is the point. A *condition column* — one outcome per row, addressed back against the table — can carry a third outcome that is neither true nor false, and the 88 rows behind it are by definition neither kept nor dropped. Three real behaviours exist: treat not-true as drop and return **412**; refuse the operation and raise; or return **500** rows, the extra 88 carrying no values at all. Two of the three are silent. The consequence to name in an interview is that restricting on a condition and restricting on its opposite no longer partition the table: the 88 rows fall out of both halves, so the counts do not add to 1,000. The fix is to resolve the third outcome to true or false yourself, before the addressing step.

go deeper

for a junior

Know that an outcome per row can be neither true nor false, and that those rows have to go somewhere. Do not answer a row-count question with a single number until you have accounted for them.

for a middle

Explain the three behaviours — drop, refuse, emit a placeholder row — and that two of them are silent. Then give the consequence: a condition and its apparent opposite do not cover the table between them.

for a senior

When a total does not reconcile, compute how many outcomes were undecidable before touching the condition itself. It either explains the gap or eliminates the whole class, and it costs one step.

for a principal

The argument worth having is whether undecidable rows should be a per-condition policy stated at each use, or a house rule the whole codebase follows. Per-use is honest about the data; a house rule is reviewable at a glance.

## A condition has three outcomes, not two An outcome computed per row can be true, false, or neither. "Neither" is not a bug and not an error state; it is the honest result when the row carried no value for the condition to decide on. Once the condition column holds a third outcome, the addressing step faces a question that a two-outcome model never poses: **the 88 rows are not marked for keeping and not marked for dropping, so what happens to them?** The answer is a design decision, and the designs in this space genuinely disagree. That is why the row count is the thing that moves. ## The three behaviours | what the design does | rows returned | how you find out | |---|---|---| | treats not-true as drop | 412 | nothing is raised; the count is simply lower than you expected | | refuses the operation | none — it raises | immediately, which is by far the kindest of the three | | emits a placeholder row for each | 500, of which 88 hold no values | nothing is raised; 88 empty rows travel downstream | The third behaviour deserves attention because it is the one people do not predict. The result is *larger* than the intuitive answer, not smaller, and the extra rows look like data until something downstream averages them or matches on them. A row count of 500 against an expectation of 412 reads like a bug in the condition, and the search goes to the wrong place. ## Why the two halves stop adding up This is the checkable consequence, and it is the one to reach for in an interview: - Restrict to *amount exceeds 100* and you get some number of rows. - Restrict to *amount does not exceed 100* and you get another number. - Under the first behaviour above, those two numbers do not add to the table's row count. Every row whose outcome was neither fails **both** conditions, because "neither" is not true in either direction. So a pair of restrictions written as a partition is not one. Aggregate each half, report both, and a quantity that should reconcile against the whole quietly does not — and no step in the chain raised anything. Anybody who has chased a total that is short by an unexplained amount has probably met this. The same asymmetry bites when a condition is negated rather than rewritten. Negating a column of outcomes flips true to false and false to true, but the third outcome is not a truth value to flip: under most designs it stays neither, so the negation is *not* the complement of the original. Two conditions that look like exact opposites can leave the same 88 rows out of both results. ## Resolving it yourself The portable move is to decide, in the code, what an undecidable row means for this particular question — because it is a question about your data, not about your tools: 1. **Choose a policy per condition.** "A row with no recorded amount is not a large-amount row" and "a row with no recorded amount must be reviewed, so keep it" are both defensible, and they give different answers. 2. **Make the third outcome into true or false before addressing.** Once the condition column holds only two outcomes, every design behaves identically and the row count stops depending on which tool is running. 3. **State the policy where a reader will see it.** The reason this class of bug survives review is that the code looks like it has considered the case when it has not. ## What the interviewer is listening for A weak answer is a single number delivered confidently. A strong answer names three things: that an outcome can be neither, that what happens to those rows differs between designs and at least two of the behaviours are silent, and that the repair is to resolve the outcome rather than to memorise a default. Adding the partition consequence — that a condition and its apparent opposite do not cover the table between them — turns a definition into evidence that you have been bitten by it. One more habit follows from all this. When a restriction returns an unexpected count, the first split to make is not "is the condition wrong" but "how many rows had an undecidable outcome". That number is computable directly from the condition column, and it either explains the gap or rules the whole class out in one step.

  • Restricting on a condition and on its negation returns 412 and 500 rows from a 1,000-row table. What explains the 88?
    Rows whose outcome was neither true nor false. Negation flips true and false but leaves the third outcome undecided, so those rows are excluded from both results. The two restrictions look like a partition and are not one, which is why the counts fall short of the total.
  • How do you make the row count the same whichever tool runs the code?
    Resolve the third outcome before the addressing step: decide whether an undecidable row counts as kept or dropped for this question, and turn the outcome into true or false accordingly. With only two outcomes in the condition column, every design agrees on the result.
  • A restriction returns more rows than the true count. What happened?
    Some designs emit a placeholder row for each undecidable outcome rather than dropping it, so the result is the true count plus the undecidable count, with the extras holding no values. The give-away is rows that are entirely empty rather than merely wrong.

A yes-or-no survey returned by 1,000 people, 88 of whom left the box blank. Counting blanks as "no" and counting them as "yes" give different totals, and both are defensible — what is not defensible is reporting a total without saying which you did. The blank is a third answer, and its treatment is a policy you choose, not a fact you look up.

saying these in an interview costs you the question

  • Assumes a condition has only two possible outcomes
  • Expects a condition and its negation to partition the table
  • Treats the excluded rows as noise rather than a bias
  • Assumes every design drops undecidable rows silently
  • Believes something would have been raised if this mattered