When a 1,000-row table is restricted to rows whose amount exceeds a threshold, what intermediate object is built, and how long is it?
answer
- two operations, not one
- an outcome computed for every row
- the intermediate is table-length
- address the rows with that column
basics
~20 sA condition column is built first: one true-or-false outcome per row, so 1,000 entries long, exactly as long as the table. That full-length column then addresses the table, and the rows whose outcome was true come back whole.
solid answer
~40 sThe one-line form is two operations. First the condition step produces a **condition column** — one outcome per row, computed for the whole table, so a 1,000-row table gives 1,000 outcomes whatever the answer turns out to be. Second, the addressing step matches that full-length column against the table's rows and returns the rows whose outcome is true, with *all* of their columns, not just the one the condition looked at. The intermediate is as long as the source, never as long as the result: at the moment it is built nothing yet knows which rows will survive. The result's row count is therefore discovered rather than declared — it is the number of true outcomes.
go deeper
Recall the shape: outcomes first, one per row, then the rows whose outcome is true. Say that the result keeps every column and that the intermediate is as long as the source table.
Explain why the intermediate cannot be shorter than the table — at build time nothing knows which rows match — and that the result's row count is the count of true outcomes, discovered rather than declared.
Show that you treat the condition column as an object with a lifetime: it can be named, inspected and reused, and reuse is where a selection starts landing on rows it was never computed for.
The angle worth arguing is whether a codebase writes the two steps separately, buying inspectability and a place to hang a check, or keeps the one-line form and buys brevity at the cost of an invisible intermediate.
## Selection by condition is two operations Keeping the rows of a table where an amount exceeds a threshold reads as one act, and in most tools it is written on one line. Two distinct things happen inside it. 1. **The condition step** produces a *condition column*: one outcome per row, computed for the whole table at once. For a 1,000-row table that is a column of 1,000 entries, each of them true or false — or, sometimes, neither, which is a story of its own. 2. **The addressing step** hands that full-length column back to the table and asks for the rows whose outcome is true. Those rows come back whole — every column of them, not only the column the condition looked at. The two steps are genuinely separable. In most designs you can assign the condition column to a name, look at it, count its true entries, and only then use it to address the table; the one-line form is the same two operations with the intermediate left unnamed. Being able to say that out loud is most of what an interviewer is checking here. ## The intermediate is exactly as long as the table The most common misreading is to picture the condition column as holding only the matches. It cannot. At the moment it is being built nothing knows which rows will match, so an outcome has to be produced for every row. Its length is a property of the **source**, not of the result: - 1,000 rows in means 1,000 outcomes, however many rows survive. - 40 surviving rows do not mean 40 outcomes were computed — 960 false outcomes were computed too, and they are as real as the true ones. - The result's row count is **discovered, not declared**: it equals the number of true outcomes, and it is unknown until the pass finishes. The outcome column is also a real object while it exists. How it is stored varies between designs — one entry per row is the constant, but some keep a byte per row, some pack the outcomes as bits, and some never build the column at all because the condition was folded into the same pass that reads the rows. Nothing makes it shorter than the table. ## What comes back, and what does not | | the condition column | the result | |---|---|---| | length | one entry per row of the source | one row per true outcome | | width | a single column of outcomes | every column the source had | | known when | only after every row is examined | after the addressing step runs | | content | true and false outcomes | the original values, unchanged | Two consequences are routinely missed. First, **restricting rows does not narrow the table**: all the columns come back, including the one the condition was computed from. Choosing a subset of columns by name is a separate act of addressing, written separately. Second, **the values themselves are untouched**. Selection by condition answers *which rows* and nothing else; whether the rows you got back sit in storage of their own or read the original's bytes is a second question the same expression quietly answers. ## Why the two-step framing pays Once the condition column is an object in your head rather than a piece of syntax, three otherwise-baffling behaviours become predictable: - **Cost.** The condition step is proportional to the whole table even when three rows survive, because producing an outcome for every row is the only way to learn which three. - **The third outcome.** An outcome can be neither true nor false. Such rows are by definition neither kept nor dropped, so the design has to choose — and the choices differ, which is how a row count moves with nothing raised. - **Matching.** A full-length column has to be lined up against the rows somehow. Designs that carry row labels may align on those labels; designs without labels match strictly by position. The same expression then selects different rows. None of that is visible to someone who thinks of the operation as "the table, restricted". ## What to say in an interview 1. Name the intermediate and its length: a column of outcomes, one per row of the source. 2. Say the result keeps every column and that its row count is the count of true outcomes. 3. Add that the outcome column can be held, inspected and reused — and that reusing it is where the interesting failures start. That answer takes fifteen seconds and shows a model rather than a habit.
- If the result has 40 rows, how many outcomes were computed?All 1,000 — one per row of the source. The 960 false outcomes cost exactly what the 40 true ones did; only the addressing step scales with the survivors. This is why the result's row count is discovered at the end of a full pass rather than known in advance.
- Why can the result's row count not be declared before the operation runs?Because it is the count of true outcomes, and an outcome exists only once the condition has been evaluated for that row. Until every row has been examined, the number of matches is unknown — even the answer "none match" requires looking at all of them.
saying these in an interview costs you the question
- Thinks evaluation stops once enough matching rows are found
- Describes the condition as a single true-or-false value
- Expects the intermediate to hold only the surviving rows
- Believes nothing is allocated between the comparison and the result
- Assumes restricting rows also narrows the table to one column