An inexact ordered match over a table of many sensors returned a full, plausible result with no error - which two preconditions went unchecked?
answer
- no error is not a check
- many entities in one table
- the identifier is just a column
- monotonic order is assumed
- prove the identifiers agree per row
basics
~20 sThat the match stayed inside one sensor, and that both inputs were in stamp order. Neither is inferred: without an entity restriction one sensor's reading attaches to another's record, and an out-of-order input is walked as if ordered, producing a plausible table.
solid answer
~50 sTwo things an inexact ordered match depends on and may not verify. First, **matching within an entity**: unless you restrict it, the match looks for the latest record at or before each stamp across the whole other side, so in a table interleaving many sensors it happily attaches sensor B's reading to sensor A's record. The identifier is never inferred from the columns - you supply the restriction, or you split both sides by entity and match each group separately. Second, **stamp order on both sides**: the match walks the inputs assuming a monotonic ordering. Some designs verify that and refuse; others document it and do not check, walking an unordered input and returning a complete, plausible, wrong table. Prove both yourself: assert the matched identifier equals the driving one, and assert monotonic order on each input before the match.
go deeper
Recall that a match by stamp over a table holding many sensors will happily pair records from different sensors, and that the operation assumes both inputs are already in stamp order.
Explain why the cross-entity result looks denser and more complete than the correct one, and why the ordered walk that makes the match cheap is exactly what makes an unordered input silently wrong.
Demonstrate the assertions: matched identifier equals driving identifier on every row, non-decreasing stamps on both inputs before the match, row count preserved, and one record traced by hand against the raw feeds.
The position to defend is that preconditions nothing verifies belong in the pipeline as assertions rather than in a comment, because the next port may land on a design that checks neither and the failure mode is a plausible table.
## Why the output looks fine Both of these failures produce a result with the right number of rows, the right columns, the right types and values in the right range. No count is off. Nothing raises. That is what makes them worth an interview question: the candidate who checks "did it error?" and "is the row count right?" has already passed both. ## Precondition one: the match must stay inside one entity A table of readings from 300 sensors, sorted by stamp, interleaves all 300 sensors' records. An inexact ordered match asks a purely temporal question - *what is the latest record at or before this stamp?* - and the answer to that question, over an interleaved table, is usually a record belonging to a different sensor. - The identifier column is just another column. Nothing about it makes the match treat it as a boundary. - The result is dense and plausible: every driving record gets a reading, because with 300 sensors there is nearly always something a few milliseconds earlier. - It is *more* plausible than the correct answer, which has larger gaps and more absences, because each sensor individually reports far less often than the table as a whole. The fix is **matching within an entity**: the match never reaches across into another sensor, account or instrument. Some designs take that as an argument on the operation. Where none exists, split both sides by the identifier and run the match per group, then recombine - the same restriction, expressed as structure instead of as a parameter. Note that sorting by identifier and then by stamp does **not** impose it: ordering the input changes which record is adjacent, not which records the match is allowed to consider. ## Precondition two: both sides in stamp order The match is defined over an ordering. Implementations walk the two inputs in that order, which is why the operation is close to linear rather than quadratic. Hand it an input that is not monotonic and the walk still completes - it simply produces the wrong record for the rows around each inversion. Whether you find out varies by design, and this is exactly the kind of claim worth stating carefully: - Some designs check monotonic order on both inputs and refuse to proceed when it does not hold. There the mistake is loud at the point you make it. - Some document the precondition and do not verify it. There you get a complete table with quietly wrong values. Because you cannot rely on which you have - and a pipeline that runs on one today may run on the other after a port - verify it yourself. It is one cheap pass per input. ## The two failures side by side | precondition | what goes wrong when it does not hold | how to prove it held | |---|---|---| | the match stays inside one entity | readings are attributed to the wrong sensor, denser and more complete-looking than the truth | carry the matched side's identifier into the output and assert it equals the driving identifier on every row | | both inputs in stamp order | the walk produces the wrong record around every inversion, with no error | assert the stamp column is non-decreasing on each input before the match, and re-assert after any step that could disturb it | A third thing is often mentioned in the same breath but is not a precondition: how old a matched value may be. That is a policy you choose, not an assumption the operation makes, and an age cap would not have caught the cross-sensor match at all - the other sensor's reading is recent, which is precisely why it was chosen. ## The assertions to write 1. Before the match, assert monotonic stamps on both inputs. If the data arrived from several files or partitions, assert it after the combine, not before. 2. Restrict the match to the entity, by argument or by splitting, and assert afterwards that the matched identifier equals the driving identifier on every row. 3. Assert the row count equals the driving side's. 4. Spot-check one entity by hand against the raw feeds - a single traced record catches all three failures at once and takes two minutes. The habit that separates a senior answer here is refusing to treat "it produced a full table" as evidence of anything. Silence is the expected behaviour of both of these bugs.
- Does sorting the table by identifier and then by stamp restrict the match to one sensor?No. Ordering changes which records are adjacent, not which records the match may consider - and it makes things worse, since the whole of one sensor's history now precedes another's. The restriction has to be expressed as part of the match, or by splitting both sides and matching each group separately.
- Would a staleness tolerance have caught the cross-sensor match?No. The wrongly attached reading came from a different sensor milliseconds earlier, so it is as fresh as any correct match would be. An age cap catches a dead feed; it says nothing about which entity the value came from.
- How do you prove after the fact that both inputs were in stamp order?Assert the stamp column is non-decreasing on each input immediately before the match, and again after any step that could disturb it - a combine of several files or partitions, or an operation that returns rows in an unspecified order. Record the assertion in the pipeline rather than running it once by hand.
saying these in an interview costs you the question
- Assumes the match infers the sensor identifier from the columns
- Believes an out-of-order input would always raise an error
- Sorts by identifier then stamp and calls that an entity restriction
- Trusts a plausible row count as evidence the match was correct
- Sorts only the driving side and leaves the other feed as it arrived
- Thinks an age cap would have caught the cross-sensor match