In a jagged schedule table with a different slot count per day, which grid operations silently break?
answer
- what does width mean here
- row lengths are per-row, not per-table
- which operations need every column present
- computing the width from row zero
- short rows overrun, long rows get truncated
basics
~20 sAnything that treats one row's length as the table's width: column reads, per-column aggregates, transposition, and neighbour lookups. A jagged table has no single column count, so those operations either read past a short row or quietly ignore the tail of a long one.
solid answer
~60 sA jagged table is a sequence of rows whose lengths differ — day one has ten slots, day two has four. Nothing about a row-of-rows representation forbids that, which is why the shape survives until an operation assumes otherwise. The tell is any code that computes a width once, usually from row 0, and then loops `for c in 0..width-1` over every row: on a shorter row it indexes past the end, and on a longer row it silently drops the tail. Column-oriented operations are the ones with no meaning at all here — a per-column aggregate, a transposition, or a `(r+1, c)` neighbour lookup all presuppose that column `c` exists in every row. The choices are to accept jaggedness and always read each row's own length, to rectangularize by padding to the maximum with an explicit absent marker, or to keep a flat values buffer alongside a row-offset array so variable row lengths live in one allocation. Whichever you pick, the constraint belongs in the type and should be checked at construction, not discovered in a column loop.
go deeper
Know that a table built as a sequence of independent rows can have rows of different lengths, and that a loop bounded by one row's length will overrun the short rows and skip the tails of the long ones.
Explain which operations actually require rectangularity — column aggregates, transposition, vertical neighbour lookups — and why a flat single-stride buffer cannot express jaggedness at all while a row-of-rows gets it for free.
Show the resolution you would ship: validate the shape once at construction, choose between staying jagged, padding with an explicit absent marker, or a flat buffer plus a row-offset array, and justify the choice from the actual length distribution.
Own where the invariant lives. Decide whether rectangular and jagged should be distinct types so the assumption cannot be made silently, and weigh the memory cost of padding across a fleet against the maintenance cost of every operation carrying a per-row length.
## What a jagged table is, and why the representation permits it A **rectangular** grid has one width shared by every row. A **jagged** one is a sequence of rows whose lengths are independent — a per-day schedule where Monday has ten bookable slots, Saturday four, and a public holiday none at all. This is not a corrupted grid; it is a faithful model of data that genuinely is not rectangular. Whether the shape is even expressible depends on the representation: - **A row of rows** permits it by construction. Each row is its own sequence with its own length, and nothing checks that the lengths agree. Jaggedness arrives for free, whether or not you wanted it. - **A flat buffer with a single stride** forbids it by construction. The addressing rule `r * stride + c` has exactly one stride, so every row is exactly the same length. You cannot represent a jagged table this way without changing the rule. That asymmetry is the whole story. The representation that makes the shape possible is also the one that offers no place to state the invariant, so "rectangular" ends up as an assumption living in the code rather than a property of the data. ## The operations that break, and how each fails **Width computed once from one row.** The archetype is `width = length(rows[0])` followed by `for c in 0..width-1` across all rows. Two distinct failures follow from the same line: on a row shorter than row 0, the index runs past the end of that row — a detected out-of-range access if you are lucky, a stale read if you are not. On a row longer than row 0, nothing is wrong at the access level and the extra slots are simply never visited. The second is worse, because it produces a plausible answer that is quietly incomplete: a report of "total booked slots" that undercounts every long day. **Per-column aggregates.** "How many bookings at the third slot of the day" presumes a third slot exists on every day. On a jagged table the aggregate is defined only over the subset of rows long enough to have that column, and the code has to decide whether a missing slot means zero, means skip, or means the question was ill-posed. Silently treating missing as zero changes averages; silently skipping changes counts. **Transposition and anything built on it.** Transposing swaps the roles of rows and columns, which requires that every column be a well-defined sequence of the same length. A jagged table has no such thing. A transpose either has to pad, producing holes that later operations must understand, or produce a jagged result whose row lengths are the column occupancy counts — a legitimate but different structure. **Neighbour lookups.** Grid algorithms reach for `(r+1, c)` and `(r, c+1)` as if both always exist. On a jagged table, the vertical neighbour may be absent because the next row is shorter, and every bounds check must consult that row's own length rather than a global width. **Multiplication-like operations.** Any operation defined over conforming dimensions is simply undefined here, and code that computes dimensions from a sample row will produce a confident wrong answer instead of an error. ## The three honest resolutions 1. **Stay jagged and mean it.** Never cache a width. Every loop over a row uses that row's own length; every cross-row operation is written in terms of "rows that have a slot here". This is the truthful model when the data really varies, and the cost is that no column-oriented shortcut is available. 2. **Rectangularize with an explicit absent marker.** Pad every row to the maximum length with a value that means "no such slot", not with a value from the data's own domain. Padding with zero is the trap: zero is a plausible booking count, so aggregates absorb the padding without complaint. The marker must be distinguishable, and every aggregate must be told what to do with it. Padding costs memory proportional to `rows * (max_length - mean_length)`, which is cheap for mild variation and ruinous for one very long day among many short ones. 3. **Flat values plus a row-offset array.** Keep every slot in one contiguous buffer, and a second array of length `rows + 1` giving where each row starts; row `r` occupies `offset[r] .. offset[r+1] - 1`, and its length is the difference. This is the compressed layout used for sparse and ragged data generally. It gives one allocation and contiguous row scans while allowing rows to differ in length. What it gives up is cheap in-place growth of a single row: lengthening row 3 means shifting everything after it and adjusting all later offsets. ## Where the invariant belongs The recurring lesson is that "this table is rectangular" is a **contract**, and contracts that live only in the reader's head get violated. Validate at construction — one pass comparing every row length to the first is linear and runs once — and make the rectangular type distinct from the jagged one so an operation that needs rectangularity cannot be handed a table that is not. A function that quietly derives the width from row 0 is asserting the contract without checking it, which is the most expensive way to hold an assumption.
- If you rectangularize by padding, what is the trap in the choice of padding value?Padding with a value from the data's own domain — zero bookings, an empty slot label — makes absent slots indistinguishable from real ones, so aggregates absorb them silently. Counts, averages and maxima all shift without any error surfacing. The pad must be an explicit absent marker outside the domain, and every operation over the table must be told how to treat it.
- How does a flat values buffer with a row-offset array compare with keeping separately allocated rows?The offset layout keeps one allocation, so a full pass is a single contiguous sweep and each row is still contiguous, while variable lengths remain expressible via `offset[r+1] - offset[r]`. What it loses is cheap independent row growth: lengthening one row shifts every later slot and requires updating all later offsets, whereas separate row storage grows one row without touching the others.
saying these in an interview costs you the question
- Computes the table width once from the first row
- Assumes every row of a table has the same length
- Pads missing slots with zero from the data's own domain
- Thinks transposition is well defined on jagged rows
- Checks bounds against a global width rather than the row's length
- Treats rectangularity as obvious rather than as a checked contract