Working With Data in One Process
The work between getting data and reporting on it, done in your own process rather than on a cluster or in a database. A candidate who thinks in whole columns writes different code.
on this pageshowhide
explore
- Tables and Arrays24 questions
- What Labels Add to a Rectangle4 questions
- A Table as Typed Columns4 questions
- Row Labels Operations Align On5 questions
- When Labels Earn Their Cost4 questions
- When a Cell Holds Many Values4 questions
- Data With More Than Two Axes3 questions
- Column Types and What They Cost20 questions
- The Type You Asked For4 questions
- One Column, Two Kinds of Thing4 questions
- Repeated Values Stored as Codes4 questions
- What a Table Actually Occupies4 questions
- The Operation That Changed the Type4 questions
- Whole Columns at a Time27 questions
- Why the Row Loop Is Slow4 questions
- Element-Wise as the Default Unit5 questions
- Stretching a Smaller Operand4 questions
- The Function the Library Cannot See Into5 questions
- Working a Column of Text4 questions
- What Cannot Be Said Column-Wise5 questions
- Getting At Rows and Columns26 questions
- Two Addressing Schemes, One Seam5 questions
- Selecting by Condition4 questions
- Whether You Got Shared Memory4 questions
- The Assignment That Landed Elsewhere5 questions
- Putting the Rows in Order4 questions
- Absent Values22 questions
- Absence in Each Kind of Column4 questions
- What Absence Does to Arithmetic5 questions
- The Default That Moved Your Denominator5 questions
- Three Dispositions, Three Biases4 questions
- Three Things That Are Not the Same4 questions
- Split, Apply, Combine25 questions
- The Shape of the Operation5 questions
- Three Things to Do With a Group4 questions
- Composite Keys and the Result's Labels4 questions
- Numbering Rows Inside a Group4 questions
- Empty Groups and Unobserved Categories4 questions
- Putting Two Tables Together25 questions
- Four Shapes of a Key Match5 questions
- Combining on Labels Instead4 questions
- Adding Rows Against Adding Columns4 questions
- Duplicate Keys and Cardinality4 questions
- Duplicate Rows and Which One Survives4 questions
- The Join That Returned Nothing4 questions
- Long and Wide20 questions
- The Long Layout4 questions
- Spreading a Column, and Its Inverse4 questions
- More Than One Level of Column Label4 questions
- Choosing the Shape Before the Operation4 questions
- The Pivot That Quietly Aggregated4 questions
- Time-Indexed Data23 questions
- Making Time the Index4 questions
- Resampling Up and Down5 questions
- Rolling and Expanding Computations5 questions
- The Comparison Wrong by Hours5 questions
- Aligning Two Differently-Sampled Sequences4 questions
- Getting Data In and Out22 questions
- The Reader's Guess5 questions
- The Chunked Pass5 questions
- What Did Not Survive Being Written4 questions
- Two File Shapes4 questions
- Bad Lines and Silent Truncation4 questions
- Saying It With a Chart22 questions
- What Varies, What It Maps To4 questions
- The Statistic the Chart Computed5 questions
- Matching the Mark to the Comparison4 questions
- Faceting Instead of More Colours4 questions
- Axis Starts and Overplotting5 questions
- When One Machine Is Not Enough18 questions
- Memory Against File Size4 questions
- Which Aggregates Allow Chunking3 questions
- Operations That Cannot Stream4 questions
- Planning Before Running4 questions
- When to Stop Optimising One Machine3 questions
- Trusting the Result21 questions
- Out-of-Order State4 questions
- Row Counts and Key Uniqueness as Checks4 questions
- Diffing Outputs When Logic Changes5 questions
- The Column That Became Text4 questions
- Finding Where Records Were Dropped4 questions
questions
295 · 13 sectionsIn a labelled table typed one column at a time, what is stored per column, and what is a row?
basics
~20 sA labelled table stores fields, not records: each named column is one run of values in a single representation. A row is not stored as a unit — it is a cross-section of every column, assembled on demand.
A dataset sits in a table with named typed columns, or in a rectangle addressed only by position - what do the names buy?
basics
~20 sNames and per-field types make a rectangle self-describing: every field says what it is, and each field can hold a different kind of value. A rectangle addressed only by position carries one representation for everything and offsets to address it by.
How can a two-axis table hold a figure recorded for every store, every month and every measure?
basics
~20 sA two-axis table cannot hold three coordinates directly, so it flattens one: repeat store and month down the rows beside a measure column, or fold the measure into the column labels. A holding with a third axis keeps all three.
A labelled table stores a per-row key beside its values. What does that key let you do, and what does nothing check about it?
basics
~20 sA row label is a per-row key stored beside the values, so a row can be named rather than counted. Nothing checks it for uniqueness, duplicates are legal, and some tabular designs carry no row identity at all.
A table of four numeric fields and one text field is converted to a uniform-type rectangle — what happens to the numbers?
basics
~20 sA uniform-type rectangle holds one representation for every cell, so the single text field decides it for all of them. The numeric values stop being packed numbers and are re-held in whatever form can also carry text.
A column holds 10 million values each stored in the same declared 8 bytes — how much does it occupy, and what is left out?
basics
~20 sEighty million bytes: the declared width times the row count. That figure is the packed value buffer only — anything stored beside it, such as an absence mask, the held-once values behind codes, or the table's row labels, is extra.
A table is built from in-memory records whose account identifier is all digits — why state that column's type?
basics
~20 sStating the column's type at construction fixes it as text before any value is stored, so leading zeros survive and long identifiers stay exact. Left undecided, digits are treated as a quantity and both are lost for good.
An average over a column of whole-number counts comes back fractional. Why is the result's representation a property of the operation?
basics
~20 sEvery operation has a result rule that fixes what its output column is made of; the inputs only feed it. An average is a division, defined over fractional values, so whole numbers in still gives fractional out.
A column of a million quantities is built and one entry is a word. Why does that one entry decide how all million are stored?
basics
~20 sA column carries one representation for all its rows, so no row can be special. The misfit forces a whole-column choice: fall back to a form that holds anything, refuse the value, or tag every row with its kind.
A column of 5 million text values is measured two ways and the figures differ tenfold — what storage form makes that possible?
basics
~20 sOne reference per row, with the values allocated separately elsewhere. The column's own buffer is a fixed slot per row, so a shallow count stops there; following the references adds every value's bytes plus its per-allocation bookkeeping.
Two numeric columns of one million values are multiplied by one written expression - what does it hand back, and how many dispatches?
basics
~20 sIt hands back a newly allocated column of one million products, one per position, and leaves both operands as they were. Your program makes one dispatch, but the million-step loop still runs - inside the library's compiled pass.
Why does adding a single number to a 1,000-value column behave differently from adding a 3-value operand to it?
basics
~20 sA one-value operand has nothing to disagree about: every rule reuses it at every position, giving 1,000 results. Three values against 1,000 positions is a real size conflict, and designs resolve it differently - by refusing, by repeating, or by filling with absent values.
A column of 200,000 customer names is folded to one case and trimmed, yet a later check still sees the old values — why?
basics
~20 sThe ordinary whole-column text operation computes a new column and returns it; it does not edit the values where they sit. If nothing captures the result, or it is captured under a name the later check does not read, the table keeps the old values and no error is raised.
Doubling two million numbers by looping over records and assigning each result back: what is paid once per record?
basics
~20 sFour fixed overheads move from once-per-column to once-per-record: a run-time type decision per value, each value handed back as a wrapped language object, a fresh record built to reach its fields, and a result container reallocated as it grows.
A colleague replaced a per-row loop by handing that same function body to a column surface. What decides whether the loop is really gone?
basics
~20 sWhat the surface hands the body. Called once per record, it is the same loop with a crossing into your code at every record. Called once with the whole column, it is one hand-off and the work stays inside compiled code.
When a table is reordered, which of a row label and a row offset still names the same record, and why?
basics
~20 sThe row label still names the same record; the offset does not. A label is an identity the row carries with it, while an offset only describes where the row happens to sit right now.
In a zero-based host, rows 2 through 5 are requested by offset and then by label — why can the two return different row counts?
basics
~20 sA range of offsets in a zero-based host is half-open — the endpoint is excluded — so it returns three rows. A range of labels includes both ends, so it returns four. Same-looking request, different rule.
An ordering step is given region ascending and amount descending — what does each key govern in the result?
basics
~20 sRegion sets the overall order and groups the rows into blocks; amount only arranges rows inside a block where the region is equal. Each key carries its own direction, so mixing ascending and descending in one step is normal.
Two full-length condition columns over the same 10,000 rows must both hold. Why does the host's scalar connective fail, and what combines them instead?
basics
~20 sTwo condition columns combine position by position into a third full-length column, one outcome per row. The host's scalar connective wants a single truth value from each side, and a whole column has none, so it either refuses or quietly answers a different question.
When a 1,000-row table is restricted to rows whose amount exceeds a threshold, what intermediate object is built, and how long is it?
basics
~20 sA condition column is built first: one true-or-false outcome per row, so 1,000 entries long, exactly as long as the table. That full-length column then addresses the table, and the rows whose outcome was true come back whole.
A cell reading zero, a cell with nothing recorded, and a cell holding text of length zero: how do the three differ?
basics
~20 sZero is a measurement whose answer was none, a text value of length zero is a value that was supplied and contains no characters, and only a cell with nothing recorded is absence: no value exists there at all.
A 10,000-row table reports 8,400 when you count one column's values — what are those two numbers, and what does their difference measure?
basics
~20 sThey answer different questions. The row count is how many records the table holds; the count of present values is how many of them recorded something in that column. The difference, 1,600, is that column's number of holes.
A numeric column has 1,500 of 10,000 cells absent, and every one is filled with zero — what does that cost?
basics
~20 sFilling with a constant stacks 1,500 identical values at one point, and where that point sits decides the damage: a zero fill in a column centred above zero drags the average down and leaves filled cells indistinguishable from measured zeros.
A column with 40 absent cells is added element-wise to a complete column of equal length — what does the result hold?
basics
~20 sForty cells of the result hold no value, and the remaining cells hold the sum. An operator meeting an absent operand carries the absence into its output rather than skipping it, substituting a zero, or raising an error.
A mean over a 10,000-row column with 1,600 cells recording nothing returns a number — which denominator did it use?
basics
~20 sThe count of present values, 8,400 — not the 10,000 rows. An aggregate that steps over the holes removes them first, so the reported number is an average of what was recorded, over a population smaller than the table.
A 10,000-row table has 50 distinct grouping-key values; how many rows come back from a collapse against a summary aligned onto every row?
basics
~20 sCollapsing to one row per group returns 50 rows, one per distinct key value. Aligning that same summary back onto its rows returns all 10,000, with each row carrying its own group's value beside it.
Computing the largest order value per customer over a long table, what does the run keep for each customer as it reads?
basics
~20 sOne running value per customer — the largest amount seen so far, replaced when a bigger one arrives. Each record is folded in and then dropped, so the room used tracks the number of customers, not the number of orders.
After grouping orders by both region and payment method, what identifies each row of the collapsed result, and how many rows are there?
basics
~20 sEach result row is identified by a pair, one value from each grouping key, and the result holds exactly one row per pair that actually occurs in the data, not one row per pair the two columns could form.
A grouped count over a table of orders returns 11 rows though the business defines 14 categories — why?
basics
~20 sA grouped result carries one row per key value that actually occurred, so the three unlisted categories had no rows in this input. They are absent from the result rather than present with a zero.
A grouped aggregation over a table is described as three phases — what happens in each?
basics
~10 sSplit, apply, combine. The split gives every row a grouping key and records which rows share each key; the apply runs one computation per key; the combine assembles those per-key answers into a result.
Two columns of 1,000 values each, taken from different tables, are added together — what decides which value pairs with which?
basics
~10 sThe tool decides, not the expression. Designs that carry a per-row identifier pair the two operands by that identifier first; designs that carry none pair strictly by position and refuse operands of different lengths.
A match of two tables on their key columns returns zero rows, though the same codes print on both sides. Why?
basics
~20 sEquality compares stored values, not the rendering. Two codes that print alike can differ by letter case, leading or trailing spaces, an invisible character, a lost leading zero, or one side holding digits as text and the other as numbers.
Matching 1,000 orders to a customer list on customer id returns 940 rows: which match shape was used, and what do the others keep?
basics
~20 sLosing 60 rows means the match kept only rows partnered on both sides. The three alternatives keep the orders table whole, the customer list whole, or both whole, and unpartnered rows then survive with the other side's columns holding the absent-value marker.
Twelve monthly pieces are stacked into one longer table - how does that differ from gluing two tables side by side?
basics
~20 sStacking adds records, so the result is longer and the hazard is two pieces whose column sets do not agree. Gluing side by side adds attributes, so the result is wider and the hazard is rows paired in an order nobody checked.
A match of 1,000 order rows against a customer reference table returns 1,340 rows - how is that possible?
basics
~20 sA key match returns, for each key value, the first side's occurrence count multiplied by the second side's. If a customer identifier appears more than once in the reference table, the orders carrying it are copied and the total grows.
A table has columns store, date, measure and amount, with a row reading store 7, 2026-03-01, footfall, 812 — what does one row represent?
basics
~20 sOne measurement: store and date say which subject and moment, measure names what was counted, amount carries the number. That is the long layout — one row per measurement, with the measure's name in a cell, not in a header.
What must be true of a table before one column's distinct values can safely become new headers filled from a second column?
basics
~20 sEach combination of row key and new header must occur exactly once. The result has one slot per combination, so two input rows sharing the same identifying values and the same header value leave one cell with two candidate contents.
When a table is widened into one column per measure using two name columns instead of one, what do the resulting column labels look like?
basics
~20 sEach output column is identified by two name parts, one from each name column, with one column per pair the data holds. Whether those parts stay separately addressable or are composed into a single string depends on the tool's label space.
A table with one row per measurement is widened to one column per measure. Which three roles must its columns be assigned?
basics
~20 sThree roles: the identifying columns that stay put as the row key, the header source whose distinct values become the new column headers, and the cell source whose entries fill those cells. Folding back reads the same declaration backwards.
One row per measurement, or one column per measure — which layout does comparing two measures against each other read more naturally?
basics
~20 sThe wide layout — one column per measure — puts both measures on the same row, so the comparison is one column-against-column step. In the long layout, one row per measurement, the two numbers sit in different rows and must be brought together first.
Daily sensor readings are re-spaced onto a monthly grid, and one month holds no readings. What can appear at that position?
basics
~20 sOne of three things, depending on the design: a row carrying an absent value, a row the same call already filled, or no row at all. Confirm which by comparing the returned row count with the months the span should contain.
What does making the timestamp the row labels of a table buy, and what does it cost?
basics
~20 sPromoting the timestamp makes it what rows are named and ordered by, so a range can be asked for by naming two moments and time operations find the ordering unasked. Only one of the record's times can hold that slot.
A column of timestamps arrives with no zone attached. What does each value actually state, and what can you not do with it?
basics
~20 sA stamp with no zone is a clock reading, not a point on the world timeline: it says what a clock showed, not when. Until the recording zone is attached, it cannot be safely compared with, or converted for, another source.
A quote feed and a price feed never share a timestamp, so how is each quote matched to the price in effect then?
basics
~20 sMatch each quote to the most recent price record whose stamp is at or before the quote's stamp - an inexact ordered match rather than an equality match. The driving side sets the row count: at most one matched record each.
Over an ordered sequence, how does a moving window of fixed length differ from a window anchored at the start?
basics
~20 sA moving window of fixed length keeps its edges a fixed distance apart, so old records drop out as new ones enter. A window anchored at the start never moves its left edge, so each answer covers everything so far.
A column of zero-padded account codes comes back as numbers with the zeros gone. What did the reader do, and what would have prevented it?
basics
~20 sEvery field in that column looked like digits, so the reader chose a numeric representation and parsed the characters into values. Leading zeros carry no numeric value, so they were discarded during the read. Stating the column as text prevents it.
A plain text file carries no type information. How does a reader decide each column's type, and why might it decide differently next month?
basics
~20 sA reader over a file with no declared types guesses each column's type from the rows it inspects, then holds every value in that column to that one representation. Different rows next month can produce a different guess.
A file larger than memory is read in pieces on one machine: what is one piece, and where may the cuts fall?
basics
~20 sOne piece is a batch of whole records handed back per turn of the loop, not an arbitrary byte range: cuts must fall on record boundaries, and the line naming the columns belongs to the first turn only.
The same dataset is stored as delimited text and as a self-describing binary columnar file: what does each read have to decide?
basics
~20 sA self-describing binary columnar file carries its own type declaration, so the read applies it. Delimited text carries characters only, so something outside the file decides what each field means — on every read, and possibly differently the next time.
Right after a reader turns a delimited text file into a table, what arithmetic tells you it dropped records?
basics
~20 sCompare rows in the table against a count you knew independently: lines in the file minus the line that names the columns, or the count the producer stated. Equal is reassurance; unequal is the alarm the read never raised.
A chart connects the totals of eight unrelated product categories with one line — what does that line claim?
basics
~20 sA connecting line claims the space between two marks was travelled: that intermediate positions exist and the value passed through them. Between unordered categories nothing lies in between, so the line asserts a path that does not exist.
In a chart, what is the difference between a colour set once for every mark and a colour bound to a column?
basics
~20 sA bound colour is a mapping: the column's values decide each mark's colour, and a scale plus a legend let a reader read them back. A fixed colour is decoration — every mark gets it, and it carries no data.
A distribution chart shows one peak at the tool's default bucket width and three peaks at half that width - which shape is real?
basics
~20 sNeither shape is a property of the data alone. The chart counted rows into buckets before drawing, so the picture is of those counts, and the bucket width - which the tool chose if you did not - decides how many peaks appear.
A bar chart's vertical axis starts at 92 rather than 0, and two bars whose values differ by 4% look three times apart — why?
basics
~20 sA bar encodes value as length from the baseline, so a baseline of 92 makes each length proportional to value minus 92 rather than to the value: 94 and 98 become lengths of 2 and 6.
A colleague asks for 'a chart of the sales data' — what must you settle before choosing between a bar, a line and a point?
basics
~20 sSettle which comparison the reader will make: magnitude, change along an ordered dimension, spread of one column, relationship between two, or rank. The comparison picks the mark, and a mark chosen for looks answers a different question than the one asked.
Each of three steps on a large table returns instantly, then one later call runs for four minutes — why?
basics
~20 sPlan-then-run execution: the three steps only recorded a description of the work, so they returned immediately. The later call is the trigger, where the whole recorded plan finally executes and every byte of work is paid at once.
An exact median over a 40 GB file cannot be answered from summaries of pieces, but a total can — why?
basics
~20 sA total is determined by the records already read, but an exact median is not: a record still unread can move the middle value. Only the total has a combine step that absorbs whatever arrives later.
A 2 GB text file of records is loaded whole and the process now holds far more than 2 GB. What sets that multiple?
basics
~20 sThe in-memory representation of the columns sets the multiple, not the row count. Values held as one runtime object per row, header and reference included, cost several times their text; fixed-width typed columns can cost less than the file did.
A large file is read in pieces of unequal row counts and each piece's mean is recorded. Why is the average of those piece means not the file's mean?
basics
~20 sAveraging piece means gives every piece one vote while the pieces hold different numbers of rows, so the answer drifts toward whatever the small pieces contained. Carry a running total and a running count instead, and divide once at the end.
What does a recorded plan buy on one machine, where there is no network round trip to avoid?
basics
~20 sThree local savings: columns the plan never references are never read or decoded, a row condition can be applied during the read instead of after it, and consecutive element-wise steps can collapse into one traversal with no intermediate buffer.
Why is a row count either side of a cleanup step a better check than scrolling its output?
basics
~20 sA row count is a claim a machine can fail on, while scrolling only samples the rows you happen to see. Counts either side of one step catch rows silently lost or multiplied anywhere in the table.
Eight chained steps take 10,000 rows in and return 9,860 with no error raised — how do you find the step that lost them?
basics
~20 sRecord the number of rows after every step in one instrumented run, then read the sequence for the first adjacent pair where it falls. Two endpoint numbers give the size of a loss and never its location.
A column of numbers now sorts with 10 before 9. What does that tell you, and why did no step raise?
basics
~20 sA column ordered character by character is being held as text rather than as numbers: '10' comes before '9' because '1' comes before '9'. Nothing raised because putting text in order is a legal operation.
Two versions of a transform return the same 40,000 rows in a different order, yet a row-by-row comparison reports 39,000 differences. Why?
basics
~20 sNothing promises that two versions emit rows in the same sequence, so a position-by-position comparison is measuring order rather than correctness. Align both outputs on the columns that identify a row, then compare the matched pairs.