skip to content

Tables and Arrays

What a labelled table is, against a plain rectangle of numbers: per-column types, row identity some tools carry and others do not, axes past two, and cells holding many values. The rest assumes it.

on this pageshow

explore

questions

24

In a labelled table typed one column at a time, what is stored per column, and what is a row?

level: juniorimportance: must knowfreq 70%

answer

  1. the holding keeps fields
  2. one representation per field
  3. a field's values are kept together
  4. a record is a cross-section, built on demand

basics

~20 s

A labelled table stores fields, not records: each named column is one run of values in a single representation. A row is not stored as a unit — it is a cross-section of every column, assembled on demand.

solid answer

~40 s

A **labelled table** — a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own — fixes one **column representation** per field: the single physical form every value in that field is stored in. What gets laid down in memory is therefore the field, one column's values kept together. Nothing lays down a record. Asking for row 500 reads one value out of each field's storage and assembles them into a container, which has to be general enough to hold all those representations at once. So per-column typing is what makes a field cheap and packable, and the record is the awkward unit: it crosses every field's storage and is built fresh each time it is asked for.

go deeper

for a junior

Recall the two halves: the type is fixed per column, and one column's values are kept together. A record is something the holding builds when asked, not something it has lying ready.

for a middle

Explain what building a record costs — one read into every field's storage plus a container general enough to hold all of their forms at once — and why that makes the record the awkward unit.

for a senior

Show where it bites: code that walks records over a wide table pays per field per record, and the repair is to keep the work on whole fields or to convert once at a boundary rather than repeatedly.

for a principal

The arguable line is where records are allowed. Decide which layers of a codebase may hand records around and which must stay field-oriented, because the assembly cost is paid at every crossing.

## Two facts that get collapsed into one A **labelled table** is a rectangle whose columns are named and typed one at a time, and whose rows may or may not carry an identity of their own. Two separate claims hide inside that description, and keeping them apart is most of this subject. - **The logical fact:** the type is *per field*. Each named column declares one **column representation** — the single physical form every value in that column is stored in. A cell does not carry a type of its own; it is a value already in whatever form its field declared. - **The physical fact:** the values are laid down *field by field*. What is kept together is one column's values, not one record's. Neither implies the other. A record-oriented holding can perfectly well give every field its own type — plenty of structures do. The tables in this family are not that: they are column-major, and the whole cost model follows from it. ## What is actually stored How one field's values are kept varies, and this is the part people assume rather than check. Three holdings are common: 1. **One run per field** — a single allocation holding that field's values back to back. 2. **A chunked column** — the field held as an ordered sequence of separately allocated runs rather than as one buffer. 3. **One slab per representation** — every field sharing a representation packed into a single two-dimensional allocation, so a field is a slice of a slab rather than a thing of its own. All three satisfy "the type is per field". None of them stores a record. ## A row is a cross-section, not a stored thing | | a field | a record | |---|---|---| | stored as a unit? | yes — a run, a sequence of runs, or a slice of a slab | no | | how many representations | one, fixed for the whole field | as many as the fields it crosses | | getting one | hand back the run or the slice | read every field's storage and build a container | | cost as the table widens | unchanged | grows with the number of fields | When you ask for row 500, nothing is looked up in a record store, because there is none. One value is read out of each field's storage at that offset and the results are assembled. Two consequences follow immediately: - the assembled container has to hold several representations at once, so it is a general-purpose holding that stores a reference per value rather than a packed run — **materialising a record throws away the packing that made the fields cheap**; - the work scales with the **number of fields**, not with the size of the table. A record out of a five-field table is five reads; out of a four-hundred-field table it is four hundred. The container is also not kept. Unless you hold on to it, it is built and discarded, and the next request for the same row builds it again. ## What the shape buys, and what it charges - **Cheap:** narrow storage, because each field picks a form that fits only its own values instead of one form covering everything in the table. - **Cheap:** saying something once for a whole field, because that field's values are already together and already in one form. - **Cheap, as a direction:** adding a field, which is one more run or one more slice and does not read the existing fields. - **Expensive:** anything that wants records, because each one crosses every field's storage and is not reusable afterwards. - **Expensive, in a degree that depends on the holding:** appending records, which has to place a value into every field's storage. ## Where designs disagree Flat statements are dangerous here, because this family genuinely diverges. - "A field is one contiguous buffer" is true of the one-run holding only. A chunked field has no single buffer to point at, and a slice of a slab has one but does not own it. - "Adding a field is cheap" holds as a *direction* in every column-major holding, but the *degree* ranges from one extra allocation to rebuilding the whole slab that every same-representation field lives in. - Whether a record is even nameable differs: designs that carry row identity can name one by its row label, while designs without it can only name a position. What does not vary is the shape of the answer: the holding keeps fields, and a record is work it does when asked. ## Reading a cost surprise When something in a table surprises you, ask which of the two facts it came from. "The values came back in the wrong form" is the logical fact — the representation. "This got slower as we widened the table" or "walking records takes minutes" is the physical one. They have different fixes, and mistaking one for the other produces changes that do not help.

  • If a record is not stored, what exactly does the holding hand you when you ask for one?
    A freshly built container holding one value taken from each field's storage. Because those fields may be in different forms, the container has to be general enough to hold all of them at once, so it is not a packed run the way a field is. It is also discarded unless you keep it, so asking again rebuilds it.
  • Does per-column typing mean two columns can never share one allocation?
    No. Some designs pack every field sharing a representation into a single two-dimensional slab, so a field is a slice of a shared allocation rather than an allocation of its own. The logical picture — one form per named field — is unchanged; only who owns the memory differs, and with it the cost of adding another field of that same form.
  • Why does widening a table hurt record-at-a-time work more than it hurts field-at-a-time work?
    Field work touches one field's storage whatever the table's width. Building a record touches every field, so its cost tracks the field count. Going from five fields to four hundred leaves a field operation unchanged and makes each record eighty times more reads, plus a larger container to build and throw away.

A filing cabinet with one drawer per field: all the surnames in one drawer, all the start dates in another. Nobody's complete file exists as a folder. You make one by pulling a single sheet from every drawer, and the more drawers there are, the longer that takes.

saying these in an interview costs you the question

  • Describes a table as a collection of record objects, one per row
  • Thinks each cell carries its own type, as in a spreadsheet
  • Assumes a record sits together in memory and is handed over as-is
  • Says every design keeps a field as one contiguous run of values
  • Treats per-column typing as a validation rule rather than a storage decision
open as a page

A dataset sits in a table with named typed columns, or in a rectangle addressed only by position - what do the names buy?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Names and per-field types make a rectangle self-describing: every field says what it is, and each field can hold a different kind of value. A rectangle addressed only by position carries one representation for everything and offsets to address it by.

open as a page

How can a two-axis table hold a figure recorded for every store, every month and every measure?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A two-axis table cannot hold three coordinates directly, so it flattens one: repeat store and month down the rows beside a measure column, or fold the measure into the column labels. A holding with a third axis keeps all three.

open as a page

A labelled table stores a per-row key beside its values. What does that key let you do, and what does nothing check about it?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A row label is a per-row key stored beside the values, so a row can be named rather than counted. Nothing checks it for uniqueness, duplicates are legal, and some tabular designs carry no row identity at all.

open as a page

A table of four numeric fields and one text field is converted to a uniform-type rectangle — what happens to the numbers?

level: middleimportance: must knowfreq 62%

basics

~20 s

A uniform-type rectangle holds one representation for every cell, so the single text field decides it for all of them. The numeric values stop being packed numbers and are re-held in whatever form can also carry text.

open as a page

A lookup by row label returned one record in March and returns four today, with the same code. What happened?

level: middleimportance: must knowfreq 66%

basics

~20 s

The label now appears on four rows. Nothing ever checked row labels for uniqueness, so a retrieval returns however many rows carry the label asked for — and the shape of the result is decided by the data, not by the code.

open as a page

An orders table holds each order's item codes in one cell, not one row per code. What does one row then stand for?

level: juniorimportance: should knowfreq 46%

basics

~20 s

One row still stands for one order, so the table's row count is the number of orders, not the number of codes. The variable-length group of codes rides inside a single cell of that row.

open as a page

A labelled table is converted to a rectangle addressed only by position — what now carries each field's meaning, and what silently breaks it?

level: juniorimportance: should knowfreq 56%

basics

~20 s

Column order becomes the meaning: after the conversion each field is identified only by its position in the rectangle. Any upstream change to field order or count silently re-points every downstream reader at the wrong values.

open as a page

When does keeping a variable-length group inside one cell beat repeating the rest of the row once per element?

level: middleimportance: should knowfreq 54%

basics

~20 s

Keep the group in the cell when downstream work is addressed to the row and the group is consumed whole — counted, tested for membership, passed along. Repeat the row when the element is what you filter, match or aggregate on.

open as a page

Why is adding a field to a million-row table usually cheap while appending rows is not, and how much does that vary?

level: middleimportance: should knowfreq 60%

basics

~20 s

Because the holding is organised by field: a new field is one more run and reads none of the others, while a new record has to place a value into every field's storage. That direction is stable; the degree ranges from one extra chunk to a full rebuild.

open as a page

A table gives each column its own type — in what three ways can the values behind one column actually be held?

level: middleimportance: should knowfreq 48%

basics

~20 s

Three holdings are common: one contiguous run per field; an ordered sequence of separately allocated chunks; or one slab per representation, in which a field is a slice. The logical picture is identical in all three; the costs are not.

open as a page

A column named for a money amount holds text, and two rows carry the same row label - what did those names guarantee?

level: middleimportance: should knowfreq 55%

basics

~20 s

Nothing. A field name and a row label are descriptions stored beside the values: no in-process holding compares a name with the field's contents, and none checks row labels for uniqueness. Anything the code depends on, the code must check.

open as a page

The same numbers are held once with column names and row labels, once as a bare rectangle - what does the labelling cost?

level: middleimportance: should knowfreq 58%

basics

~20 s

Three charges: the bytes the labels occupy, the metadata every step carries through and reconciles, and a second way of naming a row that the code must keep straight. Field names are cheap; a row labelling materialised from a field is not.

open as a page

In a three-axis holding, why must code that reduces over the last dimension change its axis number after an earlier reduction?

level: middleimportance: should knowfreq 51%

basics

~20 s

Axes are named by their place in the axis-length tuple, so a reduction that removes one drops an entry and shifts every dimension after it down a place. The number that meant measures now means something else.

open as a page

Setting an identifier column as a ten-million-row table's row labels raised its memory. What was allocated that the default labelling never allocated?

level: middleimportance: should knowfreq 52%

basics

~20 s

A full-length sequence of label values, plus whatever structure the design builds beside it to make lookup by label fast. The default consecutive labelling is usually held as a rule — a start, a step and a length — so it materialises nothing.

open as a page

A pipeline turns each element of a variable-length group in a cell into its own row, repeating the rest of that row alongside it. What does that cost, and when is keeping the group better?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Expanding raises the row count to the total element count and repeats every other value in the row once per element, so per-row figures computed afterwards weight each entity by its group size. Keep the group when the work downstream is per-row.

open as a page

The same conversion to a plain positional rectangle cost nothing on one table and doubled peak memory on another — why?

level: seniorimportance: should knowfreq 45%

basics

~20 s

How the table was already held decides it. Where every field shares one representation in a single contiguous slab, the conversion can hand back a window over existing bytes; otherwise it allocates a new rectangle and copies every value.

open as a page

A script adds twenty small derived fields one at a time to a wide table and each add runs slower than the last — what explains it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Each add is doing work proportional to the table already built, not to the new field — a shared slab rebuilt per add, or a fresh holding materialised per step. Twenty adds then cost twenty table-sized copies. Compute the fields first and attach them once.

open as a page

A result is carried through five filtering and reordering steps and must be reattached to the records it came from - what must travel with the values?

level: seniorimportance: should knowfreq 46%

basics

~20 s

An identity per record - either a row-label structure the design carries for you or an ordinary field you keep and match on. Position is not identity: the first filter or sort breaks correspondence between two holdings, and nothing reports it.

open as a page

A three-axis holding of 40,000 accounts by 60 months by 20 measures has real values in 4% of its cells — how large is it, and when does a repeated-key table cost less?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A dense three-axis holding allocates the product of its axis lengths — 48 million cells here — regardless of how many carry a value. At 4% density the flattened table, storing one address plus one value per real observation, is far smaller.

open as a page

A report selects rows over a span of ordered row labels and slowed sharply once its input began arriving shuffled. Why?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Selecting a span of labels is cheap only when the labels are in order, because then both endpoints can be found by bounded search. Shuffled labels remove that precondition, so the holding must examine every label instead, once per selection.

open as a page

Where in a codebase would you allow the conversion to an unlabelled rectangle, and what would you require at that boundary?

level: principalimportance: should knowfreq 33%

basics

~20 s

Convert late, narrow and in one declared place: at the edge of the routine that genuinely needs a positional rectangle, with an explicit ordered field list named there. Everything on either side of that seam stays named.

open as a page

A column whose cells each hold a variable-length group counts its elements fast in one table, slowly in another. Why?

level: middleimportance: nice to knowfreq 33%

basics

~20 s

The physical representation differs. Packed storage keeps the child values end to end with a run of boundary offsets, so counting is arithmetic over those offsets; a reference-per-cell fallback must visit a separately allocated object for every row.

open as a page

Your team's tables move between a design that carries row identity and one that does not. What standing rule do you set for where identity lives?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

A defensible default: identity lives in an ordinary column, and is promoted into the row labels only inside a step that performs many lookups by name, then demoted before the data is handed on. Write the rule around what you require, not what one tool offers.

open as a page