skip to content

When a column of address lines is split on a separator into several columns, what fixes how many columns come back?

level: middleimportance: should knowfreq 52%

answer

  1. a rectangle needs one width
  2. rows differ in part count
  3. widest row, or a declared count
  4. trailing positions get no value
  5. a sample's width is not the file's

basics

~20 s

Nothing in a single row can fix it: the result is a rectangle, so one width must cover every row. Designs resolve that by scanning the column and taking the widest row, by making you declare the width or the output names, or by refusing ragged input outright.

solid answer

~40 s

A split produces a different number of parts per row, but several columns is a rectangle, and a rectangle needs one width. Three resolutions travel around this family. Scan-and-widen reads the whole column, takes the largest part count and makes that many columns — which means the result's width is an output of the data and can change between a sample and the full file. Declare-it makes you name the output columns or give a maximum; short rows then get a no-value marker or an empty value in the trailing positions, and parts beyond the maximum are either dropped or all left in the final column, depending on the tool. The third simply refuses or warns on ragged input. Declare the width you expect and assert the raggedness before you split.

go deeper

for a junior

Recall that several columns is a rectangle: every row gets the same number of positions, and a row with fewer parts leaves its trailing positions with no value.

for a middle

Explain where the width comes from — a pass over the data, a declaration from you, or a refusal — and what each choice does to a job that reruns on new files.

for a senior

Show the check, not the fix: count parts per row before splitting, assert the expected width, and prove that a given position means the same field in every row.

for a principal

The call to own is whether a separator should be carrying schema at all. Splitting on a delimiter at read time is cheap and fragile; a declared shape upstream costs more and fails loudly.

## The split is ragged and the result is not Splitting one text column into several columns is two jobs wearing one name. The first is genuinely per value: cut this value wherever the separator appears. That yields three parts for one row and five for the next. The second job is to present the outcome as **several columns**, and several columns is a rectangle — every row occupies the same number of positions. Those two facts cannot both be satisfied without someone deciding a single width, and no single row contains enough information to decide it. That decision is the whole content of this question, and it is where tools in this family genuinely differ. ## Three ways designs resolve the width | Resolution | How the width is decided | What ragged rows get | What it costs you | |---|---|---|---| | Scan and widen | The tool reads the whole column first, finds the largest part count, and makes that many columns | Trailing positions filled with a no-value marker or an empty value | The width is an output of the data, so it moves between a sample and the full file, and between runs | | Declare it | You name the output columns, or give a maximum count | Same filling for short rows; parts beyond the maximum are dropped or all kept in the final column, depending on the tool | You must know the shape in advance, and guessing low truncates quietly | | Refuse or warn | The operation demands a uniform part count and complains otherwise | An error, or a warning plus one of the behaviours above | The pipeline stops — which is frequently the right outcome | Notice that the first resolution requires a pass over the data **before** the result's shape exists. That is not a detail: it means the operation cannot be planned without looking, and it means a chunked or streaming version of the same job cannot know its own output width from the first chunk. ## What the short rows get, and why "absence" is not one thing here A row with three parts in a four-column result leaves the fourth position with nothing. What goes there is design-dependent: a marker meaning there is no value, or a zero-length text value. They behave differently downstream — one is skipped by things that skip absence, the other is an ordinary value of length zero that participates in comparisons and in concatenation. Check which one your tool produced before writing anything that depends on it. A useful measurement falls straight out of this: **count the non-empty values in the last column**. That number tells you how many rows actually reached the full width, which is the raggedness of your input expressed as one integer. ## Why a data-derived width is a production hazard - **Your downstream column list becomes data-dependent.** Code that reads the third and fourth columns works on today's file and fails on tomorrow's, where no row had a fourth part. - **A sample lies.** The first ten thousand rows of a file are not a random sample of it; they are usually the oldest or the most orderly. A width measured there is routinely too small. - **The separator is doing schema work.** If which position means what depends on how many separators a row happens to contain, then position three means the city in some rows and the postcode in others, and nothing about the result's type or shape will tell you. That last point is the real failure. A width that is merely wrong announces itself eventually; a width that is right while the **meaning of each position shifts by row** produces a table that looks clean and is silently misaligned. ## The practical rule 1. **Declare the width and the names you expect.** Make the shape an input to the code, not an output of the data. 2. **Assert the raggedness first.** Count the parts per row and check the distribution before splitting. One row with six parts where you expected three is a data problem to see, not to absorb. 3. **Keep the remainder rather than dropping it.** If the tool can leave everything past the limit in the final column, prefer that to discarding it, and check that column is empty as an assertion. 4. **Normalise before splitting.** Whitespace around the separator, or a separator that varies in form, turns into extra or misplaced parts. Cleaning first makes the part count mean something. 5. **Stop if the part count is genuinely meaningful per row.** When the number of parts is real information that varies legitimately, forcing it into a fixed rectangle loses it, and the rectangular result is the wrong output shape for that data. ## The claim to be careful with "Short rows are padded with absence and long rows are truncated" is true of some designs and not of others, and both halves vary independently. State it the other way round in an interview: the width has to come from somewhere, here are the three places it comes from, and here is which one my tool uses.

  • Why can a split that takes its width from the data be a problem in a job that runs every day?
    Because the number of columns is then an output of the input. A new file whose widest row has one more part produces an extra column nobody consumes, and a file whose rows are all shorter produces one column fewer, so code addressing a fixed position breaks or, worse, reads a different field. Declaring the width makes the shape stable and turns the surprise into a check.
  • You gave a maximum of three columns and some rows have five parts. What happened to the last two?
    One of two things, depending on the tool: everything after the second separator stayed inside the final column as one value, or the excess was discarded. The two produce different data from identical input, so find out which you are on and assert on it — for instance by checking that no value in the final column still contains the separator.
  • How do you measure raggedness before committing to a width?
    Count the separators per row over the whole column and look at the distribution, not the maximum alone. A single spike at one value with a thin tail tells you to declare that width and route the tail to review; a broad spread tells you the field is not positional at all and splitting into columns is the wrong move.

saying these in an interview costs you the question

  • Assumes the tool infers a stable column count from the data
  • Expects an error when a row has fewer parts than the rest
  • Thinks the width seen on a sample holds for the whole file
  • Believes parts beyond a declared maximum are always kept
  • Assumes every design fills the short rows the same way
  • Ignores that position three can mean different fields in different rows