skip to content

Right after a reader turns a delimited text file into a table, what arithmetic tells you it dropped records?

level: juniorimportance: must knowfreq 66%

answer

  1. silence from the reader proves nothing
  2. two numbers, compared immediately
  3. rows in table against lines in file
  4. subtract the line naming the columns
  5. an equal count can still hide damage

basics

~20 s

Compare rows in the table against a count you knew independently: lines in the file minus the line that names the columns, or the count the producer stated. Equal is reassurance; unequal is the alarm the read never raised.

solid answer

~40 s

Immediately after the read, compare two numbers: how many rows the table has, and how many records you expected. The expectation has to come from outside the reader — a plain pass counting lines in the file, less the first line if it names the columns, or a record count the producing system states. Run it before any transform, because after one, a filter or a match is also a suspect. The reason this is arithmetic rather than vigilance is that readers differ in how they announce a line they could not parse: some stop outright, some write a message to a warning stream that a scheduled run discards, some record it only in a diagnostics report you have to ask for. **Silence is not evidence.**

code

pseudocode · 11 lines
pseudocode
lines_in_file    = count_text_lines(path)      # physical lines, naming line included
expected_records = lines_in_file - 1           # sound only if no value contains a line break

table      = read_delimited_text(path)
rows_taken = row_count(table)

if rows_taken != expected_records:
    fail("read produced " + rows_taken + " rows from "
         + expected_records + " expected records")

# every later step runs only past this point

go deeper

for a junior

Know that a table can come back short with no error anywhere, and that the answer is counting. Say the comparison out loud: rows in the table against lines in the file, less the line that names the columns.

for a middle

Explain why the expectation must come from outside the reader, and name the legitimate reasons lines and records differ, so that a false alarm is investigated rather than mistaken for a loss or waved away as noise.

for a senior

Show that the comparison runs before the first transform, and that you know its blind spot: a reader that restores a broken line's shape keeps the row count while losing the values inside it.

for a principal

The tradeoff is what a stopped pipeline costs against what a quietly short answer costs, and how much friction is worth accepting from a producing team to get a stated record count at all.

## Three counts, and only one of them is on your screen A **reader** — the call that turns a file into a table in one step — hides a great many decisions behind a single line of code, and how many records it actually produced is one of them. When the call returns you are holding exactly one number: the table's row count. At least two others matter, and neither is visible. - **Lines in the file.** Physical lines of text in the bytes on disk. Cheap to get with a separate pass that parses nothing. - **Records the producer believes it sent.** What the system that wrote the file would say if asked. Sometimes stated somewhere alongside the file, often not stated at all. - **Rows in the table.** What every later number will be computed from. | Count | Where it comes from | What it is evidence of | |---|---|---| | Lines in the file | a plain pass over the bytes that does no parsing | how much text there is, not how many records | | Records the producer states | the system that wrote the file | what was supposed to arrive | | Rows in the table | the object the read handed back | what you will actually compute on | The check is the comparison between the third and one of the first two. It costs a second and it is the only thing standing between a quiet loss and a published figure. ## Why it has to be arithmetic rather than attention The intuitive defence is to notice. That fails for a structural reason: **what a reader does when it meets a line it cannot parse, and where it says so, is a property of the tool and not of the subject.** Some designs raise on the first bad line. Some drop the record and emit a message to a warning stream. Some collect the problems into a report that exists only if you ask for it. Some say nothing at all. The same tool has changed which of these it does between releases. So the fact that your run printed nothing red carries almost no information. A count, by contrast, is produced by you, from a source the reader did not touch, and it is either equal or it is not. ## The comparison, step by step 1. **Establish an expectation before or beside the read.** A line count over the raw bytes, or a record count the producer states. What matters is that it does not come from the read. 2. **Take the row count immediately**, before the first selection, match or reshape. 3. **Stop on inequality.** Not log it — stop, and go and find out which of the innocent explanations below applies, or whether none does. ## Where lines and records legitimately differ A line-based expectation is cheap and slightly wrong, and knowing exactly how it is wrong is what stops a false alarm being dismissed as noise: - **The first line names the columns.** It is a line and not a row, so subtract one. An off-by-one here is the single commonest false alarm. - **A value contains a line break.** One record then occupies two physical lines, and a reader that handles it correctly gives you one row from two lines. Rows fall below lines by exactly the number of such breaks, and nothing is lost. - **A trailing blank line** at the end of the file is a line that is not a record. - **A file assembled from several extracts** may carry the naming line more than once, so some of those lines are neither the table's labels nor real records. A count the producer states is stronger than `lines - 1` for exactly these reasons, when you can get one. ## What a mismatch proves, and what an equal count does not - **Fewer rows than expected** means records did not survive the read, or the file held fewer than you thought. Either way nothing downstream is trustworthy yet. - **More rows than expected** almost always means the expectation was built wrongly — a line count taken by a pass that treated a break inside a value as a record boundary, or a producer count that excluded records it had already filtered. - **Equal counts do not prove a faithful read.** This is the honest limit of the check. Where a reader responds to a mis-shaped line by restoring the shape — padding the fields it did not find with the absent-value marker, the placeholder a tool puts in a cell with no value, or discarding a surplus field — the record survives as a row with wrong values in it. The count passes and the data is wrong. That limit is a reason to add a second cheap look at the values, not a reason to skip the count. The count is the one control that costs nothing and catches the failure mode that otherwise reaches a report with no fingerprints on it.

  • The file has 1,000,000 lines and the table has 999,999 rows. What is the least alarming explanation?
    The first line named the columns rather than carrying data, so it is a line in the file and not a row in the table. That is an off-by-one you should have predicted rather than a loss. Confirm it by checking that the table's column labels came from that line; if they did not, a record really is gone.
  • What if the table has more rows than the file has lines?
    Usually the expectation is wrong rather than the read. The line count may have been taken by a pass that treated a line break inside a value as a record boundary, or the producer's stated count may have excluded records it had already filtered. Some readers also split a line carrying surplus fields into an extra record instead of rejecting it, so establish what yours does before blaming either count.
  • You have no independent count at all. What then?
    Then the read is unverified and you should say so rather than imply otherwise. The cheap substitute is a second pass of a deliberately different kind — count lines as plain text and compare — together with setting the reader to stop on a line it cannot parse, so a loss becomes a failure instead of a silence.

saying these in an interview costs you the question

  • Assumes a read that raised nothing lost nothing
  • Checks the row count only after several transforms
  • Trusts a warning stream that a scheduled run discards
  • Compares rows to lines and forgets the naming line
  • Believes matching counts prove the values are intact
  • Eyeballs the first and last rows instead of counting