skip to content

In a zero-based host, rows 2 through 5 are requested by offset and then by label — why can the two return different row counts?

level: juniorimportance: must knowfreq 66%

answer

  1. the same bounds, two rules
  2. half-open against closed
  3. endpoint excluded by offset, included by label
  4. a one-based host counts and excludes differently

basics

~20 s

A range of offsets in a zero-based host is half-open — the endpoint is excluded — so it returns three rows. A range of labels includes both ends, so it returns four. Same-looking request, different rule.

solid answer

~50 s

The two schemes carry different endpoint conventions, and they can sit inside one tool. Under offset addressing in a zero-based host the range is half-open: it starts at the offset named and stops before the end offset, giving offsets 2, 3 and 4 — three rows. Under label addressing the range is closed, because the endpoints are values being matched rather than counts being subtracted, so labels 2, 3, 4 and 5 come back — four rows. Neither is a bug: the half-open rule makes the row count a simple subtraction and makes abutting ranges non-overlapping, while excluding a label you explicitly named would be bizarre. In a one-based host, positions start at 1, ranges include both ends, and a negative position does not address from the end at all — it removes that element.

code

pseudocode · 7 lines
pseudocode
# zero-based host, offsets count from 0
rows addressed by OFFSET, from 2 up to 5   ->  offsets 2, 3, 4        ->  3 rows
rows addressed by LABEL,  from 2 up to 5   ->  labels  2, 3, 4, 5     ->  4 rows

# one-based host, positions count from 1
rows addressed by POSITION, from 2 up to 5 ->  positions 2, 3, 4, 5   ->  4 rows
rows addressed by POSITION, -3             ->  everything EXCEPT the third row

go deeper

for a junior

Know the two rules cold and state them with the convention attached: a range of offsets in a zero-based host excludes its endpoint, a range of labels includes both ends. Being able to give the two row counts is the expected answer.

for a middle

Explain why each convention exists — subtraction and tiling for half-open offsets, matched values for closed labels — and what makes the two expressions indistinguishable when the row labels happen to be integers.

for a senior

Bring the operational angle: chunk boundaries that overlap by one row per chunk, a quarter-end day dropped from one report and not another, and the row-count assertion that turns a silent off-by-one into an immediate failure.

for a principal

Treat it as a review rule rather than a fact: ranges expressed as a start plus a count, row counts asserted at boundaries, and a stated convention for which scheme a date window is written in, so two teams cannot disagree by one row.

## The same words, two rules "Rows 2 through 5" is not one request. It is two, and inside a single tool the two return different row counts. - Under **offset addressing** in a zero-based host, a range is **half-open**: it begins at the offset you named and stops *before* the one you named as the end. Offsets 2, 3 and 4 — **three rows**. - Under **label addressing**, a range of labels is **closed**: both ends are included, because the endpoints are values being matched rather than counts being subtracted. Labels 2, 3, 4 and 5 — **four rows**. Neither convention is a mistake. The half-open rule makes offset arithmetic clean: the row count is the difference of the two numbers, two ranges that abut do not overlap, and an empty range is natural to write. The closed rule is the only sensible one for labels, because a label range says "from this row to that row", and excluding a row you explicitly named would be astonishing. ## The conventions that vary | convention | zero-based host, offsets | same host, labels | one-based host in this family | |---|---|---|---| | first position | 0 | not applicable — labels are values | 1 | | range endpoint | excluded | included | included | | rows returned by "2 to 5" | 3 | 4 | 4 | | a negative position | counts back from the end | not applicable | **removes** that element | The last row is the one that catches people moving between ecosystems, and it deserves to be said plainly: in a one-based host of this family a negative position is not an address at all, it is an exclusion. The same expression that means "the last row" in one host means "everything except that row" in another. That is not an off-by-one; it is the opposite operation. ## Where this actually shows up 1. **Chunk boundaries.** Code walks a large table in pieces by advancing a start and an end. Half-open ranges tile perfectly; closed ranges overlap by one row at every boundary, so a running total double-counts one record per chunk and the error grows with the number of chunks. 2. **Reported windows.** "The first quarter" expressed as a range of offsets quietly drops the last day; expressed as a range of labels over dates it includes it. Two reports built from the same table disagree by one row and nobody can see why from the source. 3. **Porting.** A range copied from code written for another host keeps its numbers and changes its meaning, and because both versions run and both produce plausible output, the test that catches it is a row-count assertion, not an exception. ## When the two look literally identical The conventions are easy to keep straight when labels are text and offsets are numbers, because the two expressions do not resemble each other. They become genuinely dangerous when **the row labels are themselves integers**. Then "2 to 5" is the same characters under both schemes, and which one you get is decided by which addressing surface you used — or, under a single overloaded subscript, inferred for you. A second variation to know about: a range of labels only makes sense if the labels are in a known order. Designs differ in what they do when they are not. Some require the labels to be sorted and raise otherwise; some take everything between the first occurrence of each endpoint, which is a positional answer wearing label clothes; some reject a label range over non-unique labels entirely. Do not assume; if the labels are unsorted, say what you expect. ## What to write instead - State the count, not just the bounds. "Take three rows starting at offset 2" cannot be misread; "2 to 5" can. - When a boundary matters, **assert the row count** immediately after taking the range. It is one line and it converts a silent off-by-one into a failure at the point of the mistake. - Prefer the scheme that matches the intent: a date range over dated rows is a label range, and writing it as offsets means recomputing it every time the table changes. - When porting or reviewing, read a range with the convention in mind rather than the numbers: ask which scheme resolves it, whether the endpoint is in, and whether a negative number addresses or excludes. An interviewer asking this is not testing arithmetic. They are checking whether you know that the endpoint rule is a property of the **scheme**, not of the tool, and that the two schemes differ inside the same tool you use every day.

  • Why is the half-open convention preferred for offsets in the first place?
    Because it makes the arithmetic trivial: the number of rows is the end minus the start, an empty range is a start equal to its end, and consecutive ranges tile a table with no overlap and no gap. Closed ranges force a plus-one or a minus-one at every one of those points, and that is where off-by-ones are born.
  • What is the cheapest check that catches an endpoint mistake before it reaches a report?
    Assert the row count right after the range is taken. You almost always know what it should be — a chunk size, a number of days, a stated top-n — and the assertion turns an off-by-one into an immediate failure at the line responsible, rather than a number that is quietly one row wrong three steps later.

saying these in an interview costs you the question

  • Assumes both schemes use the same endpoint rule inside one tool.
  • Says ranges are always half-open, without naming the scheme or the host.
  • Reads a negative position as counting from the end in every host.
  • Believes a label range works regardless of whether the labels are ordered.
  • Treats an endpoint mistake as harmless because the code still runs.