How do you use range's step to process a 250,000-row inventory sync in fixed batches?
answer
- The step is the batch size
- Slicing forgives an overrun; indexing does not
- Do not count the batches yourself
- Floor division loses the tail
- range(0, len(rows), size) plus rows[i:i + size]
basics
~20 sStep the start offsets: for i in range(0, len(rows), size), take rows[i:i + size]. Slicing clips at the end, so the final short batch is included with no clamp, and an empty input produces no batches at all.
solid answer
~50 sUse the range's third argument as the batch stride: `for i in range(0, len(rows), size): batch = rows[i:i + size]`. The range yields `0, size, 2*size, ...` and stops before `len(rows)`, so it produces exactly `ceil(len(rows) / size)` passes. No `min()` clamp is needed on the slice end, because list slicing silently clips out-of-range bounds — only plain indexing raises `IndexError`. Two edge cases fall out for free: an empty input gives an empty range and zero passes, and a final partial batch is produced rather than dropped, which is the bug you get from computing the count as `len(rows) // size`. Keep the offsets as a `range` object if you want to reason about them — `len(offsets)` is the batch count and `offsets[-1]` the last start, both computed arithmetically. Guard `size <= 0` yourself: a zero step raises `ValueError` and a negative one silently yields nothing.
code
python · 10 linesdef batches(rows, size):
if size <= 0:
raise ValueError("size must be positive")
return [rows[i:i + size] for i in range(0, len(rows), size)]
rows = list(range(250_000))
chunks = batches(rows, 30_000)
print(len(chunks), len(chunks[-1]), batches([], 30_000))
# 9 10000 []go deeper
Learn the shape: for i in range(0, len(rows), size) with rows[i:i + size] inside. Know that the final short batch comes out automatically and that an empty input simply produces no passes.
Explain why no clamp is needed on the slice end, and why computing the batch count with floor division drops the tail. Derive the count as ceiling division and note that len() of the stepped range already is it.
Show how you size batches against memory and the receiving system's limits, resume from a recorded offset so boundaries line up with the original run, and demonstrate that the empty-input case needs no branch.
Own the choice between an offset-stepped walk over a materialised sequence and a streaming pipeline for a sync too large to hold, and be explicit about what each costs in restartability, back-pressure and blast radius per failure.
### The idiom A recurring inventory sync between two systems has to push roughly 250,000 rows in bounded chunks, because the receiving side caps a request and because holding the whole payload in memory at once is not free. The canonical Python shape is a range whose **step** is the batch size: ```python def batches(rows, size): if size <= 0: raise ValueError("size must be positive") return [rows[i:i + size] for i in range(0, len(rows), size)] ``` `range(0, 250_000, 30_000)` yields `0, 30000, 60000, ... 240000` — nine offsets — and `rows[240_000:270_000]` returns the final 10,000 rows. That last slice asks for more than exists, and that is the point. ### Why no clamp is needed Slicing and indexing have different out-of-range behaviour, and the difference is what makes this idiom short: * `rows[i]` with `i >= len(rows)` raises `IndexError`. * `rows[i:j]` with `j > len(rows)` returns whatever exists from `i` onward, silently. So `rows[i:i + size]` never needs `min(i + size, len(rows))`. Writing the clamp is not wrong, it is just noise that suggests the author expected an exception. ### Why the count is not computed by hand The classic defect is to compute the number of batches and loop over that: ```python for b in range(len(rows) // size): # WRONG: drops the tail batch = rows[b * size:(b + 1) * size] ``` With 250,000 rows and a batch of 30,000, `250_000 // 30_000` is 8, so 10,000 rows are silently never synced — the worst possible failure shape for a data sync, because nothing raises and the row counts only diverge slowly. If you do need the count, use ceiling division (`-(-n // size)` or `math.ceil(n / size)`), but the stepped range makes the count unnecessary: `len(range(0, n, size))` already **is** the ceiling division, computed arithmetically. ### The empty input needs no branch `range(0, 0, 30_000)` is empty, so `batches([], 30_000)` returns `[]` with no special case. This is the practical payoff of range's emptiness rule: when the step cannot carry start toward stop, the range is empty rather than an error, and the loop body runs zero times. Guard clauses for "what if there are no rows" are usually redundant around this idiom. ### Reasoning about the offsets directly Keeping the range as an object rather than immediately looping over it gives you cheap answers about the plan before you execute it: ```python offsets = range(0, len(rows), size) len(offsets) # the batch count offsets[-1] # the start of the final batch offsets[3:5] # a range object -- a sub-plan, not a list 240_000 in offsets # constant-time, arithmetic ``` Indexing a range returns an `int`, and slicing one returns another `range` — so a sub-plan stays lazy. That makes it easy to log "9 batches, last starts at 240000" before any work happens, or to hand batches 3 and 4 to a different worker. ### Resuming a failed run Because `start` is an ordinary argument, resuming is a one-line change. Record the offset of the last batch that committed, and start the range there: ```python offsets = range(last_done + size, len(rows), size) ``` The stride arithmetic is identical, so batch boundaries line up with the original run — which matters when the receiving system deduplicates by batch. This is the operational argument for offset-based batching over an ad-hoc counter: the plan is reproducible from two integers. ### Sizing the batch The batch size is the only real tuning knob, and it trades three things off. Each slice is a **new list** holding references to `size` rows, so peak memory scales with the batch size, not with the total. Larger batches mean fewer round trips but a bigger unit of work to redo when one fails, and a longer transaction on the receiving side. If a sync runs on a three-week release cadence, the batch size is also the granularity at which you can change your mind mid-run — a smaller batch is a shorter commitment. ### When this idiom does not apply `range(0, len(rows), size)` requires a sized, sliceable **sequence**. A cursor, a file object or a generator has neither `__len__` nor slicing, and materialising it into a list to make the idiom work defeats the reason it was lazy. For those, `itertools.batched(source, size)` — added in 3.12 — yields tuples of up to `size` items from any iterable, with the same partial-final-batch behaviour, and `itertools.islice` covers the cases where you need to control the pulls yourself.
- What is the classic off-by-one when you compute the batch count yourself?Looping over `range(len(rows) // size)` uses floor division and drops the final partial batch — with 250,000 rows in batches of 30,000 that silently skips 10,000 rows. Nothing raises, so the two systems just diverge. Ceiling division (`-(-n // size)`) fixes the count, but stepping the range removes the need for a count at all.
- Does rows[i:i + size] need a min() clamp on the end index?No. List slicing clips out-of-range bounds silently and returns whatever exists from `i` onward, so the final short batch comes out correctly. Only plain indexing — `rows[i]` — raises IndexError past the end. Adding the clamp is harmless but signals that the author expected an exception that cannot happen.
- How would you batch a source that has no len() and cannot be sliced?Use `itertools.batched(source, size)`, added in 3.12, which yields tuples of up to `size` items from any iterable and produces a short final tuple the same way. `itertools.islice` in a loop is the equivalent when you need finer control. The stepped-range idiom needs a sized, sliceable sequence, and materialising a lazy source into a list to satisfy it defeats the point.
- How do you resume the sync from where it failed without re-sending committed batches?Record the offset of the last batch that committed and restart the range there: `range(last_done + size, len(rows), size)`. Because the stride is unchanged, the boundaries line up exactly with the original plan, which matters when the receiving system deduplicates per batch. The whole resume state is two integers.
Marking a long shelf every thirty boxes and carrying away whatever sits between two marks: the last stretch is short, and you carry it anyway without needing to have counted the boxes first.
saying these in an interview costs you the question
- Computes the count with len(rows) // size and loses the tail
- Clamps the slice end with min(), thinking slicing can overrun
- Adds a guard clause for an empty input that range already handles
- Believes each slice is a view, so batch size does not bound memory
- Passes a batch size of zero and expects no batches rather than ValueError
- Applies the stepped-range idiom to a lazy source with no length