One input arrives as files on shared storage, another over a continuous feed. Why does the transport not decide whether the input is bounded?
answer
- a property of the record set
- same source, two readings
- a pinned slice has an end
- a watched location never ends
basics
~20 sBoundedness is whether the record set has a last member, not how the records travel. A watched location that keeps filling is unbounded, and a slice of a never-ending feed pinned between two chosen boundaries is bounded.
solid answer
~50 sThe same source supports both readings, so the transport cannot settle it. A location on shared storage — storage every machine in the cluster can read — is bounded when a producer has finished writing and the file list is fixed, and unbounded when the job keeps watching it for new arrivals. A continuous feed is unbounded when read from a start position onwards with no stopping boundary, and bounded when the computation is defined over a slice pinned between a start and an end. What decides it is the *definition of the record set the computation covers*: is something still adding members to it? Boundedness also does not decide how the job runs — a finite input can be handled a record at a time, and an endless one can be served by repeated small finite runs, so all four combinations exist.
go deeper
Remember that the same place records come from can supply both a finite set and an endless one. Ask whether anything is still adding records before you call an input finite.
Explain the mapping both ways with an example each: a watched location that never ends, and a feed slice pinned between two boundaries that does. Then say why a clock-based cut is not the same thing.
Demonstrate that you make membership explicit in the design — a declared completion or a fixed pair of boundaries — rather than inferring it, and say what goes wrong in reconciliation when two runs covered different sets.
Set the expectation that every producer on the platform declares completeness in a uniform way. Inferring it from arrival timing is a per-team guess that surfaces as unexplainable disagreements between numbers.
## Boundedness belongs to the record set, not to the pipe The question "is this input bounded?" is really "is anything still adding records to the set my computation is defined over?" That set is defined by you and the source together. The source supplies records; your definition says which of them count. Transport supplies neither. This matters because the shorthand most teams use — files mean finite, feeds mean endless — is right often enough to survive for months and then wrong exactly when it is expensive. ## The same source, read two ways | Source | How the computation is defined over it | Bounded? | |---|---|---| | A location on shared storage the cluster reads | over a fixed list of files a producer has declared complete | yes — membership is settled before the read starts | | The same location | watched for as long as the job lives, picking up whatever appears | no — the producer keeps adding members | | A continuous feed | from a chosen start boundary up to a chosen end boundary | yes — both ends are fixed, and a rerun covers the same records | | The same feed | from a start boundary onwards, with no end | no — there is no last record to reach | | A table in an operational database | one export taken at a single moment | yes — that snapshot's membership cannot change | | The same table | followed continuously as it changes | no — and how such change feeds are produced is a different subject | The row that surprises people is the third: a feed with retention long enough to hold the slice is a perfectly ordinary bounded input, and pinning both ends is what makes a rerun cover the same records rather than a longer set. ## Why the shorthand is expensive when it fails - A job written as a one-shot run reads a location that is still filling. Two runs disagree, both look plausible, and nobody can say which was right, because neither covered a settled set. - A team promises a total, a global rank or a reconciled figure over something they have called finite because it arrives as files — and it never was. - A team believes it cannot re-derive a number from a feed, because "feeds are endless", when pinning two boundaries would have given them an ordinary finite computation. - A team infers the execution style from the transport and buys a system to match a property the input does not have. ## Four questions that settle it 1. Is a producer still adding records that my definition of the input includes? 2. Does anything in the data or its metadata *declare* completeness, or am I guessing from the clock? 3. Can the read reach an end without a rule I invented — and if I invented one, can I state it precisely enough that a rerun covers exactly the same records? 4. If I ran the same definition again next month, would it cover more records than it does today? A "yes" to the first or the fourth means unbounded, whatever the bytes arrive on. ## What boundedness does not settle Having an end, or not having one, says nothing about how the work is executed. Both execution styles apply to both kinds of input: - **Record-at-a-time processing** — each record is handled the moment it arrives, so nothing but the job's own retained memory carries context from one record to the next — can be pointed at a finite export and will simply run out of records. - **Repeated small finite runs** — the runtime collects whatever arrived during a fixed interval and then executes an ordinary finite job over just that slice, over and over — is a common way to serve an input that has no end at all. So all four combinations exist, and the sentence "it's a stream, so we need a streaming system" conflates a property of the data with a choice about the runtime. Which runtime to choose, and what each costs in latency floor and unit of failure, is its own subject; the point here is only that the input's ending does not decide it. One last distinction worth holding: a read that stops after a fixed duration is not the same as a bounded input. It produces a finite set of records, but the boundary is the clock rather than anything the data fixes, so a rerun covers a different set and the result is not reproducible. Pinning the boundary to something in the data — a position, a declared completion, a fixed file list — is what makes the finiteness worth anything.
- What makes a pinned slice of a continuous feed genuinely reproducible?Both boundaries have to be fixed to something in the data rather than to the clock, and the records between them must still be retained when you re-read. If the boundary is "whatever had arrived when the job started", two runs cover different sets, and if retention has already discarded the early records, the rerun cannot cover the slice at all.
- A job reads a location that is still being written. What is the cheapest fix?Make the producer declare completeness — an explicit marker, or a handover of the exact file list — and have the reader refuse to start without it. That converts a guess about membership into a fact, and it is far cheaper than reconciling two runs afterwards to work out which one saw the whole set.
saying these in an interview costs you the question
- Says files are always bounded and feeds are always endless
- Assumes an unbounded input forces one particular execution style
- Thinks a fixed slice of a continuous feed cannot be read
- Treats the transport as the answer to a boundedness question
- Calls a read that stops after ten minutes a bounded input