A stopwatch around a job's declared steps reports four milliseconds while the run takes forty minutes — what did it measure?
answer
- a number that describes nothing
- the timed region read no records
- it measured describing, not doing
- move the clock around the demand
- a running job has no elapsed time
basics
~10 sIt measured graph construction: recording the steps, and perhaps resolving names and types. No records were read inside the timed region, so the number describes the program's description of the work, not the work.
solid answer
~50 sIn an engine that records declared steps and runs none of them until a demand arrives, everything inside the timed region was bookkeeping — adding steps to the graph, resolving column names and types against a known schema, maybe listing inputs. The forty minutes belong to the call that asked for an answer or wrote the output, because that is the call that submits and runs the graph. Move the clock around that call and you get an honest wall time, but it is a single number covering planning, scheduling, reading, every point where records cross the network and the final output commit, all attributed to one line. Per-step attribution has to come from the engine's own reporting about the run, which is a separate subject. And a job that runs continuously has no elapsed time at all to measure.
go deeper
Know that the timed region only built a description of the work. The number is not a runtime, and the run happens at the call that demands an answer or writes output.
Explain what the demand's wall time actually contains — planning, waiting for capacity, reading, records crossing the network, committing output — and why a single figure on one line cannot tell you which step to change.
Design a measurement someone else could trust: repeat runs, state what was retained between them, separate first-run fixed costs, and refuse to compare a sample-style demand against a full write.
Decide what the organisation measures at all. For finite jobs a cost-per-run figure with attribution is the useful unit; for continuously running jobs there is no duration, and the contract has to be expressed as sustained throughput and acceptable lag instead.
## What the four milliseconds covered The timed region contained **assembly calls** — calls that only add a step to the graph and compute nothing. Between the start and stop of that clock the engine built the **step graph**: the ordered set of steps it derives from the program, each step naming the steps whose output it reads. Depending on the engine it may also have resolved column names and types against a schema it already knows, consulted a catalog, and in some designs listed the files it will read. None of that touches a record. So the measurement is real, it is just a measurement of a different thing: *how long it takes to describe the job*. Four milliseconds is a perfectly plausible answer to that question and tells you nothing whatsoever about the forty minutes. ## Where the clock belongs Move it around the **demanding call** — the call that asks for an answer or for the output to be written, and so is what actually makes the assembled graph run. That gives a defensible wall time for the run. Be honest about what that single number now contains: - the time the engine spends rewriting the declared graph and choosing how to run it; - waiting for capacity, if the cluster is shared or the workers have to be started; - reading the input, which for a wide scan is often most of it; - every point where records must be sent between workers because the next step needs records another worker holds; - the final commit of the output, which on some storage systems is a rename and on others is a slow copy. All of it is attributed to one line of your program, and that line is usually the least interesting one. So the second measurement answers *how long did this run take* and still does not answer *what cost the time*. Which step cost what is read from the engine's own reporting about the running job, and that is a separate subject from this one. ## Timings that are not comparable A benchmark on this kind of engine goes wrong in a few reliable ways: - **Timing assembly instead of execution** — the case in the question, and the one that produces suspiciously wonderful numbers. - **Timing a demand that does less than you think.** A demand for a handful of sample rows may let the engine stop early once it has enough, so it never touches most of the input. A row count may be answerable from metadata the storage format already carries. Neither is a measurement of the full job. - **Comparing two demands over the same graph without knowing what was retained between them.** Whether a computed result is being held for reuse changes the second number completely; pinning a result for reuse is its own subject, but a benchmark that does not state whether it happened is not reproducible. - **Ignoring the first run's fixed costs.** Starting workers, warming storage-layer caches and compiling generated code land on the first run and distort a single-shot comparison. - **Treating one run as a measurement at all.** Two runs of the same job over the same input can differ substantially for reasons that have nothing to do with your change. ## Jobs with no elapsed time The question assumes a job that ends. Two common models do not: 1. **One fixed graph kept running**, with each record passing through as it arrives. The call that submits it returns almost immediately, and the job then runs until cancelled. "How long did it take" has no answer; the health of such a job is described by sustained throughput and by how far behind its input it is running. 2. **Continuous work run as a fast succession of small finite jobs.** The call that starts it likewise returns a handle rather than a duration. Each small job has a duration, but the interesting figure is again a rate, not a single elapsed time. A candidate who reaches for a stopwatch on either of those has not noticed which kind of job they are holding. And at the other extreme, a model that runs one grouping step at a time and writes every intermediate to storage barely has an assembly phase to mis-measure: submitting the job and running it are effectively the same act. ## Answering it in an interview Say what was inside the timed region — steps recorded, no records read — and why that means the number is not wrong, just about something else. Then move the clock to the demand and immediately state the limitation of the number you now have: one figure, one line, no attribution. Finish with the caveat that a continuously running job has no duration to measure, which shows you know the split between a finite job and a running one.
- If you time the demand instead, what is that number still hiding?Everything except the total. It bundles rewriting the graph, waiting for capacity, reading the input, every point where records cross the network, and the output commit into one figure attributed to a single line. It is a fine regression signal for the whole job and useless for deciding which step to change; that attribution comes from the engine's own reporting about the run.
- Why can a demand for a small sample return far faster than the same graph written to storage?Because a sample can often be satisfied without producing everything: the engine may stop as soon as it has enough records, so most of the input is never touched. A write requires every output record to exist. Benchmarking with a sample-style demand and reporting it as the job's cost is a common and large error.
- Two runs of the same job over the same input differ by twenty per cent. Is the measurement broken?Not necessarily. Cluster sharing, worker startup, storage-layer caching and the ordinary variance of work spread over many machines all move the number. Treat a single run as a sample, repeat, report a spread rather than a point, and state what was retained between runs.
saying these in an interview costs you the question
- Reports assembly time as the job's runtime
- Concludes a step is cheap because declaring it was cheap
- Attributes the demand's whole wall time to the last declared step
- Expects the call that starts a continuous job to return when it finishes
- Compares two runs without stating what was retained between them
- Benchmarks with a sample-style demand and calls it the full cost