A single-machine data job is called slow with no further detail; which measurements name the resource that ran out and the step?
answer
- slow is a symptom, not a resource
- per step, never per run
- elapsed against summed processor time
- one core busy, others idle
- kernel time means waiting in disguise
basics
~20 sThree numbers per step, not per run: elapsed time, processor time summed across cores, and peak resident bytes. The gap between elapsed and processor time separates waiting from computing; one busy core beside idle ones separates a design's default from a bottleneck.
solid answer
~50 s'Slow' is not a resource, so the first job is to turn one duration into a sequence. Record elapsed time per step, processor time summed across all cores per step, and peak resident bytes per step. Then read them against each other. Elapsed time far above summed processor time means the run was waiting — on storage reads, or on pages moving to and from disk — and faster cores will do nothing. Summed processor time above elapsed time means several cores were genuinely working. One core pinned while the rest idle means that step ran on one core, which may be the design's default rather than a defect. Time spent in the kernel separates work from the machinery around the work. The deliverable is one sentence: which step, which resource, and what the number was.
go deeper
Recall that a program can be slow because it is computing or because it is waiting, and that those are different problems with different fixes. Timing each step separately is what tells them apart.
Explain the ratio between elapsed time and processor time summed across cores, and why a peak close to the machine's memory changes how you read a long elapsed time.
Demonstrate the discipline of instrumenting step boundaries before changing anything, and of naming the exhausted resource and the step in one checkable sentence rather than reporting a duration.
Treat per-step measurement as standing infrastructure rather than a one-off investigation: without it every later discussion about hardware or engineering effort is conducted between two unmeasured opinions.
## 'Slow' names nothing A job described as slow has been described by its symptom, not by its cause. The whole purpose of measurement here is to replace that word with a sentence of the form *step four spent fifty minutes of its fifty-two waiting on storage*, because that sentence is checkable and 'slow' is not. Everything below is in service of producing it. The first move is always the same: **stop measuring the run and start measuring the steps.** A single total for a run hides which part of it was expensive, and the expensive part is usually one step out of a dozen. ## The three numbers, and what their ratios mean For each step, three figures are enough to get a long way: - **Elapsed time** — how long the step took on the wall clock. - **Processor time, summed across all cores** — how much actual computing the step caused, regardless of how many cores did it. - **Peak resident bytes during the step** — the worst moment, not the average. The diagnosis is in the ratios, not in any single figure: | what you see | reading | what it rules out | |---|---|---| | Elapsed time far above summed processor time | the step spent its life waiting — storage reads, decompression stalls, or pages moving to and from disk | the step being limited by how fast the cores compute | | Summed processor time above elapsed time | several cores were genuinely busy at once | the step being single-threaded | | Summed processor time roughly equal to elapsed time, one core pinned | one core did the work end to end | more cores, by themselves, helping this step | | High elapsed time together with a peak close to the machine's memory | the machine was probably moving pages to keep the job alive | a pure compute limit | ## Separating the work from the machinery around it Processor time splits into time spent in your work and time spent in the kernel on that work's behalf. A step showing a core at full utilisation is not automatically compute-bound: if most of that utilisation is kernel time, the core is busy servicing page faults or copying bytes in and out of storage, which is a memory or storage story wearing a processor costume. Reading the split is what stops the diagnosis 'the core is saturated, buy faster cores' from being made on a machine that is actually short of memory. ## What varies, and why a saturated core proves less than people think This is where an honest answer has to say what differs rather than assert one model: - **Default parallelism is not a property of the subject.** Some designs run a single whole-column operation across every available core without being asked; others run one core unless the author opts in; and within one design some operations parallelise and others do not. So one busy core beside seven idle ones is a *fact to explain*, not automatically a defect. - **How many times the input is traversed differs too.** The same expression can be executed as one traversal on one design and as several on another, which changes elapsed time without changing anything about the logic. Comparing timings across two designs without knowing that is comparing two different amounts of work. - **Where the waiting happens differs by input.** Waiting on a compressed input can be decompression, which is processor time, while waiting on an uncompressed one is usually storage. Both look like 'the read is slow' from outside. ## Turning the numbers into the sentence A workable order, and it is cheap: 1. **Instrument the boundaries between steps** so that the run reports per-step elapsed time. Nothing more sophisticated is needed to find which step owns the duration. 2. **Add summed processor time and peak resident bytes** to the same per-step record. 3. **Take the one or two steps that own most of the elapsed time** and read the ratios above. The rest of the run is noise until those are understood. 4. **Write the sentence.** Which step, which resource — a core, main memory, this machine's storage, or waiting on something outside the process — and the number that says so. ## What the sentence is worth The measurement is valuable precisely because it is narrow. It does not say what to do; it says what happened, in terms nobody can argue with. A step that spent ninety-five per cent of its time waiting on storage and a step that pinned one core for an hour are two completely different situations that both arrive at the desk as 'the job is slow', and every decision that follows — of any kind, in any direction — is made badly when those two have not been told apart.
- Summed processor time comes out higher than elapsed time. Is the measurement wrong?No — that is the normal signature of several cores working at once. Four cores busy for ten minutes produce forty minutes of processor time in a ten-minute step. The ratio is a rough count of how much parallelism the step actually achieved, and a ratio near one on a many-core machine is the interesting case, not this one.
- A step shows a core at full utilisation, yet more cores do not make it faster. What else should you look at?Split that utilisation into work and kernel time. A core busy servicing page faults or copying bytes from storage looks identical from outside to a core doing arithmetic. Check the peak resident bytes for the same step: if it sits near the machine's memory, the utilisation is the cost of keeping the job alive, not the cost of the computation.
saying these in an interview costs you the question
- A core at full utilisation proves the step is compute-bound
- The job is slow, so the whole job needs reworking
- Every whole-column operation uses all cores by default
- One total duration is enough to identify the bottleneck
- Processor time can never be less than elapsed time
- Timing the run twice from the outside is measurement enough