One standing pool runs sixty pipelines for eight teams and the monthly invoice names only the pool. What must be arranged at submission time?
answer
- the invoice names only the pool
- labels attached at submission time
- meter per run, not per month
- a remainder no run caused
- no tag, no attribution, ever
basics
~20 sA run tag — a label carrying owner, pipeline and a unique run identifier, attached when the work is submitted — plus per-run usage captured while the run still exists. Without both, an untagged run leaves a remainder nobody can claim.
solid answer
~50 sA pool's invoice is one number for one pool, so splitting it needs two things arranged in advance. First, a **run tag**: a label attached at submission carrying the owning team, the pipeline and a unique identifier for this submission. Second, per-run usage metered while the run exists — capacity held times duration, input bytes read, output bytes written — captured as numbers that outlive the workers. Then the pool's bill is split in proportion to one of those usage measures. Two parts resist this. Capacity the pool held while no run was using it was caused by no run, so a usage-proportional rule leaves it over and a policy has to absorb it deliberately. And a continuous job never ends, so its unit of attribution is a slice of time rather than a run. Nothing here retrofits: an untagged submission is unattributable forever.
go deeper
Recall that a shared cluster's bill arrives as one number and that splitting it needs a label attached when work is submitted. Nothing about who ran what can be recovered after the machines are gone.
Explain the two ingredients — a label carrying owner, pipeline and submission identity, and usage metered per run — and why the split should use whichever measure the pool itself is billed on.
Show that you arrange this before it is needed: enforced tags, usage captured outside the run, and a stated policy for the share no run caused. Mention that a continuous job is attributed by time slice, not by run.
Argue who should bear the remainder and what the organisation is buying with a shared pool at all. Attribution is the mechanism that puts a cost in front of the team that can remove it; without it, every run is free at the point of use.
## What the invoice actually says A shared pool arrives as a single line: this pool, this period, this amount. Nothing in it names a pipeline, a team or a submission. Everything you will ever be able to say about who caused that amount has to have been arranged before the runs happened, because the machines that did the work no longer exist and their own record of themselves went with them. Attribution therefore has exactly two ingredients. ## Ingredient one: a label attached at submission A **run tag** is a label attached when work is submitted so the bill can later be split. What it should carry: 1. **The owner** — the team or cost centre that answers for the spend. This is the question the finance conversation asks. 2. **The pipeline or job identity** — stable across runs, so the same work can be compared week over week. 3. **A unique identifier for this submission** — so a single bad run can be isolated rather than averaged into its pipeline's monthly total. 4. **A version or deployment marker** — so a change in cost can be lined up against the change that caused it. The third and fourth are the ones teams skip and then miss. A team label answers *who pays*; only a per-submission label answers *which change made it expensive*, and the second is the one that leads to a fix. The label must also be *enforced*, not offered. If a submission without a tag is accepted, the untagged share grows quietly, and it grows fastest in exactly the places nobody is watching — ad-hoc runs, one-off investigations, a new pipeline someone stood up last week. ## Ingredient two: usage metered per run A label with nothing attached to it splits nothing. For each run you need at least: - **capacity held times duration** — worker processes, each being one process on one machine that runs pieces of the job, multiplied by how long the run held them; - **input bytes read**; - **output bytes written**, and where it was written; - **the run's start and end**, so a period can be reconstructed. Engines differ in which of these they publish and at what granularity — some expose run-level totals, others only per-step figures you have to sum — so treat the collection as something you build once and reuse, and store the results outside the run. Numbers read off a live screen are gone when the run is. ## Choosing the rule that splits the bill | Allocation rule | Splits by | What it rewards | Where it distorts | |---|---|---|---| | Held machine-seconds | capacity times duration | finishing sooner, holding less | a run that reads enormous input quickly looks cheap | | Input bytes read | volume touched | reading narrowly | a slow, idle-heavy run looks cheap | | Equal per run | count of submissions | nothing useful | one enormous run costs what a trivial one does | | Units of work completed | pieces finished | nothing useful | uneven piece sizes make it meaningless | The honest choice is to split by whichever measure the pool itself is billed on, so the internal split has the same shape as the external invoice. Where the pool is billed on machine-time, split on held machine-seconds; where it is billed on volume read, split on bytes. Mixing them produces internal numbers nobody can reconcile with the bill they came from. ## The remainder nobody caused A usage-proportional rule never sums to the invoice. The pool holds capacity while nothing is running, keeps headroom for spikes, and pays for the coordinating process that hands out work and tracks what finished. None of that was caused by any particular run. What the pool costs simply for existing is a separate question about keeping the pool at all; the decision this leaves you is only which policy absorbs the difference. The usual options are to spread it across consumers in proportion to their attributed share, or to leave it on the platform owner as the price of running a shared service. Either is defensible; leaving it silently on whichever team happened to be largest is not, because it makes their numbers untrustworthy and they will stop acting on them. ## Jobs that never end A continuous job has no finished run to price. The tag attaches to a long-lived job rather than to a submission, and the unit of attribution becomes a slice of wall-clock time: what this job held, and read, between these two instants. That also changes what a cost review looks like — instead of comparing runs, you compare periods, and a step change in a period is the signal that something was deployed. ## Why this is a senior question Attribution is not accounting hygiene; it is the only mechanism that makes waste visible to the person who can remove it. An unattributed pool is a pool where every individual run is free at the point of use, and the predictable result is the run that costs ten times its equal, running every night for three years with nobody able to name it.
- How do you attribute a job that runs continuously for months and never finishes?By time slice rather than by run. The tag attaches to the long-lived job, and you record what it held and read between two instants, then compare period against period. Add a deployment marker so a step change in a period can be lined up with the change that caused it, since there is no run boundary to blame.
- What part of a shared pool's bill can honestly be blamed on nobody, and what do you do with it?Capacity held while nothing was running, headroom kept for spikes, and the coordinating process that hands out work. A usage-proportional split leaves it over. Pick a policy and publish it: spread it pro-rata across attributed shares, or carry it on the platform owner. What fails is letting it land silently on the largest consumer.
- Why is a per-submission label worth more than a per-team one?A team label answers who pays; a per-submission label answers which run made the month expensive. Without it, one pathological run is averaged into a pipeline's monthly figure and disappears. The identifier is what lets a cost review point at a single run and at the change that shipped just before it.
saying these in an interview costs you the question
- Thinks ownership can be reconstructed later with nothing attached at submission
- Splits a shared bill equally per team regardless of what each consumed
- Assumes every unit of a pool's bill was caused by some run
- Tags the pipeline but never the individual submission, so one bad run hides
- Splits internally on a measure the pool is not actually billed on
- Produces an attribution report nobody is expected to act on