skip to content

A nightly cluster program pulls two tables out of a query service, joins and aggregates them and writes a summary back — what do you check before replacing it with three declarative statements?

level: seniorimportance: should knowfreq 52%

answer

  1. the data never needed to leave
  2. export and import buy nothing
  3. audit what the code carries silently
  4. null rules, deduplication, an outside lookup
  5. you trade debuggable code for no round trip

basics

~20 s

Check what the program's own code carries beyond the join and the aggregate — null rules, deduplication, an outside lookup, precision handling — and whether anything needs a library. If nothing does, the extraction and return trip are pure cost.

solid answer

~50 s

Start with the shape: both inputs already live inside the service, so the program's first act is to move data out and its last is to move it back, and that movement buys nothing. Then audit what the code does that the three statements would not. The join and the aggregate transfer cleanly; what hides in a program is everything else — custom null and empty-string handling, a deduplication rule, rounding and precision decisions, a lookup against an outside system, an ordering assumption. Each of those must be either re-expressed or consciously dropped. Check next whether any step needs a library the service cannot load, and whether the destination is somewhere the service can write. What also changes is operability: you trade code you can unit-test and step through for a statement whose execution you do not see.

go deeper

for a junior

Notice the round trip first: if both inputs already live in the service, the program is exporting and re-importing data for no gain. That observation alone is worth saying out loud.

for a middle

Explain what a rewrite transfers cleanly and what it does not. Joins and aggregates move; null handling, deduplication rules, rounding and outside lookups are carried by code and must be re-expressed deliberately.

for a senior

Run the audit as a list and price the trade honestly: you lose locally executable, debuggable logic and gain the deletion of a runtime, a set of machines and a round trip. Name the short list of cases where keeping the program is right.

for a principal

Note that the two directions are not symmetric. Moving into the service deletes code; moving out duplicates a business rule and adds an on-call surface, so the escape hatch deserves a higher bar than the default.

## Read the shape before reading the code The described program has a tell. Both inputs already live inside a **managed query service** — a service that owns its own storage, its own plan and its own capacity and accepts a declarative statement rather than a program. So the program's first act is an export and its last act is an import, with a join and an aggregate in between. The movement is **pure overhead**: it buys no capability, it adds two failure modes, and on most nightly pipelines of this shape it dominates the wall-clock time of the transformation it was added to serve. That observation is the beginning of the answer, not the end of it. Programs accumulate behaviour that nobody wrote down, and the rewrite is where that behaviour is silently lost. ## The audit, in order 1. **Does anything need a library?** A decoder, a model, a company parsing package, anything that must run next to the data. If one step does, it does not follow that the whole pipeline must stay — split at that step. If nothing does, the strongest argument for the cluster has just evaporated. 2. **Can the service read every input and write the destination?** Both inputs are already inside it here, but the *output* may not be: if the summary lands in a system the service cannot write, you need either a small exporter or a reason to keep the program. 3. **What does the code do besides join and aggregate?** This is where rewrites go wrong. Walk the source looking for: null and empty-string handling that differs from the service's default semantics; a deduplication rule; a rounding, precision or type-widening decision; a filter applied for data-quality reasons and never documented; a per-record or per-group call to an outside system; an assumption that rows arrive in a particular order. 4. **Is anything downstream depending on an artefact rather than the answer?** Programs often leave behind intermediate files that somebody has quietly started reading. The rewrite deletes them. 5. **What do the two bills actually say?** They are denominated differently — machines that exist against work a statement consumed — so this is a conversion, not a comparison. Do the conversion before claiming a saving. ## What genuinely changes when you cross the boundary | Concern | As a cluster program | As statements in the service | |---|---|---| | **Who owns the plan** | Your program fixes the shape; how much the engine rewrites varies by surface and by engine | The service's planner owns it entirely; you influence it only by changing the statement | | **Testing** | Unit tests over your own functions, run anywhere | Assertions over results, run against the service or a copy of it | | **Failure diagnosis** | Your stack trace, your logs, machines you can identify | Whatever the service publishes about the statement's execution | | **Dependencies** | Anything you can package | Only what the service's function surface permits, and that varies widely between services | | **Data movement** | Export, compute, import | None; the data never leaves | | **Who can change it** | Someone comfortable with a distributed runtime | Someone comfortable with the data | The first two rows are the real trade. You are giving up code you can execute on your own machine in exchange for removing a runtime, a set of machines and a round trip. For a pipeline that is honestly a join and an aggregate, that trade is usually correct — and "usually" is doing real work in that sentence, because rows three and four are where it goes wrong. ## When the answer is to keep the program There is a short list, and it is worth saying explicitly so the rewrite does not become dogma: - a step needs a library or a resource shape the service cannot provide; - the transformation calls an outside system in a way that is not expressible, or is expressible and ruinous; - the destination is outside the service's reach and the export would be as large as the input; - the logic is genuinely iterative with its own termination condition. ## Moving the other way The same audit run backwards is what makes this a boundary question rather than a migration recipe. When a transformation has to *leave* the service — because it has grown a dependency the service cannot load — you pay for the extraction you just removed, you inherit responsibility for machines, you need somebody who can operate a distributed runtime on call, and you acquire a second place where a business rule lives. The direction of travel is not symmetric in effort: moving *in* usually deletes code, and moving *out* usually duplicates it. ## How to say it A good answer opens with the shape — "the data never needed to leave" — then refuses to stop there, and names the behaviour hiding in the program as the thing that will actually break. The weak answer is either "rewrite it, obviously" or "never touch a working pipeline"; both skip the audit that decides which is right.

  • The rewrite is done and the numbers differ slightly from the old program. Where do you look first?
    At the semantics the program carried implicitly: null and empty-string handling, deduplication, rounding and type widening, and any quality filter applied in code. Small, systematic differences almost always come from one of those rather than from the join or the aggregate, which transfer cleanly. Reconcile on a single partition of input before re-running the whole thing.
  • What do you lose in testing when the logic becomes three statements?
    The ability to execute the logic in your own process against handcrafted inputs. You can still assert on results, but the assertions now need the service or a faithful copy of it, which makes the feedback loop slower and the test data harder to hold. Budget for that rather than discovering it after the cutover.
  • A single step needs a library. Does that end the discussion?
    No — it relocates the boundary rather than settling it. Run just that step where the library can load, land its output in a form the service can read, and keep the join and the aggregate declarative. You then own one small program instead of a whole pipeline, which is a materially different maintenance bill.
  • How should the two costs be compared?
    By converting, not by reading them side by side. One arrangement bills for machines that exist, whether or not anything is running; the other bills for the work a statement consumed. Turn both into cost per nightly run over a representative month, and include the movement the program was paying for.

saying these in an interview costs you the question

  • Rewrites the join and the aggregate and never audits the rest of the code.
  • Assumes the extraction step is required because the program has always had it.
  • Compares the two bills directly without converting the units.
  • Says the rewrite is free because declarative statements need no tests.
  • Keeps the program solely because the team already knows the runtime.
  • Claims moving a transformation out of the service costs the same as moving it in.