A colleague says a cluster engine will optimise whatever program you give it - when is that true, and when is it false?
answer
- rewriting needs meaning, not code
- two conditions, not one
- opaque body leaves nothing to prove
- one model rewrites almost nothing
- surface decides whether, engine decides bothers
basics
~20 sOnly a program whose steps the engine understands can be improved. Work handed over as a function body is executed close to literally, and some engine models rewrite next to nothing on either surface, so the claim is conditional on both the surface and the engine.
solid answer
~50 sRewriting requires meaning. A **plan rewriter** - the engine component that edits the declared graph into an equivalent, cheaper graph before running it - can only act on steps it understands, which means steps written on a **declared-operator surface**, where the author names operations the engine already knows. A program written as per-record function bodies gives it nothing to reason about: it can place and connect the calls, but each body still runs once per record, where it was put. The claim is conditional on the engine too. A **two-phase disk-handoff model**, which runs one grouping step at a time and writes every intermediate to storage before the next begins, rewrites almost nothing whatever surface you used. So: true for declared steps on an engine that rewrites, false for handed-over bodies, and weak on models that barely rewrite at all.
go deeper
Remember that the engine only improves steps it understands. A function you hand it to run per record is not one of those, so it runs as written.
Explain both conditions: a step must carry a meaning for anything to be proved about it, and the engine must be one that rewrites. Name the model that writes each intermediate to storage as the case where little is rewritten either way.
Demonstrate that you check the claim against the job in hand before promising a gain, and that you would not contort readable logic into named operations on an engine that will not act on them.
The strategic version is how much of a platform's performance you are willing to leave to an engine's discretion, against how much you pin explicitly - and what happens to that bet when the engine underneath is replaced.
## What "optimise" actually requires The claim that an engine improves your program hides an assumption: that the engine can *read* the program. It cannot, in general. What it holds is a **step graph** - the ordered set of steps it derives from your program, each step naming the steps whose output it reads - and the contents of each step are whatever your authoring surface put there. - Write a step on a **declared-operator surface**, naming an operation the engine already understands (keep these rows, produce these fields, group by this key, join on that key), and the graph holds a meaning. - Write a step on a **per-record function surface**, handing over a function the engine calls once per record, and the graph holds a call. Such a step is an **opaque step**: a step whose body is ordinary code the rewriter cannot look inside, so it can only be called, never reasoned about. A **plan rewriter** - the component that edits the declared graph into an equivalent, cheaper graph before it runs - works on meanings. Given a call it cannot interpret, it has no equivalence to prove and therefore no rewrite it is entitled to make. ## Where the claim is true It is true, and genuinely valuable, for a program whose work is expressed as named operations on an engine that does rewrite. There the engine can decide the *how* while you specified only the *what*: the order in which independent named steps run, how much the read is asked to produce, and which named steps collapse into a single pass over each record. Two authors writing the same named operations in different orders can land on the same plan. That is the whole point of the declared surface and the reason it is usually the default where it exists. Note what this does **not** mean. It does not mean the engine rescues a bad choice of grouping key, a bad output shape, or a job asking for far more data than it needs. Those are the author's, and they survive every rewrite. ## Where the claim is false 1. **A program written as per-record bodies.** The engine executes it close to literally. It knows the order you wrote, and it calls each body once per record in that order. Some engines will still fuse adjacent bodies into a single pass - **step fusion**, several adjacent per-record steps collapsed so no intermediate collection exists between them - but fusing calls is not the same as changing what they do. 2. **Engine models that rewrite little.** In a **two-phase disk-handoff model** - one grouping step at a time, every intermediate written to storage before the next begins - the shape of the run is fixed by the model, not chosen by a rewriter. A declared surface there still buys checking and legibility; it does not buy a cleverer plan. 3. **Mixed jobs, above the seam.** Once a handed-over body sits in the middle of an otherwise named graph, the engine's freedom on the far side of it is reduced, because it cannot show that moving work across an operation whose meaning is unknown preserves the result. 4. **Inputs with no declared shape.** If the records arriving are raw text or bytes with no known field structure, a declared surface has little to describe and the work inevitably lands in bodies. | Program | Engine that rewrites | Engine that rewrites little | |---|---|---| | All named operations | plan may differ substantially from what you wrote | runs much as written; checking and legibility remain | | All handed-over bodies | calls may be placed and fused, meanings untouched | runs as written | | Mixed | freedom on the named parts, none through the body | runs as written | ## Why the distinction is worth arguing about Two failure modes follow from believing the slogan. The first is complacency: an author hands everything to bodies, assumes the engine will sort it out, and is surprised that the job does exactly what was written. The second is cargo-culting the opposite - rewriting expressive, correct logic into a contorted chain of named operations on an engine that was never going to rewrite it, paying in readability for nothing. The honest formulation is this: **the surface decides whether there is anything to optimise; the engine decides whether it bothers.** Both halves have to be checked before you promise anyone a speed-up. ## What an interviewer is listening for A candidate who answers "yes, that's what the engine is for" has taken the marketing line. A candidate who answers "no, engines only rewrite what they understand" is closer but still absolute. The strong answer holds both conditions at once - surface and engine model - and says which one applies to the job in front of them.
- If a body cannot be rewritten, can the engine do anything at all with it?Yes, but only structural things. It can decide where the call runs, connect it to its neighbours, and on many engines fuse it with adjacent per-record steps so no intermediate collection is built between them. All of that leaves the body's own behaviour, and its position in the order you wrote, exactly as it was.
- Does the same argument apply to how the program was spelt - a query string against chained calls?No, and that is a common false lead. Those are two spellings of the same declared surface; both are turned into named operations and reach the same rewriter, so the choice between them is about authoring taste and tooling. The axis that matters is whether a step carries a meaning at all.
saying these in an interview costs you the question
- Says the engine understands and improves any program, on whatever surface
- Assumes every engine of this class has a plan rewriter worth the name
- Thinks a rewriter can fix a badly chosen grouping key or output shape
- Believes fusing adjacent calls means their behaviour was optimised
- Argues the real difference is query text against chained method calls