A team argues their data is too big for a query service and wants a cluster instead — which properties actually decide?
answer
- size is the weakest predictor
- expressibility, readability, where data sits
- the awkward input, not the large one
- libraries next to the data
- SQL is not the test either
basics
~20 sVolume is the weakest predictor. What decides is expressibility: whether the work needs arbitrary code and libraries running next to the data, whether the service can read the inputs at all, and where the data already lives.
solid answer
~50 sScanning very large tables is the ordinary daily case for a managed query service — a service that owns its own storage, plan and capacity and takes a declarative statement — so size alone rarely forces anything. Three properties do the real deciding. **Can the work be expressed** in what the service accepts, including its user-function surface? **Can the service read the inputs** — undecoded blobs, a format it has no reader for, an outside system reachable only by a network call are all disqualifying. **Where does the data already sit**, since moving it out and back is pure cost that buys nothing. If all three point inward, the honest answer is that the pipeline is a handful of declarative statements. Note that writing SQL is not itself the test: cluster execution engines accept it too.
go deeper
Remember the order of the questions: can it be expressed, can the inputs be read, where does the data already live. Size comes last and usually changes only the bill.
Explain why a large scan is the ordinary case for a service whose capacity is not yours to exhaust, and name the properties that genuinely disqualify it: a library that must run next to the data, an input with no reader, a per-record call outside.
Apply the tests to a concrete pipeline and split it at the boundary rather than moving the whole thing. Be able to say what the sandboxed function surface of a service does and does not cover, and check rather than assume.
Watch for the volume argument being used to justify a runtime the organisation wanted anyway. Insist the team names the awkward input; if they cannot, the cluster is buying capability nobody needs.
## Why volume is the wrong first question "It's too big for a query service" is the most common wrong answer in this whole subject area, and it is wrong for a structural reason: **scanning very large tables is what a managed query service is for.** A service that owns its own storage, its own plan and its own capacity is not a machine you can run out of in the way a process on a laptop runs out of memory. Its capacity is not yours to exhaust; it is the service's to supply, and the visible consequence of a large scan is a larger bill and a longer wait, not a failure. That does not mean volume is irrelevant. It means volume affects **what the work costs**, not **which arrangement can do it** — and this question is about which arrangement can do it. ## The three properties that actually decide ### 1. Expressibility Can the transformation be stated as a description of a result rather than a procedure? Most aggregation, filtering, joining and reshaping can. What resists it: - **arbitrary code** with control flow that depends on values the statement cannot see; - **third-party libraries** — a decoder, a scientific package, a machine-learning model, a company's own parsing code — that must run next to the data; - **iterative work** that revisits the same working set many times with a termination condition of its own. The qualification that matters: **many query services run user-supplied functions** in a sandboxed language runtime. What such a function may load, how long it may run and whether it may reach the network vary enormously from service to service. So "it needs code" is never a yes-or-no test. The test is whether *this* code, with *these* libraries and *these* resource needs, can run *there*. ### 2. Readability of the inputs A service can only process what it has a reader for. Genuinely disqualifying inputs include: - an opaque binary payload that needs a vendor library before there are rows at all; - scanned documents, audio, images — anything where the first step is decoding rather than selecting; - a format nobody has written a reader for, or one whose schema must be resolved by custom code; - an outside system that must be *called*, per record or per group, rather than read. Many services can read files sitting outside their own storage, so "the data is not in the service" is not automatically disqualifying either — the question is always whether it can parse what it finds there. ### 3. Where the data already lives This one is not a capability question, it is an arithmetic one. If both inputs already sit inside the service, a cluster program has to pull them out, compute and push a result back. That movement is pure cost: it buys nothing, it adds a failure mode, and it usually dominates the runtime of the transformation it was added to serve. Conversely, if the inputs are files the service cannot read and the outputs go somewhere it cannot write, the pull is the other way. ## The trap in the middle: the language is not the test A candidate who concludes "it's SQL, so it belongs in the service" has swapped one wrong test for another. Cluster execution engines — systems you hand a whole program to, which split its work across many machines — almost all accept a declarative surface as well. Writing a declarative statement tells you that the work is *expressible*, which is genuine evidence for property 1, but it tells you nothing about who should own the storage, the plan and the capacity. ## Comparing the tests | Proposed test | Verdict | Why | |---|---|---| | Input size in terabytes | Weak | Changes the bill and the wait, rarely the feasibility | | Number of joins | Weak | Joining is core competence on both sides of the line | | Needs a specific library next to the data | **Strong** | A runtime that cannot load it simply cannot do the work | | Input needs decoding before rows exist | **Strong** | No reader, no statement | | Must call an outside system per record | **Strong** | Rarely expressible, and where it is, rarely affordable | | Data already lives inside the service | **Strong, inward** | Moving it out and back buys nothing | | The team prefers writing programs | None | A staffing preference, not a property of the workload | ## How to answer this in an interview Say the weak test out loud and dismiss it, then name the strong ones and apply them to the specific workload you were handed. The line that lands is: **it is the awkward input, not the large input, that forces a cluster** — and if the team cannot name an awkward input, the pipeline they want a cluster for is very likely three declarative statements.
- One step of a pipeline needs a vendor library and the rest is plain aggregation. Must the whole thing move to a cluster?No, and assuming so is the usual overreaction. Split at the boundary: run the decoding step where the library can load, land its output in a form the service can read, and let the aggregation stay declarative. The cost of the split is an extra materialised hand-off and one more thing to schedule, which is usually far cheaper than moving the whole pipeline.
- The service can run a user-supplied function. Does that dissolve the code test entirely?It narrows it rather than dissolving it. The question becomes what the sandbox permits: which language runtimes, which packages, how much memory, how long a single call may run, whether the function may reach the network. Those limits differ sharply between services, so the answer has to be checked against the actual dependency rather than assumed.
- If both arrangements can do the work, what decides?Where the data already lives, who will operate it, and which bill you would rather defend. When inputs and outputs are both inside the service, the declarative version wins almost automatically because it removes a movement, a runtime and a set of machines someone has to own.
Hiring a van and driving the route yourself, against handing the parcel to a carrier. The carrier owns the depot, the route and the fleet, and it will happily take something enormous — what it will refuse is the awkward cargo it has no way to handle. It is the awkward load, not the heavy one, that puts you behind the wheel.
saying these in an interview costs you the question
- Picks the arrangement from the number of terabytes alone.
- Says a query service cannot handle large joins.
- Assumes every input a program can parse is one the service can read.
- Thinks writing SQL settles which system should run the work.
- Believes a service that runs user functions can run any library.