skip to content

In a data-access layer, what does it mean to project a query into a transfer type, tuple or scalar?

level: juniorimportance: must knowfreq 66%

answer

  1. the query decides the shape
  2. select list, not the mapper
  3. scalar, tuple, transfer type, grouped row
  4. the row never enters the tracked set

basics

~20 s

A projection asks the statement for only the fields a use case needs and materialises them as a scalar, a tuple or a small transfer type, so the row is never turned into a tracked mapped object.

solid answer

~40 s

Projecting means the query decides the shape of the result, not the mapper. Instead of selecting every mapped column and materialising one tracked object per row, the statement lists just the columns the use case needs, and the layer builds a plain value out of them: a single scalar, a positional tuple, a small named transfer type, or a grouped row of aggregates. Because the reduction happens in the statement, the row never becomes a tracked object - no snapshot, no entry in the identity map, no stand-ins for links to resolve later. What comes back is a read-only carrier of values, and that is both the point and the price.

go deeper

for a junior

Be able to say it in one line: the statement asks for the fields the screen needs, and you get back values - a number, a tuple, a small transfer type - not a mapped object you can save.

for a middle

Explain where the reduction happens: in the select list, before anything is materialised. Name the shapes a result can take and what the layer needs in order to build a named transfer type from a row.

for a senior

Show that you check the statement rather than the return type. A small transfer type wrapped around a wide read, or built by calling back into the layer per row, has fixed nothing.

for a principal

The tradeoff to own is duplication: every projected read path re-expresses rules the mapped model already carries. Decide how many shapes a team maintains and where each rule stays single.

## What a projection is A **projection** is a query whose result is the set of values one use case needs, rather than the objects a mapping layer knows how to build. The ordinary path is: the layer selects every column it has mapped for a class, builds one instance per row, and registers that instance so it can be changed and written back. A projection short-circuits that. The statement names a few columns; the layer converts each row into a plain result and hands it over. Nothing is registered, because there is nothing the layer intends to write. The defining detail is **where** the reduction happens. Shrinking a result after the fact - loading full objects and copying three fields into a smaller type in memory - is not a projection. That path still reads every mapped column off disk and over the wire, still builds the object, still pays for the bookkeeping. A projection is a narrower `SELECT` list, decided before anything is materialised. ## The shapes a projected result can take | Shape | What comes back | Typical use | |---|---|---| | **Scalar** | One value: a count, a maximum, a single column | A badge, a guard, an existence check | | **Tuple / positional row** | An ordered list of values with no type of its own | Ad hoc internal reads, quick reports | | **Named transfer type** | A small purpose-built type with named fields, built per row | Screens, list endpoints, exports | | **Grouped row** | One row per group whose columns are aggregates | Dashboards, totals, summaries | All four are projections. They differ only in how much ceremony the result carries, not in what the layer does with the row. ## What the layer has to know to build one To fill a named transfer type the layer needs a rule that maps the statement's output onto the target: usually the order of the select list against a constructor, or column aliases against field names. Layers differ in when that rule is checked - some fail only when the statement runs, others validate the mapping at startup - and in whether type conversion is automatic. A tuple or scalar needs no such rule, which is why they are the cheapest shapes to introduce and the easiest to break silently when someone reorders the select list. ## What projecting buys - **Fewer columns read** - narrow rows, less network and less memory per row. - **No tracking work** - no snapshot of loaded values, no identity-map entry, no change detection at flush, no growth of the working set as rows accumulate. - **No surprises afterwards** - no stand-in field that fires another statement when something touches it during rendering, and no chance of a stray change being written back. - **A shape the mapped model does not have** - aggregates and rows spanning two aggregates become expressible. ## What projecting costs - The result is a **snapshot of values with no write path**. To change the data you go back and load the mapped object by its key. - Behaviour that lived on the mapped model - derived values, rounding rules, a visibility predicate - is not available; the read path re-expresses it, so the rule now exists in two places. - Each screen tends to want its own shape, so the number of small types grows with the number of read use cases. - A projection is **not automatically fast**: it changes what each row becomes, not how many rows the engine must produce. ## Two traps worth naming early 1. **Projecting into a mapped class.** Selecting a few columns into the very class the mapper manages defeats the purpose. Layers differ in what they do with the half-filled instance - some register it as if it were fully loaded, so the unselected columns read as empty and a later flush can write those blanks back; others refuse the query outright. A projection target should be a type the mapper knows nothing about. 2. **Assuming absent means deferred.** A column the statement did not select is *missing*, not lazy. There is no field for it and nothing to trigger; the result cannot fill itself in later the way a link on a mapped object can. ## When not to project When the use case writes, load the mapped object - a projection cannot be flushed. When the value you need is produced by logic that lives on the model and is hard to re-express in a statement, weigh the duplication before copying that rule into a query. And when the projection ends up naming nearly every column anyway, you have paid for a second type and gained little; use the mapped path and be honest about it.

  • Is a single count or one column also a projection?
    Yes - a scalar result is the smallest projection there is: one value per row, or one value for the whole statement. It carries no tracking cost at all, and it is usually the right answer for a badge, a guard or an existence check, where loading objects in order to count them in memory is pure waste.
  • What can go wrong if the projection target is the mapped class itself?
    Layers differ, and the common outcomes are both bad. Some register the half-filled instance as though it were loaded, so unselected columns read as empty and a later flush can write those blanks back over real data; others reject the query. Point projections at a type the mapper does not manage.

saying these in an interview costs you the question

  • Thinks projecting means loading full objects and copying a few fields
  • Expects unselected columns to load later on first access
  • Believes a projected result can be edited and flushed back
  • Projects into a mapped class and expects it to stay untracked
  • Assumes a narrower result type must mean a faster statement