skip to content

In a data-access layer, what is an always-on filter on a mapped type, and which reads does it affect?

level: juniorimportance: must knowfreq 62%

answer

  1. declared once, applied by the layer
  2. predicate on the type, not the query
  3. live rows and current tenant
  4. only on statements the layer generates
  5. value from ambient context per unit of work

basics

~20 s

An always-on filter is a predicate declared once on a mapped type, such as live-rows-only or current-tenant-only, that the layer adds to the WHERE clause of the statements it generates for that type, so no caller has to restate it.

solid answer

~50 s

Most mapping layers let you attach a condition to a mapped type itself rather than to each query: something equivalent to `deleted_at IS NULL` or `tenant_id = :currentTenant`. From then on, every statement the layer *generates* for that type carries the condition — a query by criteria, a query over a collection-valued association, a lazily loaded reference. The parameter usually comes from ambient context (the current tenant resolved at the request boundary), not from the call site. The point is that correctness stops depending on every developer remembering the clause. The important caveat is the word *generates*: paths where the layer does not build the SQL, or does not go to the database at all, are not filtered — hand-written statements, set-based updates and deletes, and a lookup by key answered from objects already in memory.

go deeper

for a junior

Be able to say what the filter is in one sentence: a condition declared on the mapped type that the layer adds to the queries it builds, typically for live rows or the current tenant.

for a middle

Explain where the predicate is injected and where the parameter comes from, and name at least two paths the filter does not reach, such as hand-written statements and set-based updates.

for a senior

Show that you treat the filter as a safe default rather than a guarantee: you bind the tenant at the boundary, scope any switch-off to a single unit of work, and review the unfiltered paths by hand.

for a principal

Frame it as where the invariant lives — in the mapping, in each query, or in the database — and own the cost of an invisible predicate on readability, query plans, and onboarding.

## The problem it solves Once a table carries a marker column that decides whether a row still counts — a deleted timestamp, an `archived` flag, an owning-tenant identifier — every read against that table becomes conditionally correct. It is right only if it also carries the condition. In a codebase with hundreds of query sites, "everyone remembers to add `AND deleted_at IS NULL`" is not a design, it is a hope. The failure is silent and asymmetric: a forgotten predicate does not error, it simply returns rows the caller was never supposed to see, and in the tenant case those rows belong to another customer. An **always-on filter** (also called a global filter, a default scope, or a mapping-level restriction) moves the condition from the call sites to the mapping. You declare, once, that reads of this type are restricted by this predicate, and the layer appends it to the statements it builds. ## What "always-on" actually covers The rule of thumb is that the filter applies wherever **the layer itself composes the SQL**: - a query built through the layer's query API against that type; - the load of a single object by key **when it actually goes to the database**; - the load of a collection-valued association whose element type is filtered; - a lazily loaded reference resolved later, if the layer re-issues a query to resolve it. And it does not apply where the layer is not composing the statement, or is not issuing one at all: - a statement you wrote yourself and handed to the layer to execute; - a set-based update or delete that the layer translates straight to one SQL statement; - a lookup by key served from objects the unit of work already holds; - an object handed back from a cache that lives beyond one unit of work. Those holes are the whole reason the topic is interesting, and they are worth studying separately; the point to hold here is that "always-on" means *always on the statements the mapper writes*, not *always on every row that can reach your code*. ## Where the parameter comes from A live-rows filter is usually parameterless — the predicate is a constant condition. A tenant filter is not: it needs a value, and that value must not come from the caller, because the whole purpose is that callers cannot forget or falsify it. Layers therefore read it from **ambient context** established at the boundary: the authenticated principal's tenant, resolved when the request is accepted and bound to the unit of work for its lifetime. Two consequences follow immediately. First, code that runs outside a request — a scheduled job, an importer, a startup task — has no ambient tenant, and the design has to say what happens then. Second, the value must be bound per unit of work, not per process; a value cached in a process-wide slot will leak between concurrent requests. ## Filter versus writing the predicate by hand | | Always-on filter | Predicate at each call site | |---|---|---| | Default when someone forgets | Correct (filtered) | Wrong (leaks rows) | | Visible in the query source | No — it is invisible at the call site | Yes | | Exceptions (an admin view of everything) | Needs an explicit switch-off | Just omit the clause | | Applies to generated statements only | Yes | N/A — you write the statement | | Reviewability | One declaration to audit | Every query is a place to get it wrong | The trade is real: the filter buys a safe default and pays with invisibility. A newcomer reading a query cannot tell from the query text why it returned fewer rows than the table holds, and a slow plan may be caused by a predicate that appears nowhere in the code they are reading. Teams mitigate that by keeping the set of filters very small and documented, and by logging generated SQL when diagnosing. ## Switching it off Every real system eventually needs one screen that sees everything: a support tool that must show a cancelled account, a purge job that must find rows to erase. Layers that support the filter also support disabling it, and the sane scope for disabling is **one unit of work** — the narrow, explicit, short-lived block that has the exception — never a process-wide setting. A filter disabled globally at startup is the same as having no filter, with the added harm that the declaration still exists and reassures readers. ## Practical rules 1. Declare the filter on the mapped type, not in a base query helper that a caller can bypass. 2. Bind the tenant value at the boundary, once per unit of work, from the authenticated principal. 3. Decide, deliberately, what a missing ambient value does — refusing the work is usually safer than silently filtering on nothing. 4. Know the paths the filter cannot reach, and treat those as the places to review by hand. 5. Never treat the filter as a security boundary on its own: it narrows the statements the mapper writes, and anything that reaches the database another way is outside it.

  • Why should the tenant value for such a filter come from ambient context rather than a method parameter?
    Because a parameter can be forgotten, defaulted, or supplied from user input — the three ways a tenant filter is defeated. Binding it at the boundary from the authenticated principal means no call site can choose a different value, and code review can check one binding instead of every query.
  • What is the main cost of making the predicate invisible at the call site?
    Readability and diagnosis. A query's text no longer explains its result set or its plan, so a developer debugging "why is this row missing" or "why is this index not used" must know the filter exists. Keeping filters few, named, and documented, and reading the generated SQL when in doubt, is the usual mitigation.
  • Does an always-on filter make a soft-deleted row unreachable?
    No. It narrows the statements the layer composes. A hand-written statement, a set-based update, a key lookup answered from objects already tracked in the unit of work, or an entry from a cache spanning units of work can all still surface the row.

saying these in an interview costs you the question

  • Thinks the filter applies to hand-written SQL executed through the layer
  • Passes the current tenant as a query parameter at each call site
  • Stores the ambient tenant in a process-wide slot shared by concurrent requests
  • Believes a declared filter alone makes hidden rows unreachable
  • Disables the filter globally so one admin screen works