skip to content

Why can two services that both implement 'documents not tagged archived' return different documents for the same corpus?

level: seniorimportance: should knowfreq 40%

answer

  1. not-X is relative to something
  2. name the collection you subtract from
  3. corpus, tenant, page or shard
  4. unindexed is not the same as untagged
  5. a stored complement fails open

basics

~20 s

Because a complement is meaningless until its universe is fixed. 'Not tagged archived' names a set only relative to a chosen collection — the whole corpus, one tenant, one shard, or the current page — and two services that pick different universes implement different filters.

solid answer

~60 s

Complement is not a property of a set; it is a **binary operation against a universe**: `U \ A` depends on `U` just as much as on `A`. Two services can agree perfectly on `A`, the set of documents tagged `archived`, and still disagree because one takes `U` as the whole corpus, another as the documents the caller is allowed to see, a third as the current result page, and a fourth as a single index shard. Each returns a correct complement of a *different* universe. A second source of disagreement is what 'not tagged' means for a document whose labels were never indexed: absent is not the same as known-to-be-absent, and the two services may classify it differently. The practical discipline is to write exclusions as an explicit difference `U \ A` with `U` named in the same expression, rather than as a bare 'NOT archived', and to fix where in the pipeline the complement is taken — before or after paging, before or after the visibility filter — because those orders do not commute.

go deeper

for a junior

Remember that a complement always needs a universe: 'not tagged archived' only names a set once you say which collection it is taken from.

for a middle

Explain the difference form and why it is preferred to a bare negation, and name the universes a request might plausibly be using between the client and one shard.

for a senior

Diagnose the disagreement: compare the two services' pipeline order, show how excluding after paging shortens pages, and say which stage the exclusion belongs in and why.

for a principal

Decide the invariant: exclusions carry their universe explicitly, and anything governing visibility is defined positively, because a stale complement fails open while a stale positive set fails closed.

## Complement is a two-argument operation wearing a one-argument notation Writing `A'` or saying 'not archived' looks like a property of `A`. It is not. The complement of `A` is `U \ A` for a chosen universe `U`, and changing `U` changes the answer while `A` stays put. Every disagreement in this question comes from that hidden second argument. ## The universes a real filter might be using - **The whole corpus** — every document the system stores. - **The tenant's documents** — a partition of the corpus, which is what most product copy actually means. - **The visible set** — documents this caller is permitted to see, itself a filter. - **One shard or partition** — what a node can answer for locally, before results are merged. - **The current result page** — what has already been narrowed, sorted and cut to fifty rows. - **The indexed set** — documents whose labels have actually been ingested. Each is a legitimate universe and each yields a different 'not tagged archived'. The two services in the question are not buggy in isolation; they are answering different questions with the same words. ## A concrete disagreement Take a corpus of 100 documents of which 10 are tagged `archived`. Service one takes the corpus as the universe and returns 90. Service two applies the caller's visibility filter first, leaving 40 documents of which 4 are archived, and returns 36. Service three excludes after paging: it takes the first 50 rows of a sorted result, removes the 6 archived ones it finds there, and returns 44 rows — a page that is now short, and whose missing rows do not come back on any later page. All three computed a correct complement. Only one of them computed the filter the product promised. ## Order matters, because these steps do not commute | Pipeline order | What the user gets | |---|---| | Exclude, then page | Full pages; excluded documents never occupy a slot | | Page, then exclude | Short, ragged pages; total count no longer matches | | Visibility filter, then exclude | Exclusion applied within what the caller may see | | Exclude, then visibility filter | Same final set, but counts computed midway are wrong | Restricting and then complementing is not the same as complementing and then restricting. In general, complement does not survive being moved across another operation unchanged — that is exactly what De Morgan's rewrites exist to handle for unions and intersections, and paging (a cut of a sorted sequence) is not a set operation at all, which is why moving it across a complement is the most damaging of the four orders above. ## Absent, or unknown? A document whose labels were never indexed carries no `archived` label as far as the query engine can see, so a naive complement includes it in 'not archived'. Whether that is right is a product decision, not a mathematical one: set algebra has no third state, so a system that wants one must model it explicitly — for example by keeping a separate set of documents whose labelling is known complete and intersecting with it. Ecosystems differ in how their query layers treat a missing attribute in a negation, so an engineer who assumes one behaviour transfers a wrong assumption when they move; the safe habit is to state the intended treatment in the filter rather than inherit it. ## The stability problem A complement is not stable as the corpus changes. Add a document, and it joins `U \ A` automatically unless it is tagged. That makes a stored complement — a materialised 'not archived' list, a cached exclusion result, a permission set defined by negation — go stale in a direction that **adds** rows rather than dropping them, which is the dangerous direction for anything protecting visibility. A positive definition ('documents tagged active') fails closed as the corpus grows; a negative one fails open. ## What to do 1. Express an exclusion as an explicit difference with the universe written down: `visible_to_caller \ archived`, never a bare `NOT archived`. 2. Fix the pipeline position of the exclusion and test it, especially against paging and counts. 3. Decide in the contract what an unindexed or unknown document means for the filter. 4. Prefer positive definitions for anything that controls access, and treat any stored complement as a cache that must be invalidated on ingest.

  • Why does excluding after paging produce short pages that never recover their missing rows?
    Paging cuts a sorted sequence to a fixed window, and removing rows from that window leaves it under-full. The removed slots are not backfilled from later rows, because the next page starts after the original cut. Exclusion must happen while the result is still a set, before it is ordered and cut, or the page size becomes a function of the data.
  • Why is a stored 'not archived' list riskier than a stored 'active' list?
    Both go stale, but in opposite directions. A newly ingested document belongs to the complement automatically, so a stale negative list grows to include documents nobody vetted — it fails open. A positive list omits the new document until something adds it, so it fails closed. For anything that controls visibility, the failure direction decides which definition is acceptable.
  • How would you model 'we do not know whether this document is archived'?
    Set algebra has only membership, so a third state must be carried by a second set: keep the documents whose labelling is known complete and intersect the exclusion with it. The unknown ones then fall out of the answer explicitly instead of being swept into the complement by default, and the choice becomes visible in the expression.

saying these in an interview costs you the question

  • Treats 'not tagged archived' as an absolute set of documents.
  • Applies the exclusion after paging and calls it the same filter.
  • Assumes an unindexed document is simply one with no labels.
  • Believes a stored complement stays correct as documents arrive.
  • Expects services with different universes to return the same exclusion.
  • Defines an access-control set by negation and calls it equivalent.