skip to content

Fields the Filter Trusts

The date and status a filter checks are copied out of the document, so whoever writes the file also decides which scoped queries it can answer. Interviewers probe past the reflexive tenant answer.

on this pageshow

explore

questions

4

In RAG retrieval, why doesn't a metadata filter reading a document's own status field exclude a planted passage?

level: juniorimportance: must knowfreq 58%

answer

  1. ask where the value came from
  2. the predicate inherits the field's provenance
  3. the author wrote the status too
  4. narrowing drops rivals, not the plant

basics

~20 s

A metadata predicate is only as trustworthy as where its value comes from. When status, date or type are copied out of the submitted document, the passage's author sets them too, so narrowing removes honest rivals and keeps the plant.

solid answer

~50 s

A retrieval predicate looks like a claim about provenance, but it inherits the provenance of the field it reads. Take a field-service assistant answering from equipment service bulletins, narrowed on `status`, `effective_date` and equipment model. If those come from front matter copied verbatim out of each bulletin submitted through a supplier portal, then whoever submits a bulletin also chooses its status and its date. A plant satisfies the predicate by construction. Worse, the filter is working for the person who planted it: honest bulletins that were never marked current, or that carry a different model string, get dropped, and the plant sits in the small surviving set. The construction only dies where the predicate reads a value the submitter never touches, such as one the pipeline stamps at receipt or one derived from the caller's session. Narrowing is not protective in itself; the question is always who authored the value.

go deeper

for a junior

Be ready to say, in one sentence, that a filter is only as trustworthy as the source of the value it reads, and that fields copied out of a submitted document are written by whoever submitted it.

for a middle

Explain the ingest path concretely: front matter parsed into the index payload, the predicate reading that payload, and the difference between a value parsed from the artefact and one stamped by the pipeline or derived from the session.

for a senior

Show the second-order effect an interviewer is listening for - the narrowed set is not just no safer, it is enriched, because untidy legitimate documents fail the predicate while a purpose-written one passes it exactly.

for a principal

Be able to frame this as a provenance question rather than a query question, and to say what a retrieval predicate can honestly be claimed to establish when its inputs are supplied by the parties being filtered.

## The claim a filter appears to make A retrieval query in a RAG application usually does two things: it narrows a candidate set with a predicate over metadata, then it ranks what survives by vector similarity. The narrowing step reads structured fields sitting beside each chunk in the index payload - a status token, an effective date, a document type, an equipment model. Because those fields look like database columns, they carry an unearned air of authority: a reader assumes something checked them. Nothing checked them. The index payload is populated at ingest, and a very common ingest design parses those fields out of the artefact itself - YAML or key-value front matter at the top of the file, embedded document properties, a header table, sometimes the filename. The value in the index is a copy of a value in the document. ## Two provenances, and only one of them resists an author The useful discipline is to trace every field the predicate reads back to the moment it entered the payload. | Field the predicate reads | Who put the value there | | --- | --- | | status parsed from the file's front matter | whoever submitted the file | | effective date declared inside the document | whoever submitted the file | | equipment model declared inside the document | whoever submitted the file | | receipt timestamp stamped by the ingest stage | the pipeline | | submitting account recorded by the portal | the portal's authentication, not the file | | a scope derived from the caller at query time | the session | Everything in the first group is authored by the same person who authored the passage. A predicate over those fields is not a check on the document; it is a restatement of what the document says about itself. ## The bulletin case Consider an assistant that answers technicians' questions from equipment service bulletins, where suppliers submit bulletins through a portal and the pipeline copies their front matter into the index. The application narrows to bulletins whose status is current, whose effective date falls in a window, and whose equipment model matches the unit the technician named. That predicate exists to keep obsolete guidance out of the answer, and for honest traffic it does exactly that. A planted bulletin fills the same fields in. It declares itself current, dates itself inside the window, and names the model the technician will ask about. It is eligible by construction, because eligibility was defined by fields it wrote. ## The filter does the attacker's sorting The second-order effect is the part candidates miss. Narrowing removes candidates, and the plant is never one of the removed. Real corpora are untidy: legitimate bulletins have blank status fields, dates in a different format, model strings written three ways. Those are the documents the predicate drops. The plant's metadata is complete and exactly conformant, because it was written to satisfy a predicate rather than to describe a document. So the narrowed result set is smaller, cleaner, and contains proportionally more of the planted content than the unnarrowed one did - and a technician reading three results instead of fifty tends to read the shortness as vetting. ## Direction of the claim Be precise about what each artefact proves. A status field records what its author wrote in that field, not that anyone approved anything. A declared effective date records what its author typed, not that the guidance is in force. Passing the predicate proves a passage carried conformant values, not that it is correct, current or from the party it claims. Getting these backwards is the misconception the whole topic exists to correct. ## Where the construction stops working It stops exactly where authorship stops. A value the ingest stage assigns at receipt is out of the submitter's pen, although they still choose when to submit. A value derived from the caller's session at query time is out of reach entirely - the person writing a document cannot influence a field that is computed from whoever is asking. That asymmetry, not the act of narrowing, is what separates a predicate that resists a document author from one that serves them. ## What this is not It is not an argument about how filtering is implemented - whether the predicate is applied before or after the vector search, and what that does to recall, is a retrieval-engineering question with its own answers. It is also not about the model being tricked: no instruction has been obeyed here and no refusal has been bypassed. The retrieval stage has simply been asked to sort on data the adversary supplied.

  • Does it help if the submission portal validates that status is one of a fixed set of tokens?
    Validation constrains the vocabulary, not the authorship. An enum check proves the submitter chose a legal token, which is what a plant would do anyway - it needs a conformant value to stay eligible. The only thing schema validation removes is a malformed plant, and a malformed plant was never the threat.
  • Which fields in that index payload does the submitter not control?
    The ones created outside the file: the receipt timestamp the ingest stage stamps, the portal account the upload authenticated as, the connector or source path the pipeline recorded. Note what they establish though - they identify a submission, not a review, so they are evidence about who and when, never about whether the content is correct.
  • If the predicate reads a field the pipeline derives from the caller's session instead, what changes for the attacker?
    They lose the pen. A document author can write anything into a file, but they cannot write a value that is computed from whoever is asking the question at query time. That is why the same narrowing operation can be inert against a document author or a gift to one, depending entirely on the field it reads.

It is an application form with an approved box the applicant ticks themselves. The box is filled in, and it tells you only what the applicant wanted written there.

saying these in an interview costs you the question

  • Says filtering on metadata always reduces retrieval risk
  • Treats a status value as evidence that someone approved the document
  • Assumes the vector store validates or verifies metadata on insert
  • Confuses narrowing a candidate set with vetting it
  • Thinks the ranking model can tell a conformant plant from a real bulletin

context

open as a page

Why can narrowing a retrieval query on document-supplied metadata raise a planted chunk's odds of being returned?

level: middleimportance: should knowfreq 45%

basics

~20 s

Narrowing removes candidates, and a planted chunk is never among them: it satisfies the predicate by construction, while legitimate documents with blank or non-conformant metadata are dropped. Fewer rivals means a better chance at the top-k slots.

open as a page

In a service-bulletin retrieval index, how do you separate submitter-authored metadata from pipeline-assigned metadata?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Trace each field back to the moment it entered the index payload. Fields parsed out of the artefact - front matter, embedded properties, filename - were written by the submitter. Fields stamped at receipt or derived from the caller were not.

open as a page

In a corpus whose status flags were copied from submitted files, what can a retrospective review prove?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Very little from inside the index. A planted row and a legitimate one carry self-declared values that pass the same predicate identically, so separation depends on records made outside the documents - and those establish submission events, not correctness.

open as a page