In a corpus whose status flags were copied from submitted files, what can a retrospective review prove?
answer
- the sorting field has no discriminating power
- look outside the documents for evidence
- submission is not review
- a sample bounds the sample
- unverified is not verified-clean
basics
~20 sVery little from inside the index. A planted row and a legitimate one carry self-declared values that pass the same predicate identically, so separation depends on records made outside the documents - and those establish submission events, not correctness.
solid answer
~50 sStart by saying what the index cannot settle. If currency was asserted by each file about itself, then a planted bulletin and a real one are indistinguishable on exactly the field you would sort on, and no amount of querying separates them. What a review can work with is evidence created outside the artefacts: which account submitted each file and when, whether that account maps to the supplier the document names, whether an out-of-band copy exists at the supplier. That establishes submission provenance, never content correctness. Sampling is similarly bounded - confirming twenty bulletins against external copies supports the claim that twenty were confirmed, not that the corpus is clean, because a plant written to look ordinary is the case sampling is worst at finding. And be clear that changing what the query narrows on affects future retrievals only; it re-classifies nothing already indexed.
go deeper
Understand that if a document declared its own status, then reading that status back out of the index tells you nothing new about whether the document is genuine.
Explain where usable evidence lives instead - submission records, account history, comparison against a copy held outside the pipeline - and what each of those actually establishes.
Demonstrate the assurance discipline: state findings with counts and selection method, refuse to turn a sample into a corpus-wide claim, and separate the forward-looking query question from the historical provenance one.
Be ready to tell an owner that the unverified subset is not measurable from inside, to scope verification by consequence rather than by convenience, and to defend saying so rather than producing a reassuring number.
## The chair this question is asked from You have inherited a retrieval corpus - equipment service bulletins, ingested from a supplier portal over several years, with front-matter status and effective dates copied into the index payload and a query that narrows on them. Someone now asks a reasonable question: which of these bulletins are legitimately current? The honest first move is to explain why the corpus itself cannot answer. ## Why the index cannot separate the two populations A planted bulletin and a genuine one both carry values their authors wrote. Both pass the predicate. The plant, if it was written competently, is the more conformant of the two, because conformance was its design goal while the genuine document's metadata is an accident of whoever prepared it. So the field you would naturally sort on is precisely the field with no discriminating power. Querying harder does not help: every query reads the same self-declared values. This is worth stating plainly to whoever asked, because the instinct is to run a report. A report over authored fields returns the attacker's assertions in tabular form. ## What can be established, and what it means Everything usable was created outside the artefacts: - **Submission records.** Which portal account uploaded each file, and when. This establishes a submission event. It does not establish that the content is correct, and it does not establish that the account belongs to the supplier the document names - only that a credential authenticated. - **Account provisioning history.** Whether the credentials that submitted were issued to the parties they claim to represent, and whether any were shared, dormant, or issued outside the normal process. This narrows the population worth looking at; it does not classify documents. - **Out-of-band comparison.** Fetching the supplier's own copy of a bulletin and comparing. This is the only step that speaks to content at all, and it is expensive, partial, and only as good as the channel you fetched over. - **Structural anomalies.** Rows whose metadata is uniformly perfect in a corpus that is otherwise untidy are worth a look, but this is a lead, not a verdict: a careful supplier produces clean front matter too, and a careful plant can be written to look average. ## What sampling can honestly be claimed to buy A spot-check bounds nothing about the unsampled remainder when the population you are hunting was authored to be unremarkable. The defensible sentence is "n bulletins were confirmed against an external copy, and m discrepancies were found", with n and the selection method stated. The indefensible sentence is "the corpus is clean". If the sample was chosen by the same conformance signal the plant optimises for, the sampling is anti-correlated with the thing it was meant to find, and saying so is part of the answer. ## The distinction that gets missed Two different problems are in the room and they have different remedies and different evidence. | Question | What it is about | | --- | --- | | What will future queries narrow on? | The query and the ingest path - forward-looking, decidable today | | Which rows already in the index were planted? | Provenance of past submissions - historical, often undecidable | Changing what the predicate reads - moving to a value the pipeline stamps, or a scope derived from the caller - changes the first. It re-classifies nothing in the second. Re-indexing changes even less: the same values are parsed out of the same files and land in the same slots. Neither is a review, and presenting either as one is the error the question is testing for. ## How to close it Say what you can support and mark the rest unknown. A corpus of this shape has an unknown subset, its size is not estimable from inside, and the honest deliverable is a scoped verification of the documents that matter most - the ones covering equipment where following the wrong bulletin has the worst outcome - plus a clear statement that everything else is unverified rather than verified-clean. Interviewers are listening for whether you will say the words "we cannot tell, and here is the boundary of what we can tell" instead of producing a reassuring number.
- A technician reports a bulletin that should not be current. Is that a finding about one document or about the corpus?About the trust model. One wrong document is an incident; the finding is that currency was asserted by each file about itself, which means the same condition is unfalsifiable across every other row. Report it as a property of the ingest and query path, with the single document as the demonstration rather than the scope.
- Does re-indexing or re-embedding the corpus help?No. Re-embedding changes vector representations, and re-indexing re-parses the same files into the same payload slots. Neither touches provenance: the declared status that made a row eligible is read out of the same document and lands in the same field. It is a rebuild, not a review.
- How would you word the assurance statement at the end of such a review?Narrowly and with counts: how many documents were compared against an external copy, how they were selected, what discrepancies were found, and an explicit statement that the remainder is unverified. Avoid any sentence that converts a sample into a property of the corpus, and name the fields whose values remain self-declared.
saying these in an interview costs you the question
- Claims a clean spot-check means a clean corpus
- Runs a report over self-declared fields and calls it an audit
- Treats an ingest timestamp as proof that content is current
- Thinks re-indexing or re-embedding removes planted documents
- Says changing the query re-classifies rows already indexed