skip to content

In a service-bulletin retrieval index, how do you separate submitter-authored metadata from pipeline-assigned metadata?

level: seniorimportance: should knowfreq 40%

answer

  1. one question per field: what wrote it
  2. parsed, assigned, or derived
  3. uniform formatting means code wrote it
  4. timing influence is not authorship
  5. only the query-time class is unreachable

basics

~20 s

Trace each field back to the moment it entered the index payload. Fields parsed out of the artefact - front matter, embedded properties, filename - were written by the submitter. Fields stamped at receipt or derived from the caller were not.

solid answer

~50 s

Ask one question per field: what wrote this value, and when. Anything parsed from the uploaded artefact - front-matter status, a declared effective date, an equipment model, an embedded document property, a filename fragment - is the submitter's prose in a structured slot, and a predicate over it is a restatement of the document. Anything the ingest stage stamps at receipt (arrival timestamp, connector, authenticated portal account, content hash) was created outside the file, so a document author cannot write it, although they still choose when and from which account to submit. Anything derived from the caller's session at query time is unreachable from inside a document altogether. Then say what each class establishes: an arrival timestamp proves when bytes landed, not that guidance is current; an account proves which credential uploaded, not that anyone reviewed it. Only the third class narrows against the author rather than for them.

code

json · 15 lines
json
{
  "filter": { "status": "current", "effective_date": { "gte": "..." } },
  "candidates": [
    {
      "chunk_id": "c-8812#2", "score": 0.81, "passes_filter": true,
      "parsed_from_file": { "status": "current", "effective_date": "...", "equipment_model": "..." },
      "assigned_at_ingest": { "received_at": "...", "submitted_by": "portal-acct-4417", "connector": "supplier-portal" }
    },
    {
      "chunk_id": "c-1140#7", "score": 0.79, "passes_filter": false,
      "parsed_from_file": { "status": null, "effective_date": "...", "equipment_model": "..." },
      "assigned_at_ingest": { "received_at": "...", "submitted_by": "portal-acct-0102", "connector": "supplier-portal" }
    }
  ]
}

go deeper

for a junior

Know the basic split: some metadata is copied out of the uploaded file and some is created by the pipeline, and only the second kind was not written by the person who submitted the document.

for a middle

Practise the trace on a real payload - name for each field whether it was parsed from the artefact, stamped at receipt, or derived from the caller, and say which of those a predicate can safely narrow on.

for a senior

Show the judgment: name residual influence such as submission timing, resist calling normalisation a provenance change, and state precisely what each field lets you claim rather than declaring the index clean or dirty.

for a principal

Own the framing that a predicate's assurance is bounded by the weakest provenance among the fields it reads, and be able to defend that bound to someone who wants the filter counted as a check.

## The audit is per field, not per index An index payload looks homogeneous - one flat object of metadata beside each chunk - which is exactly why this analysis gets skipped. The fields have wildly different origins, and the only way to know what a predicate over them establishes is to trace each field back to the moment it was written. Three classes are worth separating. **Parsed from the artefact.** Front matter at the top of the file, embedded document properties, a header table, a fragment of the filename, a value extracted from the body by a rule. Whatever the submitter wrote is what the index holds. A predicate over these fields tests conformance of a self-description. **Assigned by the pipeline at receipt.** The arrival timestamp, the source connector or bucket, the authenticated account the upload came from, a content hash, an internal document id. These are created outside the artefact. A document author cannot write them - though note the residual influence: they choose when to submit, so an arrival timestamp is under their timing control, and an account field is only as meaningful as the portal's account provisioning. **Derived from the caller at query time.** A scope computed from the session or the requesting identity. Nothing written inside a document can influence a value produced from whoever is asking. This is the only class an author is structurally locked out of. ## Reading a result set Given candidate rows and their metadata, the mechanical check is: for each key in the predicate, which of the three classes is it in? A predicate composed entirely of class-one fields narrows using the adversary's own text. That is not a subtle failure - it is the filter doing the sorting on the plant's behalf, because untidy legitimate rows fail a conformance test that a purpose-written row passes exactly. The tell in a payload is often structural rather than semantic. Fields that mirror the document's own vocabulary, in the document's own casing, with the document's own optionality (sometimes present, sometimes null) came out of files. Fields that are uniformly populated across every row, in one canonical format, were produced by code. ## What each class proves, stated carefully - A declared status proves what its author typed in that slot. It does not prove an approval happened. - A declared effective date proves what its author typed. It does not prove the guidance is in force. - An arrival timestamp proves when bytes reached the pipeline. It does not prove the content is current, and it does not prove anybody read the content. - A submitting account proves which credential authenticated an upload. It does not prove a human vetted the file, and it does not prove the account belongs to the party the document names. - A content hash proves the bytes were not altered after receipt. It says nothing about whether they should have been trusted before it. - A session-derived scope proves something about the asker, which is why an author cannot aim at it. Getting the direction of any of these backwards encodes the exact misreading that makes the family work. ## The intermediate cases that decide interviews Normalisation at ingest is the interesting middle ground. If the pipeline rewrites a declared value - canonicalising a model string, mapping declared status tokens through a lookup, refusing a date it cannot parse - the stored field is no longer a verbatim copy, and some malformed plants stop being eligible. That is real, and it is worth naming, but it is a change of format, not of provenance: a submitter who writes a well-formed conformant value still gets that value into the payload. Do not overclaim it. Cross-referencing is the other case. A field populated by looking the document up against a register held outside the submission path has a different provenance from a field the document declared, even if the two carry the same name. The question is never what the field is called; it is what wrote it. ## Why this is the senior version of the question Anyone can say a filter is only as good as its source. The production skill is doing the field-by-field trace on a real payload, resisting the pull of names that sound authoritative, naming the residual influence in the second class rather than declaring it clean, and stating precisely what each surviving field lets you claim.

  • The pipeline canonicalises declared model strings at ingest. Does that move the field out of the authored class?
    No. Normalisation changes format, not provenance. It ejects malformed entries, which removes sloppy plants along with sloppy legitimate rows, but a submitter who writes a well-formed conformant value still lands that value in the payload. The field is still a self-description, just a tidier one.
  • An arrival timestamp is stamped by the pipeline. How much influence does a submitter still have over it?
    They control timing, not content. They choose the moment of upload, so they can place a document inside a recency window by submitting at the right time - but they cannot back-date it, and they cannot make it claim anything else. That residual influence is worth naming rather than filing the field as untouchable.
  • What does a submitting-account field actually establish about a bulletin?
    That an upload authenticated as that credential. It does not establish that a person reviewed the file, that the account belongs to the supplier the document names, or that the content is accurate. It is evidence about a submission event, which is a different claim from evidence about the document.

saying these in an interview costs you the question

  • Treats every field in the index payload as equally trustworthy
  • Calls an ingest timestamp proof that content is current
  • Reads a submitting account as evidence of review
  • Declares pipeline-assigned fields wholly out of the submitter's influence
  • Assumes normalisation at ingest changes who authored a value

context