skip to content

In a reconstruction built from repeated samples, does span agreement prove recall?

level: seniorimportance: nice to knowfreq 33%

answer

  1. the test is asymmetric
  2. disagreement rules out, agreement does not rule in
  3. a peaked distribution has several causes
  4. format-shaped spans converge on their own
  5. no sampling variation, no signal

basics

~20 s

No. Disagreement between independent samples is strong evidence a span was generated rather than read. Agreement only proves the output is stable, which both genuine recall and a strongly format-shaped guess produce. Agreement is necessary for recall, never sufficient.

solid answer

~50 s

The test is asymmetric and most people use it as though it were symmetric. If five independent samples return five different order numbers for the same field, the field is being invented - a span actually present in the context does not vary that way. That direction is close to conclusive. The reverse is weak: agreement shows the model's output for that span is sharply peaked, and a heavily constrained format peaks all by itself. Dates, order-number patterns, boilerplate closings and common names converge across samples with no source text behind them at all, and those are exactly the spans a triager most wants to believe. Two caveats decide whether the test is usable: it needs genuine sampling variation, so a near-deterministic decoder erases the signal entirely, and independent means separate runs, not one run asked twice.

go deeper

for a junior

Recall that a model can produce confident text with nothing behind it, so a recovered value is not automatically real. Knowing that repetition across separate runs is the crude test is enough at this level.

for a middle

Explain why copying from context is low-entropy while filling in is not, and therefore why disagreement is the informative direction. Be able to state what agreement alone establishes about the output distribution.

for a senior

Show that you use the test asymmetrically, check the decoder before trusting stability, and keep a stable-but-unconfirmed bucket instead of promoting it. Naming format-constrained spans as the weak case is what separates a real answer here.

for a principal

Own the standard of evidence a programme applies to reconstructions, including whether confirming a span against a live customer record is a step your organisation is willing to take and who authorises it.

## The problem being solved You have been handed four hundred captured auto-replies from an unattended triage workflow and a reconstructed passage somebody assembled from them, said to be another customer's ticket - a name, an order number, dates, an internal resolution note. Your job is to say which parts are evidence and which parts are the model producing plausible text. Nothing in the capture distinguishes them: a recalled span and a confabulated span arrive in the same register, in the same reply shape, past the same screening layer. The only lever you have is repetition. Ask for the same region across many independent runs and look at where the samples disagree. ## Why the test works in one direction A span that is genuinely present in the run's context is being copied. Copying is a low-entropy operation for a model: across samples, the same characters keep coming back, because the context makes that continuation overwhelmingly likely. A span the model is filling in is being drawn from a much flatter distribution. There is no order number in front of it, so it produces something order-number-shaped, and next time a different one. Variation across independent samples is therefore strong evidence of generation. This is the reliable half of the test, and it is the half that lets you strike spans off a reconstruction with confidence. ## Why the reverse inference is weak Agreement establishes that the output distribution for that position is sharply peaked. It does not establish why. A distribution peaks for at least three reasons that have nothing to do with a source text being present. The span may be format-forced: a date in an obvious format, a reference number with a fixed prefix, a closing line the surrounding style makes near-inevitable. It may be a high-frequency default: a very common given name, a round quantity, a canonical phrasing. Or it may be shaped by the request itself, if the captured replies were elicited in a way that suggested the answer's form - in which case every sample agrees because every sample was pushed the same way, and the agreement measures the request, not the context. So the correct reading is: disagreement rules a span out, agreement fails to rule it out. That is much less than confirmation, and a reconstruction presented as verified because it was stable is overclaiming. ## The two conditions that make the test usable at all First, the decoder has to vary. If the deployment decodes near-deterministically, every sample agrees on everything and the signal is identically zero - you will observe perfect stability across a recalled span and an invented one alike. A triager who reports high sample agreement without knowing whether the deployment samples at all has measured nothing. Where the decoder cannot vary, any variation has to come from the request side, and variation introduced that way changes the conditions between samples, which weakens the comparison. Second, independent has to mean separate runs. Asking twice inside one context is not two samples: the second answer is conditioned on the first, which is the strongest possible pressure toward agreement, and an invented span becomes self-consistent the moment it is on the page. ## What actually confirms a span Only something outside the captured set. Checking a recovered order number against the real record confirms it. Structural corroboration helps a little - a recovered date that is consistent with a recovered reference number, both stable, is better than either alone - but consistency is also what a fluent generator produces, so this raises confidence rather than settling it. In practice a triaged reconstruction ends up in three buckets: confirmed against a source, struck out by disagreement, and stable-but-unconfirmed. The third bucket is usually the largest and it is the one people quietly promote into the first. ## The trap worth naming in an interview The spans that most reward stability testing are the ones the test handles worst. Names, dates and identifiers are the payoff of the whole exercise and also the most format-constrained things in the passage - the categories where a model converges on a plausible value with no source at all. The prose in between, which nobody cares about, is where sample agreement is actually informative. ## What an interviewer is listening for The asymmetry, stated as an asymmetry; the decoder caveat; the difference between separate runs and repeated turns in one context; and a willingness to leave spans in the unconfirmed bucket rather than rounding stability up to proof.

  • The workflow decodes near-deterministically. What happens to your sample-disagreement test?
    It stops measuring anything. Every sample agrees on every span, recalled or invented, so stability carries no information at all. Any variation then has to be introduced from the request side, which changes the conditions between samples and weakens the comparison. Reporting high agreement without knowing the decoder's behaviour is reporting a constant.
  • Which spans in a reconstruction should you trust least, and why is that awkward?
    The most format-constrained ones - identifiers, dates, names, boilerplate. A model converges on a plausible value for those with no source text behind it, so their stability is nearly uninformative. It is awkward because those are precisely the spans that make the finding worth filing, while the unremarkable prose between them is where agreement actually means something.
  • Is asking the same question twice in one run two independent samples?
    No. The second answer is conditioned on the first, which is the strongest available pressure toward agreement, and an invented span becomes self-consistent as soon as it is in the context. Independence here means separate runs over the same hidden text, which is also what a per-ticket session reset happens to give you for free.

Five witnesses giving five different licence plates tells you nobody read the plate. Five witnesses giving the same plate might mean they read it - or that it is the format everyone guesses.

saying these in an interview costs you the question

  • Treats sample agreement as confirmation of recall
  • Ignores whether the deployment samples at all
  • Counts repeated turns in one context as independent samples
  • Trusts identifiers and dates because they looked stable
  • Says a screening layer would have caught invented data

context