skip to content

When no change your team can make fixes an AI-powered feature's wrong answers, what must that conclusion rest on, and what is recorded instead of a fix?

level: seniorimportance: should knowfreq 38%

answer

  1. exhaustion is not evidence
  2. it is a negative claim, so measure
  3. a rate on a fixed input set
  4. record what was ruled out
  5. every limit needs a re-check condition

basics

~20 s

It rests on evidence, not exhaustion: a measured failure rate on fixed inputs, and every cheaper layer changed and re-measured one at a time. In place of a fix, record the limit, its conditions, its containment, and when to re-check.

solid answer

~50 s

The conclusion is a claim about the world, so it needs measurement behind it. What it must rest on: - a **stable failure rate over a fixed input set**, not one memorable reproduction; - each cheaper layer changed and re-measured separately: complete material, minimal wording edits, settings restored to tested values, the task decomposed into smaller asks; - the rate surviving all of them. What gets recorded in place of a fix is a durable statement of the limit: the conditions that produce it, the measured rate and the input set that measured it, the layers that were changed and ruled out, what the product now does when it hits the limit, and the condition that reopens the question — a new model version, richer material, or a rate that moves. Without that record the same investigation is repeated by the next engineer, at full cost.

code

yaml · 20 lines
yaml
limit:
  summary: totals disagree with the document on long statements
  conditions:
    - source document longer than roughly forty pages
    - answer requires exact arithmetic across many rows
  measured:
    inputSet: fixed-set-014        # 25 recorded user requests
    repeatsPerInput: 20
    unacceptableRate: 0.31
  ruledOut:
    - change: complete and current material supplied
      rate: 0.29
    - change: three separate minimal wording edits
      rate: 0.28
    - change: settings restored to tested values
      rate: 0.31
    - change: task split into extract-then-total
      rate: 0.30
  containment: totals computed in ordinary code; step only extracts rows
  recheckWhen: new model version, or richer source material available

go deeper

for a junior

Be ready to say why a wrong answer is not automatically a limitation of the generative model, and what the cheap checks are that come first. Knowing that this conclusion needs evidence is more important than knowing how to gather it.

for a middle

Explain the mechanics of establishing it: a fixed input set, repeated calls, a measured unacceptable rate, and one changed layer at a time with a fresh rate for each attempt. Interviewers want the measurement, not the impression.

for a senior

Show the judgment about when to stop and what to write down. Name the false positive of incomplete material, insist on a mechanism for why the task is structurally hard, and describe the record that stops the next engineer repeating your investigation.

for a principal

Own the standard. Decide how much evidence a limit claim needs before it is accepted, ensure containment is ordinary engineering rather than resignation, and make re-check conditions real so limits do not become permanent product constraints by neglect.

## Exhaustion is not evidence *We tried everything and it still gets it wrong* is a statement about a team's stamina, not about a generative model. The dangerous version of this conclusion is reached after an afternoon of unrecorded wording edits, because it is comfortable: it converts an open defect into a fact of nature and closes the ticket. It is also the conclusion that is most expensive to be wrong about, because a limit stays in the product long after the real defect would have been fixed, and every future complaint about the same behaviour is waved away by pointing at it. So the bar is deliberately higher than for an ordinary defect. An ordinary fix proves itself by making the failure go away; this conclusion has to prove that a whole class of fixes does *not* make it go away, which is a negative claim and needs correspondingly better evidence. ## What the conclusion has to rest on 1. **A measured rate, not an anecdote.** Assemble a fixed set of inputs that exhibit the behaviour, run each of them repeatedly, and record how often the answer is unacceptable. The number is the thing being explained, and every later experiment is judged against it. 2. **Each cheaper layer changed and re-measured, separately.** Complete and current material supplied; the instruction wording edited minimally, more than once; every setting restored to the value that was actually tested; the task decomposed into two smaller asks. Each attempt gets its own rate on the same input set. 3. **The rate surviving all of them.** Not *no attempt helped a bit*, but no attempt moved the rate beyond the variation between two identical batches. 4. **A plausible mechanism.** The task should be one that is structurally hard for a generative step rather than merely unfamiliar: exact arithmetic over a long document, faithful counting, strict ordering, or a guarantee that something is absent. A confident ceiling claim with no mechanism is usually an unfound defect. 5. **The alternatives considered.** Could the task be done in ordinary code, or narrowed so the generative step handles only the part it is good at? A ceiling reached without asking that question is a design decision disguised as a measurement. ## The record that replaces a fix | Field | What it holds | Why it matters | | --- | --- | --- | | Conditions | The input shapes and circumstances that produce the failure | Lets the next reader tell a new complaint apart from this known one | | Measured rate | How often it fails, on which fixed input set, at what repeat count | Turns a vague weakness into something comparable later | | Ruled out | Each layer changed, the attempt made, and the rate it produced | Stops the whole investigation from being repeated at full cost | | Containment | What the feature now does when it lands in this territory | Records that the product responded, not just that the team gave up | | Re-check condition | What would make this worth revisiting | Keeps the limit from becoming permanent by neglect | The record belongs where an engineer investigating the next complaint will find it, and it should be linked from the original complaint. A limit recorded only in a chat thread is a limit that will be rediscovered. ## Containment and reopening A recorded limit is usually accompanied by a change in the surrounding product, and that change is ordinary engineering: doing the hard part in deterministic code, narrowing what the feature offers to do, constraining the shape of the output so a bad answer is rejected rather than displayed, checking a claim against a source of truth before showing it, or asking the user to confirm before anything irreversible happens. None of these fix the generative step. They limit what its weakness can cost. The re-check condition matters as much as the rest. Model versions change, the material available to the feature improves, and the task itself may be reshaped by a later product decision. A limit with no re-check condition quietly becomes a permanent product constraint that nobody ever revisits, and teams routinely carry those for years past the point where they were true. ## How this conclusion goes wrong - **Concluding it from a single input.** One hard example proves the feature struggles with one example. - **Skipping the material check.** Incomplete or stale material imitates a ceiling almost perfectly, and it is the single most common false positive here. - **Testing wording once.** One rewrite that did not help is weak evidence against wording in general. - **Recording the conclusion without the evidence.** A note saying *this is a model limitation* with no rate, no inputs and no ruled-out list cannot be checked, defended or revisited by anyone later.

  • Which false positive most often looks like a capability ceiling but is not one?
    Material that never arrived, arrived out of date, or was cut to fit a size limit. From the outside it is indistinguishable: the answer is confidently wrong at a stable rate whatever the wording. The check is free, because you read the material exactly as it was sent rather than as the code was meant to send it, so it belongs ahead of any ceiling claim.
  • Why does a recorded limit need a re-check condition rather than a review date?
    A date fires whether or not anything changed, so it is either ignored or wastes an investigation. A condition ties the re-check to something that could genuinely move the rate: a new model version, better source material, or a reshaped task. It also tells the next reader what would falsify the limit, which is what makes the record a claim rather than an excuse.

saying these in an interview costs you the question

  • Declares a limit after an afternoon of unrecorded wording edits
  • Concludes a ceiling from one memorable failing example
  • Records the verdict without the rate or the inputs behind it
  • Never checks whether the material actually reached the step
  • Leaves the limit with no condition that would reopen it