skip to content

Testing AI-Powered Features

Testing the product built around a model, not the model itself: acceptance criteria for variable output, suite seams and repeats, and the ship call. Interviewers probe where that line falls.

on this pageshow

questions

page 1 of 2

A feature's answer text comes from a generative step. In what distinct ways can that step fail to deliver an answer?

level: juniorimportance: must knowfreq 56%

answer

  1. Failing to speak is not one event
  2. Four shapes, not one error branch
  3. Deadline, volume refusal, outage, oversized input
  4. Only some are worth a second attempt
  5. One catch-all case proves almost nothing

basics

~20 s

Four delivery failures matter: nothing comes back inside the deadline, the call is refused for volume, the provider is unavailable, or the input exceeds what the step accepts. Each strands the user differently, so each earns its own test case.

solid answer

~40 s

Sort them by what the caller observes. **No answer inside the deadline** — the feature's own waiting budget expired while the call was still open, so it must decide without an answer. **Refused for volume** — the provider declined because too many requests arrived, and the same request would likely succeed shortly after. **Unavailable** — the provider is not answering at all, and trying again inside the same interaction rarely helps. **Input too large** — the request is rejected before any generation happens, deterministically, so waiting never fixes it; the feature must shorten or split the material, or say so plainly. A single "it errored" test case proves an error was caught, not that any of these four leaves the user somewhere sensible.

code

pseudocode · 20 lines
pseudocode
CASES for a feature whose text comes from a generative step:

no_answer_inside_deadline:
  given the step returns nothing within budget_ms
  expect fallback panel shown, entered text still present

refused_for_volume:
  given the step declines the call as over the allowance
  expect a retry control, enabled after the suggested pause

provider_unavailable:
  given the call to the step cannot complete at all
  expect fallback panel, manual route offered, no retry loop

input_too_large:
  given material exceeds the accepted size
  expect a size message, and assert no call was made

anti_case:
  one shared "it errored" branch asserting only "no crash"

go deeper

for a junior

Be ready to name the failures without prompting: no answer inside the deadline, a refusal for volume, an unavailable provider, and material larger than the step accepts. Naming the four is the whole ask at this level.

for a middle

Explain what separates them mechanically — which are transient, which are deterministic, and which the feature can see coming before it spends a call. An interviewer expects you to say which of the four a second attempt could plausibly fix.

for a senior

Show how the four become four cases with four different assertions, and how you keep the feature's own waiting budget shorter than anything it depends on. Talk about what production tells you afterwards: how often each of the four actually fires.

for a principal

Own the policy question of how much degradation the product will pay for. Handling all four properly costs design work and spend on extra calls, and deciding which failures deserve a queued promise and which deserve an honest apology is a product tradeoff, not an implementation detail.

A feature that gets part of its output from a generative step depends on something that can be slow, busy, absent, or unwilling to accept the request at all. Those are four different events, they strand the user in four different ways, and a suite that folds them into a single "the step errored" case proves only that an error was caught somewhere. ## The four ways it fails to deliver **No answer inside the deadline.** The feature set a budget — the longest it is willing to make someone wait — and that budget expired while the call was still open. The budget has to be the feature's own: a feature that sets none inherits whatever the underlying connection allows, which is usually far longer than any person will sit and watch. The distinguishing property of this failure is that the work may still be happening. It may finish after the feature has given up, which matters if the work is chargeable or has already changed something. **Refused for volume.** The provider took the connection and declined the work because too many requests arrived in some window — from this feature, from the rest of the product, or from everyone sharing an allowance. The distinguishing property is that the identical request would probably succeed a short time later, and the refusal frequently carries a hint of how long to wait. This is the only one of the four where waiting is a real remedy. **Unavailable.** The provider is not answering, or is answering with a hard failure that has nothing to do with this particular request: a dropped connection, an expired credential, a dependency of the dependency that is down. The distinguishing property is duration — it typically outlasts one interaction, so trying again inside the same session is close to useless and the feature has to degrade properly. **Input too large.** The request was rejected before any generation began, because the material exceeds what the step accepts. The distinguishing property is that it is deterministic: the same input is rejected the same way forever. The feature has to shorten, split or summarise the material, or say plainly that this document is too big — and it can usually detect the problem itself before spending a call at all. | Failure | Would a second attempt help? | What the feature owes the user | | --- | --- | --- | | No answer inside the deadline | Sometimes | A bounded wait, then an honest fallback | | Refused for volume | Yes, after a pause | A pause that respects any suggested wait | | Unavailable | Rarely within one session | A route that does not need the step | | Input too large | Never | A specific message about size | ## Why each one earns its own test case 1. **The user-visible outcome differs.** A busy provider deserves "try again in a moment"; an outage deserves an alternative route or an honest stop; an oversized document deserves a message about size that tells the person what to change. One shared branch means one shared message, and one shared message is wrong for at least three of the four. 2. **The remedy differs.** Only two of the four are worth a second attempt, and one of them must never be attempted twice at any budget. A single branch tends to acquire a single retry rule, which is then either too eager or too timid. 3. **A catch-all case passes for the wrong reason.** If the assertion is "an error was handled", it is satisfied by any handling at all, including handling that renders an empty panel. The case then survives every change that quietly breaks the path. 4. **The failures arrive separately in production.** Volume refusals cluster at peak, outages arrive as an incident, oversized inputs arrive with one particular customer's documents. Cases named after the failures make an incident readable — "the size path is the one that regressed" — instead of "the error case failed". ## Where the boundary sits Two nearby things are not this subject. An answer that arrives and is thin, wrong or off-tone is judged against whatever the feature promised about its output; that is a quality question with different cases and different assertions. A decline the product deliberately designs — the feature choosing not to answer something it should not answer — is intended behaviour rather than a delivery failure. This material is only about the step failing to produce anything to judge at all. ## A practical starting list - Name the four cases after the failures, not after the code path they happen to share. - Fix the feature's own waiting budget in the case rather than trusting an inherited default. - Measure size locally before spending a call, and assert that no call was made. - Assert something positive in each case: the specific message, the working alternative, the created work item. - Record which of the four occurred, so production can tell you which is actually common. It is rarely the one the team expected.

  • Which of these failures can the feature detect before it calls the generative step at all?
    Oversized material, in most cases. The feature can measure what it is about to send against the accepted size and refuse locally, with a message about size, instead of spending a round trip to be told. The other three are only observable by making the call, though a very recent outage can be remembered briefly so the feature stops hammering a provider it just found unavailable.
  • Why should the test case pin the feature's own deadline rather than an inherited one?
    Because the person's patience is the constraint the product owns. A feature that sets no deadline of its own inherits whatever the underlying connection allows, usually far longer than anyone will wait, and the degraded path then never runs — the request simply hangs. A case that fixes the feature's budget, say four seconds, and asserts a fallback view inside it, is asserting a product promise rather than an inherited default.
  • How is a delivery failure different from an answer that arrives but is unusable?
    A delivery failure means nothing came back to judge, so the feature's own behaviour is the whole subject: what the person sees and what they can still do. An answer that arrives and is poor is judged against what the feature promised about its output, which is a different question with different cases. Keeping them apart stops a suite from reporting a quality problem as an outage.

A parcel that never arrives, one refused at a busy counter, a closed depot, and a box too big for the slot all leave the doorstep empty — but only one of them is fixed by calling back this afternoon.

saying these in an interview costs you the question

  • Treats every failure of the step as one branch
  • Believes trying again will eventually fix an oversized input
  • Leaves the wait unbounded and calls it patience
  • Tests only the path where an answer arrives
  • Shows the same message for a busy provider and an outage
open as a page

For a feature whose answer text comes from a generative step, what makes an acceptance criterion adjudicable?

level: middleimportance: must knowfreq 64%

basics

~20 s

An adjudicable acceptance criterion names one observable property of the response, says where to look for it, and fixes the decision rule in advance, so two testers reading the same response record the same pass or fail.

open as a page

Before writing test cases for an AI-powered feature, why agree how often each acceptance criterion may miss?

level: middleimportance: must knowfreq 57%

basics

~20 s

A criterion with no stated allowance is read as absolute, so the first response that misses it forces an unplanned argument. Agreeing the allowance up front makes a rare miss an expected, recorded outcome rather than a crisis.

open as a page

When a test of a feature that answers from a generative step fails and will not reproduce, what must that test run itself have recorded?

level: middleimportance: must knowfreq 62%

basics

~20 s

Write the evidence at failure time, because a repeat may never reproduce it: the exact input, the material and instruction text given to the generative step, the model identifier and generation settings, the raw output, and a correlation id.

open as a page

Which layers can own a wrong answer from an AI-powered product feature, and how do you decide which one it is?

level: middleimportance: must knowfreq 62%

basics

~20 s

Five layers can own it: the ordinary product code around the generative step, the instruction text sent to it, the material supplied with it, the configuration, and the generative model's own capability ceiling. Attribute only after evidence separates them.

open as a page

In a product feature whose answer text comes from a generative step, which parts can a test still assert exactly?

level: middleimportance: must knowfreq 62%

basics

~20 s

Everything except the generated wording stays exactly assertable: the dispatch path and its inputs, the response's non-text fields, citation rendering, the empty and cut-short states, and permission refusals. Map that surface first, then write looser checks over the text.

open as a page

When the generative step inside a feature returns nothing usable, what must the test case assert the user can still do?

level: middleimportance: must knowfreq 58%

basics

~20 s

Assert one of three outcomes and assert it works: a manual route the person can finish themselves, a queued promise that really arrives later, or an honest stop. A dead end that merely looks handled is a defect.

open as a page

In a product feature backed by a generative step, what can a test case put in that step's place, and what does a failing case prove?

level: middleimportance: must knowfreq 58%

basics

~20 s

Three seams exist: replay a response recorded from a real call, return an answer the test author wrote, or make the real call. The first two make a failure point at your own code; the third does not.

open as a page

For a feature whose answers come from a generative step, why does a test suite pass rate misstate readiness, and what replaces it?

level: middleimportance: must knowfreq 58%

basics

~20 s

A pass rate describes one judged sample of varying output, not the feature itself: re-judge the same cases and the number moves. Report a measured residual error rate per user action, with its sample size, judging rule and split by consequence.

open as a page

How many repeats should a test case run against a generative product feature, and what does a four-in-five pass rule claim?

level: middleimportance: must knowfreq 60%

basics

~20 s

Repeat count follows the success-rate difference you need to detect; three to five repeats separate only coarse differences. A four-in-five rule claims the feature's per-attempt success rate is high, not that any single user attempt will succeed.

open as a page

In acceptance criteria for an AI-powered feature's answer, why separate must, should and must-never?

level: juniorimportance: should knowfreq 48%

basics

~10 s

Because the three carry different obligations. A must failing means the promise was not kept and the case fails; a should failing is recorded but does not block; a must-never failing stops everything regardless.

open as a page

Why must a failing test keep the generative step's raw answer, not just the value the product parsed from it?

level: juniorimportance: should knowfreq 44%

basics

~20 s

Only the raw answer separates a bad answer from good text the product mishandled. With just the derived value, a deterministic parsing, trimming or fallback bug in the team's own code reads as variation in the generative step.

open as a page

For a feature that streams a generated answer into the page, which presentation states need exact expected results?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Four states have fixed expected results: the streamed pieces must assemble and terminate cleanly, a length ceiling must show the defined cut-short marker, nothing to answer from must show the defined empty state, and a caller without permission must be refused outright.

open as a page

Why is re-running a test case until the build turns green not a repair when the product's answer comes from a generative step?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Re-running draws another sample from the same varying feature; it changes the reported result, not the feature. Under a pass rule that already tolerates the expected variation, a repeated red is evidence the success rate is genuinely below the bar.

open as a page

The same build of an AI-powered feature answers one user correctly and another wrongly. What differences do you rule out first?

level: middleimportance: should knowfreq 45%

basics

~20 s

Rule out deterministic differences before blaming any layer: the literal inputs each user sent, the material each account's request assembled, the variant and settings served, and the client each used. Only then consider run-to-run variation.

open as a page

How do you write test cases so an eloquent generated answer that misses the product's promise still fails?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Derive every criterion from the product's stated promise before any response is read, then adjudicate each response against that list only. Fluency, extra detail and pleasing tone are not criteria, so a well-written answer that omits a promised item fails.

open as a page

Failing tests of a generative feature save real customer text as evidence. How should that evidence be redacted, kept and read?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Apply per-field redaction in the writer, so the unredacted form never reaches the store; keep enough shape that the record still explains the failure; set and enforce a retention window; and default read access to the people investigating that feature.

open as a page

A failing test saved the generative step's input and output but nothing about the call itself. What can an investigation not conclude?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Without the supplied material, the instruction version, the model identifier and the generation settings, an investigation cannot say which wording or model produced the failure, whether a repeat is the same call, or whether a later edit fixed anything.

open as a page

An AI-powered feature returned a wrong answer. Which cheap experiments separate the possible causes, and in what order do you run them?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Run the cheap deterministic checks first: replay the recorded call, compare the shown answer with the returned one, and read the instructions and material exactly as sent. Only then vary one layer at a time and re-measure the failure rate.

open as a page

When no change your team can make fixes an AI-powered feature's wrong answers, what must that conclusion rest on, and what is recorded instead of a fix?

level: seniorimportance: should knowfreq 38%

basics

~20 s

It rests on evidence, not exhaustion: a measured failure rate on fixed inputs, and every cheaper layer changed and re-measured one at a time. In place of a fix, record the limit, its conditions, its containment, and when to re-check.

open as a page

In a feature where a generative step picks which internal action to run, how do you assert that dispatch exactly?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Observe the hand-off between the decision and the action: record the selected action's name and its argument values, then assert those against exact expected values. Pin the request so the case repeats, and assert nothing about the text that came back.

open as a page

A feature's degraded path runs only when its generative step fails. How do you keep it exercised as the feature changes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Make the failure cases first-class: run them on every change rather than in an occasional batch, assert something positive instead of the absence of an error, review the fallback view whenever the primary view changes, and watch how often each fallback fires in production.

open as a page

A feature retries its generative step once before falling back. What should the test case assert about that retry?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Assert the retry's price, not only its success: the total wait the person can now face, a hard ceiling on attempts, that only transient failures are retried, and that the extra calls are counted somewhere people read.

open as a page

A test suite replays one recorded response from its product's generative step: how do you keep that recording honest?

level: seniorimportance: should knowfreq 37%

basics

~20 s

Stamp the recording with its capture date and the configuration it was captured under, and refresh whenever any of that changes. Comparing it word for word against a fresh real answer proves nothing, because the real step never repeats itself.

open as a page

For a generative step in a product, what defect does each test seam - replay, script, or live call - let through?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Replay hides any change in what the real step now returns, including a shape the parser cannot handle. An authored answer hides everything the real step actually does. A live call hides shapes it happens not to produce that day.

open as a page

Not every wrong answer from a generative product feature costs the same. How do you separate cheap misses from expensive ones?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Classify each wrong answer by what it costs rather than how often it happens: whether the user notices at the time, whether it can be undone, and whether the feature acted or only proposed. Then set a separate tolerance for each class.

open as a page

You sign off a feature whose answers come from a generative step and will sometimes be wrong. What observation reverses that decision, and who watches?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Wrong answers are forecast, so one bad output reverses nothing. Write a rate on an observable signal — edits, discards, repeat attempts, escalations — with a review level, a revert level, a time window, a named person watching, and a fallback that already exists.

open as a page

In a generative product feature, which evidence shows a varying test result comes from the feature rather than the test?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Pin the generative step to one fixed recorded answer and repeat: if the result still moves, the test is the source. Then read which assertion moved - generated wording varies by design, a stored identifier does not.

open as a page

Calling the real generative step is slow, costly and sometimes fails for reasons outside your team - what seam policy do you set?

level: principalimportance: should knowfreq 40%

basics

~20 s

Keep the merge gate deterministic and put the rich live cases on a schedule and before a release, with a named owner. Add one minimal real call to the gate so a total outage of the step cannot pass unnoticed.

open as a page

A product feature's generated answers are wrong in 4% of attempts and will not reach zero. How do you make the ship call?

level: principalimportance: should knowfreq 42%

basics

~20 s

Compare against the process the feature replaces, not against zero. Establish whether the figure is trustworthy, what each wrong answer costs, whether errors are detectable by the user and by the team, and what reverting takes. Then record the number, its conditions and who accepted it.

open as a page

showing 1–30 of 32