skip to content

Case Design

What a tester can write a case against when output is not fixed: properties an answer must hold, the parts still assertable exactly, and the degraded path. 'Looks right' is not a criterion.

on this pageshow

questions

12

A feature's answer text comes from a generative step. In what distinct ways can that step fail to deliver an answer?

level: juniorimportance: must knowfreq 56%

answer

  1. Failing to speak is not one event
  2. Four shapes, not one error branch
  3. Deadline, volume refusal, outage, oversized input
  4. Only some are worth a second attempt
  5. One catch-all case proves almost nothing

basics

~20 s

Four delivery failures matter: nothing comes back inside the deadline, the call is refused for volume, the provider is unavailable, or the input exceeds what the step accepts. Each strands the user differently, so each earns its own test case.

solid answer

~40 s

Sort them by what the caller observes. **No answer inside the deadline** — the feature's own waiting budget expired while the call was still open, so it must decide without an answer. **Refused for volume** — the provider declined because too many requests arrived, and the same request would likely succeed shortly after. **Unavailable** — the provider is not answering at all, and trying again inside the same interaction rarely helps. **Input too large** — the request is rejected before any generation happens, deterministically, so waiting never fixes it; the feature must shorten or split the material, or say so plainly. A single "it errored" test case proves an error was caught, not that any of these four leaves the user somewhere sensible.

code

pseudocode · 20 lines
pseudocode
CASES for a feature whose text comes from a generative step:

no_answer_inside_deadline:
  given the step returns nothing within budget_ms
  expect fallback panel shown, entered text still present

refused_for_volume:
  given the step declines the call as over the allowance
  expect a retry control, enabled after the suggested pause

provider_unavailable:
  given the call to the step cannot complete at all
  expect fallback panel, manual route offered, no retry loop

input_too_large:
  given material exceeds the accepted size
  expect a size message, and assert no call was made

anti_case:
  one shared "it errored" branch asserting only "no crash"

go deeper

for a junior

Be ready to name the failures without prompting: no answer inside the deadline, a refusal for volume, an unavailable provider, and material larger than the step accepts. Naming the four is the whole ask at this level.

for a middle

Explain what separates them mechanically — which are transient, which are deterministic, and which the feature can see coming before it spends a call. An interviewer expects you to say which of the four a second attempt could plausibly fix.

for a senior

Show how the four become four cases with four different assertions, and how you keep the feature's own waiting budget shorter than anything it depends on. Talk about what production tells you afterwards: how often each of the four actually fires.

for a principal

Own the policy question of how much degradation the product will pay for. Handling all four properly costs design work and spend on extra calls, and deciding which failures deserve a queued promise and which deserve an honest apology is a product tradeoff, not an implementation detail.

A feature that gets part of its output from a generative step depends on something that can be slow, busy, absent, or unwilling to accept the request at all. Those are four different events, they strand the user in four different ways, and a suite that folds them into a single "the step errored" case proves only that an error was caught somewhere. ## The four ways it fails to deliver **No answer inside the deadline.** The feature set a budget — the longest it is willing to make someone wait — and that budget expired while the call was still open. The budget has to be the feature's own: a feature that sets none inherits whatever the underlying connection allows, which is usually far longer than any person will sit and watch. The distinguishing property of this failure is that the work may still be happening. It may finish after the feature has given up, which matters if the work is chargeable or has already changed something. **Refused for volume.** The provider took the connection and declined the work because too many requests arrived in some window — from this feature, from the rest of the product, or from everyone sharing an allowance. The distinguishing property is that the identical request would probably succeed a short time later, and the refusal frequently carries a hint of how long to wait. This is the only one of the four where waiting is a real remedy. **Unavailable.** The provider is not answering, or is answering with a hard failure that has nothing to do with this particular request: a dropped connection, an expired credential, a dependency of the dependency that is down. The distinguishing property is duration — it typically outlasts one interaction, so trying again inside the same session is close to useless and the feature has to degrade properly. **Input too large.** The request was rejected before any generation began, because the material exceeds what the step accepts. The distinguishing property is that it is deterministic: the same input is rejected the same way forever. The feature has to shorten, split or summarise the material, or say plainly that this document is too big — and it can usually detect the problem itself before spending a call at all. | Failure | Would a second attempt help? | What the feature owes the user | | --- | --- | --- | | No answer inside the deadline | Sometimes | A bounded wait, then an honest fallback | | Refused for volume | Yes, after a pause | A pause that respects any suggested wait | | Unavailable | Rarely within one session | A route that does not need the step | | Input too large | Never | A specific message about size | ## Why each one earns its own test case 1. **The user-visible outcome differs.** A busy provider deserves "try again in a moment"; an outage deserves an alternative route or an honest stop; an oversized document deserves a message about size that tells the person what to change. One shared branch means one shared message, and one shared message is wrong for at least three of the four. 2. **The remedy differs.** Only two of the four are worth a second attempt, and one of them must never be attempted twice at any budget. A single branch tends to acquire a single retry rule, which is then either too eager or too timid. 3. **A catch-all case passes for the wrong reason.** If the assertion is "an error was handled", it is satisfied by any handling at all, including handling that renders an empty panel. The case then survives every change that quietly breaks the path. 4. **The failures arrive separately in production.** Volume refusals cluster at peak, outages arrive as an incident, oversized inputs arrive with one particular customer's documents. Cases named after the failures make an incident readable — "the size path is the one that regressed" — instead of "the error case failed". ## Where the boundary sits Two nearby things are not this subject. An answer that arrives and is thin, wrong or off-tone is judged against whatever the feature promised about its output; that is a quality question with different cases and different assertions. A decline the product deliberately designs — the feature choosing not to answer something it should not answer — is intended behaviour rather than a delivery failure. This material is only about the step failing to produce anything to judge at all. ## A practical starting list - Name the four cases after the failures, not after the code path they happen to share. - Fix the feature's own waiting budget in the case rather than trusting an inherited default. - Measure size locally before spending a call, and assert that no call was made. - Assert something positive in each case: the specific message, the working alternative, the created work item. - Record which of the four occurred, so production can tell you which is actually common. It is rarely the one the team expected.

  • Which of these failures can the feature detect before it calls the generative step at all?
    Oversized material, in most cases. The feature can measure what it is about to send against the accepted size and refuse locally, with a message about size, instead of spending a round trip to be told. The other three are only observable by making the call, though a very recent outage can be remembered briefly so the feature stops hammering a provider it just found unavailable.
  • Why should the test case pin the feature's own deadline rather than an inherited one?
    Because the person's patience is the constraint the product owns. A feature that sets no deadline of its own inherits whatever the underlying connection allows, usually far longer than anyone will wait, and the degraded path then never runs — the request simply hangs. A case that fixes the feature's budget, say four seconds, and asserts a fallback view inside it, is asserting a product promise rather than an inherited default.
  • How is a delivery failure different from an answer that arrives but is unusable?
    A delivery failure means nothing came back to judge, so the feature's own behaviour is the whole subject: what the person sees and what they can still do. An answer that arrives and is poor is judged against what the feature promised about its output, which is a different question with different cases. Keeping them apart stops a suite from reporting a quality problem as an outage.

A parcel that never arrives, one refused at a busy counter, a closed depot, and a box too big for the slot all leave the doorstep empty — but only one of them is fixed by calling back this afternoon.

saying these in an interview costs you the question

  • Treats every failure of the step as one branch
  • Believes trying again will eventually fix an oversized input
  • Leaves the wait unbounded and calls it patience
  • Tests only the path where an answer arrives
  • Shows the same message for a busy provider and an outage
open as a page

For a feature whose answer text comes from a generative step, what makes an acceptance criterion adjudicable?

level: middleimportance: must knowfreq 64%

basics

~20 s

An adjudicable acceptance criterion names one observable property of the response, says where to look for it, and fixes the decision rule in advance, so two testers reading the same response record the same pass or fail.

open as a page

Before writing test cases for an AI-powered feature, why agree how often each acceptance criterion may miss?

level: middleimportance: must knowfreq 57%

basics

~20 s

A criterion with no stated allowance is read as absolute, so the first response that misses it forces an unplanned argument. Agreeing the allowance up front makes a rare miss an expected, recorded outcome rather than a crisis.

open as a page

In a product feature whose answer text comes from a generative step, which parts can a test still assert exactly?

level: middleimportance: must knowfreq 62%

basics

~20 s

Everything except the generated wording stays exactly assertable: the dispatch path and its inputs, the response's non-text fields, citation rendering, the empty and cut-short states, and permission refusals. Map that surface first, then write looser checks over the text.

open as a page

When the generative step inside a feature returns nothing usable, what must the test case assert the user can still do?

level: middleimportance: must knowfreq 58%

basics

~20 s

Assert one of three outcomes and assert it works: a manual route the person can finish themselves, a queued promise that really arrives later, or an honest stop. A dead end that merely looks handled is a defect.

open as a page

In acceptance criteria for an AI-powered feature's answer, why separate must, should and must-never?

level: juniorimportance: should knowfreq 48%

basics

~10 s

Because the three carry different obligations. A must failing means the promise was not kept and the case fails; a should failing is recorded but does not block; a must-never failing stops everything regardless.

open as a page

For a feature that streams a generated answer into the page, which presentation states need exact expected results?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Four states have fixed expected results: the streamed pieces must assemble and terminate cleanly, a length ceiling must show the defined cut-short marker, nothing to answer from must show the defined empty state, and a caller without permission must be refused outright.

open as a page

How do you write test cases so an eloquent generated answer that misses the product's promise still fails?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Derive every criterion from the product's stated promise before any response is read, then adjudicate each response against that list only. Fluency, extra detail and pleasing tone are not criteria, so a well-written answer that omits a promised item fails.

open as a page

In a feature where a generative step picks which internal action to run, how do you assert that dispatch exactly?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Observe the hand-off between the decision and the action: record the selected action's name and its argument values, then assert those against exact expected values. Pin the request so the case repeats, and assert nothing about the text that came back.

open as a page

A feature's degraded path runs only when its generative step fails. How do you keep it exercised as the feature changes?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Make the failure cases first-class: run them on every change rather than in an occasional batch, assert something positive instead of the absence of an error, review the fallback view whenever the primary view changes, and watch how often each fallback fires in production.

open as a page

A feature retries its generative step once before falling back. What should the test case assert about that retry?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Assert the retry's price, not only its success: the total wait the person can now face, a hard ceiling on attempts, that only transient failures are retried, and that the extra calls are counted somewhere people read.

open as a page

When a feature's generative step is swapped and its wording shifts, which test cases should fail and which must not?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Nothing on the exact-assertion surface should move: dispatch, response fields, citation rendering, states and permissions keep passing. Only the wording-reading checks may fail, and their output must name the text as what changed rather than the surrounding behaviour.

open as a page