skip to content

Calling the real generative step is slow, costly and sometimes fails for reasons outside your team - what seam policy do you set?

level: principalimportance: should knowfreq 40%

answer

  1. Two defensible positions, one measurable input
  2. Outside-caused reds destroy a gate
  3. Tier the suite, own the schedule
  4. The false-red rate decides the line

basics

~20 s

Keep the merge gate deterministic and put the rich live cases on a schedule and before a release, with a named owner. Add one minimal real call to the gate so a total outage of the step cannot pass unnoticed.

solid answer

~50 s

Split the suite into tiers and decide per tier what may block a merge. Everything that gates a merge uses a deterministic seam - authored answers and replayed recordings - because a check that goes red for reasons the team cannot fix stops being read within weeks. The rich live cases run on a schedule against a deployed environment and again before a release, with a named owner and a response time for their failures. Keep one exception inside the gate: a single real call asserting only that the step was reachable and answered in the expected shape, which closes the hole where a fully deterministic suite is green through a total outage. Then justify the split with numbers - the live set's rate of red results caused outside the team, cost per pipeline run, and a date the policy is reviewed.

code

pseudocode · 11 lines
pseudocode
policy "seams for the generative step":

  merge gate     : authored + replayed cases only, plus one reachability
                   call asserting only that the step answered in shape
  scheduled      : the live set (12 cases), nightly, deployed environment
                   red -> ticket owned by the feature team, never a merge block
  before release : live set must have run green within 24 hours of the tag

  evidence kept  : share of live-set reds caused outside the team,
                   cost per pipeline run, age of every recording,
                   date this policy was last reviewed

go deeper

for a junior

Know that calling a real generative step inside a pipeline costs money and time, and that it can fail for reasons that have nothing to do with the change being reviewed.

for a middle

Explain why a check that goes red for outside reasons stops being read, and describe a suite arranged in tiers - deterministic cases blocking merges, real calls running elsewhere - rather than one that treats every case the same.

for a senior

Be ready to defend a concrete arrangement: which cases block a merge, which run on a schedule, what happens when the scheduled set goes red, and the measured failure rate that made you draw the line where you did.

for a principal

Own the trade-off out loud - a deterministic gate that can be green through a total outage, against a real gate nobody trusts - and bring the numbers that pick between them plus the date the choice gets reviewed.

## The decision, stated honestly The real call is the only seam that exercises the real behaviour of a product's generative step. It is also slow, it costs money on every pipeline run, and it goes red for reasons the team cannot fix: the provider refusing calls for volume, a bad network, a step that is simply slower today. Both of those facts are true at once, and a lead has to choose an arrangement rather than wish the tension away. Two positions are defensible. **Position A - no real calls in the merge gate.** Every case that can block a merge uses a deterministic seam. Real calls live in a separate, named set that runs on a schedule and before a release, and its failures go to an owner rather than to whoever pushed last. The argument: a gate that goes red for reasons outside the team's control stops meaning anything within about two weeks, and a pipeline nobody believes is worse than no pipeline. **Position B - a thin live set inside the gate.** A handful of live cases block a merge, because a suite with no live case at all can be entirely green while the step is unreachable, misconfigured, or refusing every request. That is a total product outage a deterministic suite is structurally incapable of noticing. The argument: some floor of reality has to be enforced where it is still cheap to act on, and a team that never feels the real call stops designing for it. ## What picks between them - **How often the live set goes red for outside reasons.** This is a measurable rate, not a feeling. Above a few percent of runs, Position B corrodes the gate. - **What the product does when the step is unavailable.** A feature with a well-exercised alternative path can afford to learn about an outage from monitoring; one with no alternative path cannot. - **Cost per pipeline run multiplied by runs per day.** A number, not an instinct. Some teams run the pipeline forty times a day. - **How fast the step changes underneath.** A step whose behaviour moves monthly needs its real calls more often and closer to the release decision. - **Whether a deployed environment exists** to run the scheduled set against. If it does not, the pipeline is the only place a real call can happen, and the arrangement is forced. - **The blast radius of shipping the break.** An internal tool and a customer-facing flow do not deserve the same ceremony. ## Where most teams land | Tier | Seam | Blocks a merge | Runs when | |---|---|---|---| | Most of the suite | Authored answers | Yes | Every pipeline run | | One shaped path per feature | Replayed recording | Yes | Every pipeline run | | The live set | Real calls, full checks | No | On a schedule, and before a release | | Reachability check | One real call, minimal assertion | Yes | Every pipeline run | The fourth row is the compromise that gives both arguments most of what they want. A single live case that asserts almost nothing - the step was reachable and returned something of the expected shape - costs one call, almost never goes red for content reasons, and closes precisely the hole Position B worries about. The rich live cases stay out of the gate. ## The evidence the arrangement needs A policy is not a paragraph in a document. It is a paragraph plus the numbers that justify it, kept current. 1. **The live set's rate of red results caused outside the team.** This is the number that decides whether it may block a merge, and it has to be measured rather than asserted. 2. **Cost per pipeline run and runs per day**, so the bill is a figure a lead can trade against everything else. 3. **Proof that the deterministic seams still reach every branch of the team's handling.** Otherwise "the gate is deterministic" is hiding a gate that tests very little. 4. **A maximum age between the last green live run and a release.** The live set is allowed out of the merge gate only if it is inside the release decision. 5. **A named owner and a response time for the scheduled set.** An unowned scheduled set is decoration: red for a month, and nobody surprised. 6. **A review date.** Every input above moves - the price, the stability, the release cadence - so the policy is reviewed on a schedule rather than when it finally hurts. ## How the arrangement fails It fails when the live set is quietly made green: assertions loosened until nothing can fail them, or the set skipped whenever the pipeline is busy. Both keep the dashboard tidy and destroy the only evidence the whole arrangement was resting on. If the live set is not worth reading, delete it and say so. An honest Position A is defensible; a decorative live set is not.

  • Which single number tells you the live set must not be allowed to block a merge?
    The share of its red results caused by something outside the team - the provider refusing calls, a network fault, a slow step. Once that share is more than a few percent of runs, the team starts re-running and ignoring rather than reading, and the gate has stopped being a gate. Measure it before arguing about it.
  • A scheduled live set has been red for three weeks and nobody noticed. What is missing from the policy?
    A named owner and a stated response time. Moving cases out of the merge gate only works if something else makes their failures land on a person; otherwise the tier is decoration and the arrangement's whole justification - that reality is checked somewhere - is false. The other missing piece is a rule that a release cannot go out behind a stale live result.

saying these in an interview costs you the question

  • Puts the whole live set in the merge gate and lives with the noise
  • Removes real calls entirely and cannot notice a step-wide outage
  • Sets the policy on instinct with no measured failure rate or cost
  • Leaves the scheduled live set unowned, so its failures go unread
  • Loosens live assertions until nothing can fail and calls that evidence