skip to content

When a team proposes moving a new framework onto a tech radar's 'Trial' ring for a pilot project, what exit criteria should be defined up front so the trial doesn't drag on indefinitely, and what typically happens if the criteria aren't met?

level: middleimportance: should knowfreq 55%

answer

  1. time-boxed trial
  2. bounded blast radius pilot
  3. quantifiable exit criteria set upfront
  4. promote/demote/extend outcomes
  5. prevents Trial limbo/shelfware

basics

~20 s

Before starting a trial, agree on a deadline and clear pass/fail signals — like production stability for a set period, a cap on incidents, and team feedback above a bar. If those aren't met by the deadline, the trial ends and the technology moves back out (to Assess or Hold) instead of just lingering.

solid answer

~50 s

A well-run Trial needs three things set before it starts: a time box (e.g., one quarter or one release cycle), the specific system it will run on (real, but with a bounded blast radius), and measurable exit criteria — things like incident count, on-call burden, performance against a target, developer ramp-up time, and a written retro. At the deadline, the reviewing body checks the evidence against those criteria: a clear pass moves it to Adopt (or extends to a second trial team for broader validation); a clear fail moves it to Hold with rationale documented so the question doesn't get re-litigated every quarter; ambiguous results extend the trial with tightened criteria rather than defaulting to Adopt by inertia. Without this upfront agreement, trials tend to become permanent by default — the pilot team keeps using it, nobody revisits the ring, and the radar loses credibility as a living document.

go deeper

for a junior

Understands a trial should have some kind of endpoint, but may not have designed exit criteria or run a review themselves.

for a middle

Can propose concrete, measurable exit criteria and describe the promote/demote/extend outcomes, likely from having participated in a trial review.

for a senior

Has actually set exit criteria and pilot scope for a real trial, including choosing a pilot system with appropriately bounded blast radius, and can discuss trade-offs when criteria are ambiguous.

for a principal

Designs the org-wide review cadence and governance discipline that keeps trials from stalling across many concurrent proposals, including how to handle politically uncomfortable demote decisions.

## Why Trial turns permanent on its own The core problem this practice solves is that 'Trial' can silently turn into a permanent, ungoverned state if nobody defines in advance what success or failure looks like. Left unstructured, a pilot team simply keeps shipping features on the new framework because it works well enough for them, nobody schedules a review, and eighteen months later the organization has a de-facto second standard technology that was never formally promoted, never got the operational investment (training material, shared libraries, on-call runbooks) a true Adopt-ring choice would receive, and never got the scrutiny that would have caught it if it were actually a poor fit. Defining exit criteria up front converts an open-ended 'let's see how it goes' into a decision with a deadline. ## The four things agreed before code is written Mechanism, concretely: before code is written, the proposing team and the reviewing body (architecture guild, principal engineer group) agree on four things. 1. **First, a time box** — commonly one quarter or one release cycle, long enough to hit a real production incident or two, short enough that the review doesn't get indefinitely postponed. 2. **Second, the specific system the trial will run on**, chosen deliberately for a bounded blast radius: an internal tool, a new low-traffic service, a batch job with no customer-facing SLA — never the checkout flow. 3. **Third, quantifiable exit criteria**: examples include a cap on production incidents attributable to the framework, latency within a defined margin of the current default technology, a developer with no prior exposure shipping a feature within a target number of days, and the on-call engineer diagnosing a failure using existing observability tooling without framework-specific expertise. 4. **Fourth, an owner and a calendar date for the review meeting**, so revisiting it isn't left to chance. ## The three outcomes at the deadline At the deadline, the evidence is compared against the criteria and one of three outcomes follows: - **Promote to Adopt** (sometimes gated on a second, independent team also trialing it, to rule out one-team bias). - **Demote to Hold**, with the rationale written down so the same pitch doesn't get re-litigated from scratch next quarter by someone unaware it was already tried. - **Extend the trial** with tightened criteria, used sparingly, because open-ended extension is exactly the failure mode this process exists to prevent. ## What the discipline protects Why this matters, beyond just tidiness: a radar's value comes entirely from being a trustworthy, current signal. If Trial items never get resolved, the radar stops answering the question people actually consult it for — 'is this safe to build on for my next project?' — because a stalled Trial item gives no real signal either way. It also protects the trialing team: without a documented exit review, a team that picked a framework which quietly underperformed has no forcing function to admit it and migrate off before the sunk cost grows large. ## The cost of rigid criteria **Trade-offs:** defining rigid criteria up front costs some flexibility — sometimes a technology's real value only becomes clear after conditions the criteria didn't anticipate (say, a scaling event the trial's low-traffic pilot never triggered), and a team can feel unfairly graded against a benchmark set before they understood the technology. There's also a genuine cost to running the review discipline itself: someone has to chase the retro write-up, schedule the review meeting, and make the promote/demote call even when it's politically uncomfortable (telling a team their months-long pilot didn't meet the bar). The alternative — no criteria, informal 'we'll know it when we see it' — is cheaper up front but is exactly what produces permanent, ungoverned Trial limbo. ## Failure modes in production - **Criteria that are vague enough to be unfalsifiable** ('the team likes it') — the most common. Nothing forces a hard decision, so the trial just continues by default. - **A review meeting that gets scheduled but perpetually rescheduled** — a second. Delivery pressure eats into calendar time, until the trial has effectively become permanent Adopt-by-neglect without ever passing a real review. - **Choosing a pilot system with too small a blast radius to actually generate meaningful evidence** — a third. A toy internal tool that never sees real load tells you nothing about how the technology behaves under a production incident, so the 'evidence' at review time is thin and the promote/demote decision ends up being made on vibes anyway. ## Putting it together A concrete instance: a team wanting to trial a new event-streaming platform for asynchronous processing might scope the trial to a single non-critical notification pipeline, agree on a one-quarter window, set exit criteria around message-delivery latency and operator runbook completeness, and pre-schedule the architecture guild review — so that regardless of outcome, the organization walks away in one quarter with either a promoted Adopt-ring default backed by evidence, or a documented Hold decision, rather than an open-ended experiment nobody remembers agreeing to.

  • Why is it important to require a second independent team's trial before promoting a technology to Adopt, rather than relying on the pilot team's own results?
    A single team's positive experience can reflect that team's specific skills, use case, or tolerance for rough edges rather than the technology's general fit — requiring a second, independent trial checks for that bias and validates the technology generalizes beyond the original champions before the whole org is asked to standardize on it.
  • What should happen if a trial's exit-criteria review lands in ambiguous territory — some criteria met, some not?
    The reviewing body should extend the trial with explicitly tightened, narrower criteria targeting exactly what was ambiguous, rather than defaulting to either promotion or demotion by inertia — and this extension should itself carry a new deadline so it doesn't become the same open-ended limbo the original criteria were meant to prevent.
  • How should the choice of pilot system for a trial balance getting real evidence against limiting blast radius?
    Pick a system real enough to generate genuine production signal — real users, real on-call, some meaningful load — but non-critical enough that a failure doesn't harm customers or revenue; an internal admin tool, a new low-traffic service, or a batch job with no tight SLA are typical choices, while the customer checkout path or anything handling regulated data is not.

Like a probationary hire with a defined 90-day review and specific performance goals set on day one — without that upfront agreement, 'probation' quietly turns into permanent employment by default, with nobody ever formally deciding it worked out.

saying these in an interview costs you the question

  • Says a trial should run 'until we're confident' with no defined deadline
  • Can't name a single measurable (not just anecdotal) exit criterion
  • Thinks one team's positive experience alone is sufficient to promote to Adopt
  • Has no answer for what happens when criteria aren't met — implies it just quietly continues
  • Would pilot a brand-new technology directly on a customer-critical system

context