skip to content

What does operational acceptance testing check before a system is accepted into live running?

level: seniorimportance: should knowfreq 36%

answer

  1. Accepted by whoever will run it
  2. Criteria about operating, not features
  3. Prove it by doing it, not by having it
  4. The least-exercised path is the recovery path
  5. Rehearse the half-finished failure

basics

~20 s

Operational acceptance testing checks that the team who will run the system can run it: install and upgrade, backup and restore, failover, rollback, monitoring and alerting, access management and the written procedures -- each rehearsed rather than assumed.

solid answer

~50 s

It is acceptance by the people who will operate the system rather than by the people who asked for the feature, and its criteria are operability criteria. The core list: install and upgrade on a like-for-like setup; restore from a real backup, not merely proof that a backup file exists; failover and recovery; **rollback rehearsed under partial failure**, because that path is the least-exercised code in any system; alerts that fire on injected faults and point at an action; access and secret rotation; capacity and latency budgets stated as percentiles; and the runbooks driven by someone who did not write them. The output is evidence -- what was exercised, on which build, with what result -- not an assurance. Skipping it is how organisations discover at 3 a.m. that the documented rollback was never run and leaves half the data behind.

go deeper

for a junior

Know that acceptance is not only about features: before go-live someone checks that the system can be installed, backed up and restored, monitored and recovered, and that the people who will run it agree they can.

for a middle

Be able to list the operability criteria and explain why each must be demonstrated rather than asserted -- especially why a written backup proves nothing until a restore has been performed on a clean instance.

for a senior

Show production judgement: rehearse recovery and rollback under partial failure, inject faults to confirm alerts fire and are actionable, time each procedure, and state honestly how the rehearsal setup differs from the live one.

for a principal

Own the standard: which operability criteria are mandatory before any system is accepted, who is entitled to refuse, how often recovery is re-rehearsed as the system changes, and how that cost is funded against feature delivery.

### Acceptance by the people who will hold the pager Functional acceptance asks whether the system does what the business asked. **Operational acceptance asks whether the people who will run it can run it.** The signatories are the operations or platform group, the criteria are operability criteria, and the level exists because a system can meet every functional criterion and still be unrunnable in practice. ### The checklist worth being able to recite **Install and upgrade.** Deploy from nothing on a setup matching the live one, then upgrade from the currently running version -- upgrade paths break in ways clean installs never show, particularly around data migration. **Backup and restore.** Testing that a backup was written is testing nothing. The criterion is a *restore*: take yesterday's artefact, restore it onto a clean instance, and demonstrate the data is complete and usable. Untested backups fail at the moment they are needed, and the failure is unrecoverable. **Failover and recovery.** Kill an instance, sever a dependency, and watch what the system does. Note the recovery time and compare it against whatever the business was told. **Rollback, rehearsed under partial failure.** This is the item people skip and the item that hurts. A rollback path is written once and executed almost never, so it rots silently. Rehearse it *while something is half done*. In a document e-signing rollout: begin the upgrade, let 3 of 41 in-flight signing rounds land with the archive write committed and the audit record not yet written, then roll back. Does the rollback restore a consistent state, or leave documents archived with no audit trail -- signed in one store and unknown in the other? Does it reject the operation, or complete silently while corrupting the reconciliation? You want the answer in a rehearsal, not during an incident. **Monitoring, alerting and diagnosability.** Inject a fault and check that an alert fires, reaches the right rota, and names an action. An alert with no runbook entry trains people to close it. Also verify log volume and retention -- a system that quintuples log output on a bad day and fills the disk has turned a small incident into an outage. **Capacity and performance budgets.** State them as percentiles against a defined load: the signing round completes within 4.3 seconds at the **92nd percentile** at 180 concurrent rounds. Averages hide the tail that generates the complaints, so a criterion phrased as an average is barely a criterion at all. **Access, secrets and expiry.** Who can reach production, how access is granted and revoked, how secrets rotate, and what expires -- certificates, tokens, licences. A certificate that lapses 40 days after go-live is an operational defect that was fully visible before release. **Procedures.** Have the runbooks driven by someone who did not write them, at the pace of a real incident. Every assumed step the author left out is discovered in the rehearsal instead of during the outage. **Data lifecycle.** Retention, archival and deletion. In a signing product this is often a hard external obligation, and it is operational work rather than feature work, so it falls between the two stools unless it is an explicit criterion. ### Why it is a separate acceptance Three reasons. The signatory is different -- the run team, not the sponsor, and they can legitimately refuse a system that works perfectly. The criteria are different in kind: not "the copy reaches every signer" but "a restore completes within the stated window". And the failures it finds are invisible to functional testing, because they only appear when something is broken, being upgraded, being rolled back or being recovered -- states functional runs deliberately avoid. ### Doing it credibly Rehearse on a setup that resembles the live one in the ways that matter -- data volume, network topology, identity integration, storage class -- and say plainly which ways it does not, since every rehearsal is an approximation and the honest statement of the gap is part of the evidence. Prefer injected faults to imagined ones. Time each procedure, because "we can restore" and "we can restore in 20 minutes" are different promises. And produce the same artefact as any other acceptance: what was exercised, on which build, observed result, who witnessed it, when. ### The one-sentence answer "Operational acceptance is the run team's acceptance: install, upgrade, restore, failover, rollback under partial failure, alerting, access and the runbooks -- each rehearsed against stated operability criteria and recorded as evidence, not asserted."

  • Why insist on rehearsing rollback while an operation is half-completed rather than from a clean state?
    Because a clean rollback is the case the author already imagined. The damage comes from partial states: a record committed in one store and not the other, an in-flight operation interrupted mid-sequence. Rolling back from there is where inconsistency, silent data loss and unresolvable reconciliations appear, and the rollback path is written once and almost never run, so nothing else keeps it honest.
  • How do you acceptance-test a written operating procedure?
    Have someone who did not write it drive it end to end, at the pace of a real incident, with the author present only to observe and forbidden from prompting. Every step the author assumed, every missing credential and every stale path surfaces immediately. Record the elapsed time as well, because a procedure that works but takes three hours is a different promise from one that takes twenty minutes.
  • Which operational acceptance criteria do teams most often forget?
    Restore, as opposed to backup. Then log volume and retention under a bad day, expiry of certificates, tokens and licences, secret rotation, access revocation, and data retention and deletion obligations. They are forgotten because they belong to neither the feature backlog nor the functional criteria, so unless someone writes them as operability criteria they are discovered in production.
  • How should a capacity criterion be phrased so it can be accepted or rejected?
    As a percentile against a stated load and operation: the signing round completes within 4.3 seconds at the 92nd percentile with 180 concurrent rounds. That names what is measured, under what conditions, and where the bar is. An average conceals the slow tail users complain about, and a bare number with no load figure cannot be reproduced by anyone else.

saying these in an interview costs you the question

  • Assumes the rollback works because the procedure exists
  • Verifies backups are written but never performs a restore
  • Treats monitoring and alerting as outside acceptance scope
  • Rehearses a runbook with its author driving and prompting
  • States a capacity criterion as an average rather than a percentile
  • Skips the upgrade path because a clean install succeeded

context