skip to content

Quality Practices & Metrics

Quality work outside the pre-merge suite: human review, pairing, checks that run against live traffic, and the numbers a team reports. Interviewers probe what you trust beyond a green build.

on this pageshow

questions

25

In what order should you read a code change under review, and why does the order matter?

level: juniorimportance: must knowfreq 72%

answer

  1. Cheapest mistake to undo goes last
  2. Judge the approach before the details
  3. Four widening passes, not one sweep
  4. Intent, contract, edge cases, tests
  5. Ask which test would fail

basics

~20 s

Read the intent first — what problem the change claims to solve — then the contract it exposes, then the edge cases and error paths, then the tests. That order settles whether the change is right before you argue about how it is written.

solid answer

~40 s

I read in decreasing order of how expensive the mistake would be to undo. **Intent first**: the change description, the requirement behind it, and whether this is the smallest change that satisfies it — if the approach is wrong, nothing below it matters. **Contract second**: the signatures, inputs, outputs, error results and stored shapes the change exposes to other code, because those are what callers depend on and the hardest thing to retract later. **Edge cases and error paths third**: empty and boundary inputs, partial failure, retries, two callers at once. **Tests last**: do they assert the behaviour that was actually claimed, and would any of them fail if the change were subtly wrong? Cosmetic remarks come after all of that, and most of them should not be a human's job at all.

go deeper

for a junior

Be ready to describe your own reading order out loud and defend it. Interviewers mostly want to hear that you read what the change is for before you read how it is written, and that you open the tests rather than trusting a green result.

for a middle

Explain why the order is what it is: reversibility. Show that you can say which parts of a change are expensive to retract later, and give an example of a defect that only an adversarial pass over the error paths would find.

for a senior

Demonstrate that you adapt the order to risk. Say how you triage a change that is too big to read, how you decide which surfaces get a careful pass, and how you turn a vague worry into the concrete question of which test would fail if it were real.

for a principal

Own the question of what human attention should be spent on at all. Be able to argue which review findings a team should stop producing by hand, and how the reading order changes when a change crosses a boundary many teams depend on.

## Why a review needs an order at all Reviewing is reading under a budget. Attention is finite and it decays fast, so whatever you look at first gets your best judgement and whatever you look at last gets a glance. A reviewer who opens the change and reads it from the first altered line to the last spends that best judgement on whichever file the diff tool happened to sort first. A deliberate reading order spends it on the parts where being wrong costs the most. The ordering principle is *reversibility*. Rank the aspects of a change by how expensive the mistake is to undo once it is merged, and read them in that order. ## Pass one — intent Before reading any code, read what the change claims to do: the description the author wrote, the requirement or defect it points at, and the conversation that produced it. Then ask three questions. Is this the right problem to solve now? Is this approach a reasonable way to solve it? Is this the smallest change that solves it, or has unrelated work ridden along? This pass is the only one that can save the whole change. A design decision caught here costs a conversation; the same decision caught after the code has been written, tested and read line by line costs a rewrite and a demoralised author. If the change has no intelligible description, that is the first finding — a reviewer who cannot state what the change is for cannot judge whether it works. ## Pass two — the contract Next read the surface the change exposes to everything outside it: function and message signatures, parameter and return types, the shape of persisted or transmitted data, error results and what a caller is expected to do with them, defaults, and units. This is the part with the widest blast radius and the shortest half-life of correction — an awkward internal loop can be tidied next month by whoever touches it, but a wrong parameter order, a nullable that should not be, or a stored field with the wrong precision spreads to every caller and then has to be unwound everywhere at once. Read the contract as a caller who has not read the implementation. If you cannot tell from the signature and its documentation what happens on an empty input, a duplicate call, or a downstream failure, that is a finding regardless of how the body is written. ## Pass three — edge cases and error paths Only now read the body, and read it adversarially rather than sequentially. Walk the boundaries: zero, one, many; empty, maximum, and just-over-maximum; absent optional values; a collection that arrives unordered. Walk the failure paths: what happens when the dependency times out, returns a partial result, or returns success but with nothing useful in it. Walk concurrency: what happens if this runs twice at once, or if a retry replays it after the first attempt already half-succeeded. Most of what review catches that a test suite does not lives here — not defects in the paths someone thought to write a test for, but the paths nobody thought about at all. A test suite can only assert what its author imagined; a second reader is a second imagination. ## Pass four — the tests Read the tests last, and read them as evidence rather than as code. Two questions do most of the work. Do these tests assert the behaviour the change claims in pass one, or do they merely exercise it — calling the new path and asserting something incidental like "no exception was thrown"? And, for each risk you found in pass three, would any existing test fail if that risk were real? If the answer is no, the useful review comment is not "add a test" but "which test would have failed if this were wrong?" — that question names a missing oracle instead of a missing file. Some reviewers deliberately read tests first when the intent is unclear, using them as documentation of what the author believes should be true. That works, at a cost: you inherit the author's framing of what matters and are less likely to notice the case they never considered. ## What the order is not The order is a default, not a ritual, and it interacts with size. Past a few hundred changed lines a reader cannot hold the contract in mind while walking the branches, and the passes collapse into skimming; at that point the honest review comment is that the change should be split. It also assumes the mechanical layer is already handled elsewhere — formatting and other machine-checkable properties should not be consuming a human's first pass at all. What a reviewer brings is judgement about intent, contract and the unimagined case, and the reading order exists to spend that judgement where it is scarcest.

  • You have twenty minutes and the change is far too large to read properly. What do you actually do with the time?
    Spend it on the highest-risk surface rather than spreading it thinly: the contract, anything that writes or migrates stored data, anything touching money or access. Then say plainly in the review what you did and did not read, so the approval does not imply coverage it never had, and ask the author to split the remainder. An honest partial review is more useful than a full one nobody performed.
  • Should you read the tests first instead, since they describe the intended behaviour?
    Sometimes — when the change description is thin, the tests are the clearest statement of what the author believes should be true, and reading them first orients you quickly. The cost is anchoring: you adopt the author's model of the problem and become less likely to spot the case nobody considered. A common compromise is to skim test names early for orientation and read the assertions properly at the end.
  • What do you do when the change is correct but you would have designed it differently?
    Separate the two claims. If the difference affects the exposed contract or something expensive to change later, raise it as a design point with the reason and the cost. If it is an internal preference that a future reader could reverse cheaply, say so explicitly as a preference and let the change proceed. Reviewers who cannot tell those apart make every review a negotiation about taste.

It is the order a building inspector uses: does the building belong here at all, then is the structure sound, then do the fittings work, then is the paint tidy. Nobody argues about paint before checking the foundations.

saying these in an interview costs you the question

  • Starts at the first line of the diff and reads straight through
  • Treats formatting remarks as the substance of a review
  • Approves without opening the tests at all
  • Never reads the change description or the requirement behind it
  • Assumes passing tests mean the behaviour is correct
  • Critiques the implementation before deciding the approach is right

context

open as a page

Why does a defect found in production usually cost more to fix than the same defect found during development?

level: juniorimportance: must knowfreq 68%

basics

~20 s

A production defect costs more because the code change is only a fraction of the work. Someone must notice it, reproduce it and re-learn code written weeks ago, and the team then pays for data repair, support and an unplanned release.

open as a page

What are the driver and navigator roles in pair programming, and why rotate them?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The driver types and handles the mechanics. The navigator stays a step ahead — unhandled cases, naming, the remaining plan — and keeps talking. Rotating the keyboard on a timer keeps both people active and able to explain the change.

open as a page

What are the common quality ownership models, from an independent test group to whole-team quality?

level: juniorimportance: must knowfreq 64%

basics

~20 s

Four shapes recur: an independent test group that verifies finished work, testers embedded in each delivery team, whole-team quality where every engineer tests and no dedicated tester exists, and a small coaching group that builds testing skill in others.

open as a page

How do you measure a test suite's flake rate, and why is its trend more useful than its level?

level: middleimportance: must knowfreq 60%

basics

~20 s

Flake rate is the share of runs or cases giving a different verdict on the same unchanged revision. Measure it by recording every attempt, not just the final one, and track the trend: an acceptable level depends on suite size.

open as a page

Before a staged rollout starts, what stopping rule and blast-radius limit must already be defined?

level: seniorimportance: must knowfreq 64%

basics

~20 s

Written-down abort criteria with an observation window per stage, a named decider, and a default of reverting when evidence is ambiguous; plus a bounded exposure - cohort, data, irreversible side effects - and a rehearsed revert path.

open as a page

What is a synthetic check, and how does it differ from a test in the build pipeline?

level: juniorimportance: should knowfreq 57%

basics

~20 s

A synthetic check is a scripted transaction run against the live system on a fixed interval by a scheduler outside it, alerting when it fails. A pipeline test runs once per change, against a build.

open as a page

What is a feedback budget for a test suite, and what happens when the suite outgrows it?

level: juniorimportance: should knowfreq 47%

basics

~20 s

A feedback budget is the longest a check suite may run before its verdict stops changing what a developer does next. Past that limit people switch tasks, batch changes and stop running the suite locally.

open as a page

What makes a code review checklist effective, and why do review checklists decay over time?

level: middleimportance: should knowfreq 38%

basics

~20 s

An effective review checklist is short, drawn from defects this team actually let through, and made of items only a human can judge. Checklists decay because they only ever grow, the system's failure modes move on, and long lists get ticked ritually instead of read.

open as a page

What distinguishes an informal read, a walkthrough, a technical review and a formal inspection?

level: middleimportance: should knowfreq 44%

basics

~20 s

They differ in structure and purpose. An informal read is one peer's unstructured pass; a walkthrough is author-led and aims at understanding; a technical review judges conformance to a specification; a formal inspection adds named roles, preparation, entry and exit criteria and a logged defect list.

open as a page

What are the four cost-of-quality categories, and how does spending move between them?

level: middleimportance: should knowfreq 46%

basics

~20 s

Cost of quality splits spending into prevention (stopping defects being made), appraisal (looking for them), internal failure (rework before release) and external failure (everything after release). The argument is that prevention and appraisal spending buys down the two failure categories.

open as a page

When does an ensemble (mob) session earn the cost of the whole team on one change?

level: middleimportance: should knowfreq 44%

basics

~20 s

When the constraint is shared understanding rather than typing throughput: onboarding, a change nobody can safely make alone, or a decision that would otherwise be relitigated later. It costs the whole team's hour and pays only when knowledge is the bottleneck.

open as a page

What is traffic shadowing, and what does diffing the shadowed responses tell you?

level: middleimportance: should knowfreq 38%

basics

~20 s

Traffic shadowing duplicates live requests to a new version running alongside the current one and discards its responses. Comparing the two responses on real inputs shows where the new version behaves differently, before any user sees it.

open as a page

Who should sign off that a defect is fixed, and what makes that sign-off a rubber stamp?

level: middleimportance: should knowfreq 44%

basics

~20 s

Someone other than the author, who reproduces the original steps against a build containing the fix and records what they saw. It becomes a rubber stamp when the approver only reads the change description, never runs the case, and approves in seconds.

open as a page

What does a tester-to-developer ratio tell you about a team, and what does it hide?

level: middleimportance: should knowfreq 46%

basics

~20 s

A tester-to-developer ratio describes how testing labour is staffed, not how quality is produced. The same ratio fits a fast embedded team and a badly queued separate group, so read it as a symptom that needs explaining, never as a target to hit.

open as a page

What do time to detect and time to repair a broken shared-branch build measure, and how do you shorten each?

level: middleimportance: should knowfreq 41%

basics

~20 s

Time to detect runs from a change landing to a human seeing a credible red signal; time to repair runs from that signal back to green. Shorten detection with earlier checks and addressed notifications, repair with revert-first and clear ownership.

open as a page

Reviews on your team are approved within minutes, yet a silent data-corruption defect shipped. How do you make review effective?

level: seniorimportance: should knowfreq 54%

basics

~20 s

Fast approval is a symptom, not a success. Measure change size against reading time, find where the defect could have been caught, then shrink changes and aim review at the risky surfaces — rather than asking reviewers to try harder.

open as a page

Why is the 100:1 late-defect cost curve contested, and how should you cite it responsibly?

level: seniorimportance: should knowfreq 33%

basics

~20 s

The order-of-magnitude curve rests on small, old datasets from very different projects, and the round multiplier is usually quoted third-hand without its context. Cite the direction and the mechanism, back the size with your own measured repair cost, and give a range.

open as a page

How do you choose which work to pair on rather than pairing uniformly?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Pair where risk and unfamiliarity are highest: hard-to-reverse changes, wide blast radius, areas only one person knows, open design work, stuck investigations. Leave well-understood, mechanical, verifiable work to one person, and decide per work item, not per person.

open as a page

How can real-user telemetry act as a test oracle for behaviour no assertion covers?

level: seniorimportance: should knowfreq 49%

basics

~20 s

By turning properties that must hold for every real transaction into continuously evaluated assertions: per-record invariants, comparisons against a prior baseline, funnel step ratios and user-distress signals. Together they judge behaviour you could never enumerate as cases.

open as a page

An independent test group runs a 340-case regression pack after each two-week freeze, and defects reach authors nine days later. How would you change that?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Measure where the nine days actually go before reorganising anything, then shorten the loop incrementally: move a fast subset of the pack in front of the merge, embed a tester or two, and keep the group for the work only it can do.

open as a page

What must a quarantine policy for unreliable tests contain so that quarantine does not become permanent?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Quarantine takes an unreliable case off the blocking path while still running it. Staying temporary needs an evidence-based entry rule, a named owner, a timebox that expires into repair or deletion rather than renewal, and a capped, published list.

open as a page

How would you build the economic case for moving quality work earlier without leaning on a defect-cost multiplier?

level: principalimportance: should knowfreq 38%

basics

~20 s

Measure your own recent escaped defects, price the repair in hours and consequences, and compare that against the cost of the specific earlier check that would have caught them. Argue at the margin, state the range, and name what would prove you wrong.

open as a page

How do you justify pair programming's cost when the evidence for it is contested?

level: principalimportance: should knowfreq 38%

basics

~20 s

Do not quote a multiplier — published evidence on effort and defect rates is mixed and study conditions vary widely. Argue from a bounded local trial with signals agreed up front, and be explicit about what pairing does not replace.

open as a page

Which suite-health indicators would you publish as team targets, and how do you keep them from being gamed?

level: principalimportance: should knowfreq 38%

basics

~20 s

Publish a small paired set: flake-rate trend beside quarantine size and age, runtime beside the count of blocking cases, pass rate beside flake rate and time to repair. Keep them as team-level indicators, never thresholds with consequences.

open as a page