When is deviating from the canonical test pyramid shape the right call?
answer
- The shape encodes an assumption
- Ask where the risk actually lives
- Thin glue services invert the logic
- Alternative shapes are contested heuristics
- Escaped defects decide, not a picture
basics
~20 sDeviate when the system's real risk does not live inside components. For thin glue services, data pipelines and configuration-heavy systems, weight the boundary-crossing level instead — the invariant is to catch each risk at the cheapest level that can see it.
solid answer
~50 sThe pyramid encodes an economic argument, not a law, so the question is whether its assumption holds for your system: that most risk lives inside components and can be caught narrowly. Where a service is mostly orchestration, mapping and configuration — a thin layer over stores and other services — narrow tests largely assert that stand-ins behave as configured, and the defects that actually reach production are boundary defects. There the sensible shape is bulkier in the middle, which is what the trophy and honeycomb models argue for; other teams describe a diamond around legacy code with coarse characterisation tests. **These alternatives are contested and the evidence is largely experiential, so present them as competing heuristics rather than findings.** What survives every variant: prefer the cheapest level that can catch a given risk, keep feedback time bounded, and let escaped defects — not a picture — decide the shape.
go deeper
Know that the pyramid is a default rather than a rule, and that some systems have most of their risk at a boundary rather than inside a component. You are not expected to argue the alternatives.
Be able to name the assumption the shape rests on — that risk is reachable narrowly — and give one concrete system type, such as a thin orchestration service, where that assumption is weak.
Show that you would classify real escaped defects by the cheapest level that could have caught them, and that you price a heavier boundary layer in provisioning and run time before recommending it.
Own the debate without joining a camp: state that the competing shapes are experience-based and contested, keep the feedback budget and cheapest-level rule as the invariants, and address the incentives that actually determine a suite's shape.
## What the pyramid assumes The pyramid's argument is: cost and diagnosis difficulty rise with scope, so push checking down. That is unconditionally true. The *shape* it recommends adds a second, conditional claim: **that most of your risk is reachable from the narrow level.** When that second claim is false, following the shape produces a wide, fast, cheap base that fails to see the defects that actually escape. So deviation is not rebellion; it is checking whether the assumption holds for the system in front of you. ## Where the assumption breaks **Thin orchestration services.** A service whose job is to receive a request, map it, call two dependencies and combine the answers has almost no interesting logic in isolation. Narrow tests over it mostly assert that a stand-in returns what it was told to return. The genuine risks — mapping errors, serialization mismatches, null and default handling across a boundary, timeout and retry behaviour — all live at the boundary. **Data pipelines and ingest paths.** Correctness here is largely a property of schema, ordering, deduplication against real storage semantics and late-arriving data. On a fleet telematics ingest, the rules worth guarding are things like whether a retried delivery produces a duplicated side effect, and whether readings arriving out of order still produce a monotone odometer. Some of those can be pinned narrowly; several are only real against a genuine store. **Configuration-dominated systems.** When behaviour is driven by feature settings, routing rules or policy documents rather than by branches in code, the code paths are few and the combinations are many. The risk is in the combination, which the narrow level cannot see at all. **Legacy without seams.** Where no unit boundary exists to test against, coarse outside-in characterisation tests are the only affordable net, and a diamond or even a temporary cone is a rational *transitional* shape while seams are carved. ## The competing shapes, honestly presented Several alternative pictures circulate. One argues for weighting the integration band most heavily on the grounds that those tests give the best confidence per unit of cost for typical applications. Another draws the suite as a honeycomb with a thick service-level middle for orchestration-heavy services. A third describes a diamond where legacy constrains the base. **None of these rests on strong empirical evidence; they are structured experience reports, and the debate between them is genuinely unresolved.** A principal-level answer says that plainly instead of picking a camp. What every variant agrees on is the underlying rule — catch each risk at the cheapest level that can see it — and they differ only in where they believe risk usually lives. ## How to decide, concretely 1. **Start from escapes, not from a picture.** Classify the defects that reached production over a real period by the level that could have caught them cheapest. The distribution *is* your recommended shape. If most escapes are boundary defects, a wide base is not virtue, it is misallocation. 2. **Set a feedback-time budget.** Decide what a developer will wait for before pushing, and what the merge path may cost. The budget constrains the shape from above regardless of where risk lives; if the ideal shape breaks the budget, that is an argument about parallelism and environment cost, and it should be made explicitly. 3. **Check what a narrow test would actually assert.** If the answer is 'that the stand-in was configured correctly', the test is theatre and the level is wrong for that risk. 4. **Account for the cost of the middle.** A heavier boundary layer means real dependencies to provision, state to reset between tests, and slower runs. Deviating upward is a real spend; be able to name what it buys. 5. **Re-decide periodically.** Shape follows architecture. A system that grows genuine domain logic should grow a base; one that is decomposed into thin services should expect its middle to thicken. ## The organisational half Shape is also a function of who writes tests. If checking is owned by a group separate from the people designing components, the suite drifts upward regardless of doctrine, because the outside is the only seam that group can reach. Likewise, if provisioning a real dependency is slow or requires a ticket, engineers will write a narrow test whether or not it covers the risk. A principal who wants a different shape usually has to change an incentive — testability of the code, or the ease of standing up a boundary — rather than issue a guideline. Publishing a target picture without changing either is the most common way this decision fails.
- What evidence would convince you to thicken the boundary-crossing level for a given service?A classification of real escaped defects showing most could only have been caught against a genuine dependency, plus an inspection of existing narrow tests showing they largely assert stand-in configuration. Together those say the risk is at the boundary and the base is not seeing it. I would also price the change: provisioning, state reset and added run time.
- How do you avoid this argument becoming licence for a slow, top-heavy suite?Keep the two invariants explicit: a bounded feedback-time budget, and the rule that a risk is caught at the cheapest level that can see it. Deviation shifts weight into the middle for named risks; it never justifies moving checks to the assembled-system level because writing them there is easier.
- Why is naming a specific alternative shape less important than the reasoning?Because the alternatives are experience reports rather than measured findings, and interviewers can tell the difference between someone reciting a picture and someone reasoning about where risk lives, what each level costs, and what the escape data says. The reasoning transfers to a system none of the pictures anticipated.
Choosing a suite shape is like staffing inspections on a production line: you place inspectors where defects are actually introduced, not evenly along the belt because the diagram looks balanced.
saying these in an interview costs you the question
- Presents an alternative shape as empirically proven
- Uses deviation to justify a slow top-heavy suite
- Applies one shape to every service in a portfolio
- Ignores the provisioning cost of a heavier middle level
- Publishes a target shape without changing testability or incentives