skip to content

Why is an ice-cream-cone shaped test suite treated as an anti-pattern?

level: middleimportance: should knowfreq 57%

answer

  1. The pyramid drawn upside down
  2. Manual scoop above broad automation
  3. Slow feedback, poor localisation
  4. Flakiness multiplies across the top
  5. Escape by pushing checks downward

basics

~20 s

An ice-cream-cone suite inverts the pyramid: a broad layer of end-to-end automation, often with manual checking piled on top, over a thin base of narrow tests. Feedback becomes slow, failures stop naming a cause, and instability spreads.

solid answer

~50 s

The ice-cream cone is the pyramid upside down — most of the checking sits in assembled-system tests, with a scoop of manual regression above them, and only a sliver of narrow tests underneath. It is an anti-pattern because every cost the pyramid tries to minimise is maximised: the suite takes tens of minutes so it runs late and rarely; a failure names a journey rather than a component, so diagnosis dominates the team's time; and each test touches so many moving parts that nondeterministic failures accumulate until people re-run rather than investigate. It also duplicates heavily — the same branch gets exercised through the slowest available path dozens of times, while a whole class of rules has no direct test at all. Teams usually arrive there by automating an existing manual script rather than by testing behaviour where it lives.

go deeper

for a junior

Know the picture and the name: most tests at the assembled-system level, a manual scoop on top, few narrow tests underneath. Be able to say that feedback is slow and failures are hard to trace.

for a middle

Explain the mechanics — run time, failure localisation, compounding nondeterminism, duplication with coverage holes — and describe how a team drifts into the shape by automating an existing manual script.

for a senior

Show a migration you would actually run: freeze the top, let real failures queue the narrow tests, delete only on replacement, and fix the testability problem that made narrow tests expensive.

for a principal

Frame it as an incentive problem, not a taste one. Be ready to discuss why the shape recurs when checking is owned separately from design, and what you change organisationally so the cheap level is also the easy one.

## The shape An **inverted pyramid** puts most automated tests at the assembled-system level and few at the narrow level. The **ice-cream cone** is that inverted triangle with a scoop on top: a large body of manual regression checking sitting above the automation. Both names describe the same pathology — checking concentrated where it is slowest, most fragile and least diagnostic. ## How teams get there The shape is rarely chosen; it accretes. - **Automating the manual script.** A team already has a regression script written as user steps. Turning each step into an automated one produces top-level tests exclusively, because the script never mentioned components. - **Testing added after the code.** When checking is written by people who did not design the units, or written weeks later, the only stable seam left is the outside of the system. - **Code that resists narrow tests.** If behaviour is tangled with wiring and global state, a narrow test is expensive to write and a top-level one is cheap — so the incentive points the wrong way, and the design never gets the pressure that would fix it. - **A visible-progress bias.** A top-level test demos well: it looks like the product working. Narrow tests are invisible to anyone outside the team, so they get cut first under schedule pressure. ## Why it hurts **Feedback arrives too late to be cheap.** A suite dominated by assembled-system tests runs in tens of minutes at best. It therefore runs after the change is finished rather than during it, so a defect is found once context has been lost, and often after other changes have merged on top. **Failures do not diagnose.** Each failure names a journey. The candidate causes include every component on the path plus the environment, data, deployed versions and timing. The dominant cost of the cone is not authoring — it is the recurring hours spent turning a red journey into a named cause. **Instability compounds.** Independent flakiness rates multiply across a broad top layer: many long tests, each touching many moving parts, produce a suite that is red for unrelated reasons on most runs. The team's rational response — re-run and move on — quietly converts genuine failures into ignored noise, which is worse than having no suite, because it looks like coverage. **Duplication with holes.** Top-level tests re-walk the same happy path repeatedly, so branches near the entry point get exercised dozens of times while error handling, boundary conditions and invariants buried a layer down have no direct test at all. A cone can have a large test count and a genuinely thin risk coverage. **Cost of change.** Because top-level tests couple to end-to-end structure, a small refactor breaks many of them at once, and each break costs a full investigation. The suite becomes a reason not to change the system — the exact inversion of what tests are for. ## Getting out Inverting a cone is a gradual, defect-driven job rather than a rewrite. 1. **Stop widening the top.** Agree that new checking goes in at the lowest level that can catch the risk. The shape stops worsening immediately, which is more than most cleanups achieve in a quarter. 2. **Use failures as the queue.** Each time a top-level test goes red for a real reason, write the narrow test that would have caught it at the level where the rule lives, then decide whether the top-level case still earns its seconds. 3. **Delete on replacement, not on principle.** A top-level test is removed when something cheaper covers its risk — not because a target ratio says the tip should be thinner. 4. **Fix the testability that caused it.** If a rule cannot be tested narrowly, that is a design finding: extract the behaviour from the wiring so a narrow test becomes the cheap option. 5. **Keep a deliberate thin tip.** The goal is not zero assembled-system tests. A few journeys that prove wiring, configuration and deployment are exactly what the base cannot see. ## The honest caveat A cone is not always incompetence. Around a legacy system with no seams, broad outside-in tests may be the only safety net available, and building them first is a reasonable *transitional* strategy — the classic advice is to wrap the system in coarse characterisation tests before carving it up. The anti-pattern label applies when the shape is the destination rather than the scaffolding, and when nobody is converting the scaffolding into cheaper tests as seams appear.

  • Is a broad outside-in suite ever the right starting point?
    Yes, around legacy code with no seams. Coarse characterisation tests written from the outside pin current behaviour so it can be refactored safely, and that is standard advice. It becomes an anti-pattern when the shape is treated as the destination and nobody converts it into cheaper tests as seams appear.
  • Why does a large test count not prove a cone is well covered?
    Because top-level tests re-walk the same paths. Dozens of journeys hit the entry-point branches repeatedly, while invariants and error branches a level down have no direct test. Counting tests measures effort; what matters is which distinct risks have a check somewhere.
  • What is the first change you would make on a team living with a cone?
    Freeze the top: new checking goes in at the lowest level that can catch the risk. That single rule stops the shape worsening without asking for a rewrite, and it makes each genuine top-level failure the natural trigger to write the narrow test that should have existed.

It is a smoke alarm that only sounds once the whole house is alight: technically it is monitoring everything, but by the time it speaks nobody can tell which room started it.

saying these in an interview costs you the question

  • Calls any use of end-to-end tests an ice-cream cone
  • Proposes deleting the top layer before replacing its coverage
  • Says a high test count means coverage is fine
  • Blames flakiness on tooling rather than on scope
  • Cannot name a case where broad outside-in tests are justified

context