How do you detect that an equivalence class you assumed was uniform actually hides sub-partitions?
answer
- The class table is a hypothesis about code
- Look for what the code switches on
- Unreached branches under a passing suite
- Escapes name the class you split wrongly
- Compare effects, not only responses
basics
~20 sTreat uniformity as a hypothesis and look for evidence against it: unexecuted branches under a one-value-per-class suite, outcomes that vary with retries, caching, size, locale or load, and every escaped defect, which is a report that some class holds two behaviours.
solid answer
~50 sA class is a claim that the software treats a set of inputs identically, and that claim is about the implementation even though you derived it from the specification. Falsify it deliberately. Run the partition suite with a coverage measurement and look for handler branches no representative reached — an unvisited branch inside a class you called uniform is a split waiting to be named. Ask what the code switches on beyond the field values: a retry or duplicate-submission path, a cache hit, a size or encoding threshold, a per-tenant configuration, timing under concurrency. Each such switch is an axis that cuts through your classes. Then mine history: every escaped defect landed inside a class you believed uniform, so it names the sub-partition you missed. Finally, strengthen the assertions — a class can look uniform in the response while differing in its side effects.
go deeper
Know that an equivalence class is an assumption rather than a guarantee, and that a defect found inside a class you thought was uniform means the class needs splitting, not just another value.
Be able to name concrete axes that split a class — retries, caches, size thresholds, locale, configuration — and explain why field values alone do not express them.
Show the routine: coverage measurement scoped to the handler, reading what the code switches on, clustering real inputs by outcome and effect count, and mining escapes. Be ready to walk one incident end to end.
Own how class tables stay alive across teams — who reviews them when the specification changes, how escapes feed back into them, and how much production observation the strategy relies on instead of design-time analysis.
## The hypothesis you are actually making Equivalence partitioning is derived from the specification, but the coverage claim it makes is about the implementation: *these inputs travel the same path and produce the same behaviour, so one of them stands for all of them*. The specification cannot guarantee that. It describes intent; the code is free to branch on anything it likes. A **hidden sub-partition** is a class you named as one thing that the running system treats as two. Senior work on this technique is not producing the class table. It is knowing that the table is a hypothesis, and having a routine for falsifying it. ## Where the splits come from The axes that cut through a class are rarely the ones the specification talks about: - **Repeat and retry paths.** A request that the system has seen before may take an entirely different route from a first-time request, even though both carry identical field values. - **Caching and memoisation.** First call and second call differ, and the class you defined on the input has silently become two behaviours separated by state. - **Size and shape thresholds.** A payload above some internal limit is streamed, chunked or paged rather than handled in one piece. - **Encoding, locale and calendar.** Text that normalises differently, dates near a daylight-saving shift, or a locale-specific format can split what looked like one class of strings. - **Configuration and flags.** A per-tenant setting or a feature flag makes two deployments of the same class behave differently. - **Concurrency and load.** A path only reached when two requests overlap is a sub-partition of every class, and it appears at volume rather than in a quiet test run. ## A worked escape A hospital appointment scheduler had one valid class for booking requests: a whole-number lead time from 1 to 90 days with a known clinic code. One representative, 37 days, passed every build for months. In production, at a **1,200-request-per-minute** peak, some patients received two confirmation messages for a single appointment. The class was not uniform. Clients retried on a slow response and re-sent the same booking, and the handler recognised the repeat only after the confirmation had already been dispatched — a **duplicated side effect** on a request whose field values were identical to the one in the suite. The input domain had a second dimension nobody had written down: *first submission* versus *repeat submission*. Two sub-classes, one of them never tested, because from the specification's point of view they were the same booking. The fix at test-design level was not more values from 1 to 90; every one of those would have passed. It was to add the missing axis to the class table — first submission, repeat submission with the same identifier, repeat submission with a different identifier — and to assert the count of confirmations rather than the presence of one. ## The detection routine 1. **Measure what the representatives reach.** Run the partition suite under a coverage measurement scoped to the handler. Branches that no representative executed are candidate sub-partitions: something in the code distinguishes inputs your table says are the same. This is a white-box cross-check of a black-box design, and it is one of the highest-value uses of a coverage measurement, because it asks a specific question rather than chasing a percentage. 2. **Read the switches, not the rules.** Walk the path and list every condition that is not a field value from the specification: state lookups, feature flags, size checks, clock reads, cache probes. Each is an axis to test against your class list. 3. **Cluster real traffic.** Sample production inputs and group them by observed outcome, latency band and side-effect count. Two clusters inside one of your classes is direct evidence of a split, and it also finds classes you never wrote down because no field in the specification names them. 4. **Mine the defect history.** Every escaped defect entered through some input. Ask which class that input belonged to; by definition it was a class you believed uniform. Escapes are the cheapest source of true sub-partitions you have, and a pattern across several of them usually names an axis, not a one-off. 5. **Strengthen the oracle.** Two members of a class can return identical responses while differing in what they wrote, emitted or scheduled. If the assertion looks only at the reply, whole sub-partitions are invisible by construction. Assert counts of effects, not just their presence. ## What to do once you find one Split the class and give each sub-class its own expected behaviour and representative, rather than adding a second value to the same class — a second value with the same expectation asserts the very uniformity you have just disproved. Record the axis in the class table so it survives the next specification change, and revisit sibling classes: an axis such as retry handling or caching rarely cuts through only one class. Finally, be honest about the limit. This routine narrows the gap between the class table and the code; it does not close it. A class list is a living artefact whose job is to be wrong in ways you can find quickly, not to be provably right.
- Why is adding a second value from the same class the wrong response to an escape inside it?Because it repeats the disproved claim. A second value carrying the same expected behaviour asserts that the class is uniform, which the escape has already contradicted; if the split is along an axis the field values do not express — a repeat submission, a cache hit — every value in the range passes. The right move is to name the axis, split the class, and give each sub-class its own expected result and representative.
- How does a coverage measurement help here without becoming a coverage-percentage exercise?You are asking a targeted question, not chasing a number: under a suite of exactly one representative per class, which branches in the handler did nothing execute? Each such branch distinguishes inputs your table calls identical, so it is a candidate sub-partition to name. The output is a list of questions about the class table, not a percentage — and a suite at a high percentage can still leave that list non-empty.
- Which sub-partitions will this routine still miss?Ones that need a state or timing combination the environment never produces — an overlap window that only appears at production concurrency, a path behind a configuration no test deployment uses, or behaviour that depends on data volume built up over months. Those tend to surface through production observation rather than design work, which is why traffic clustering and defect archaeology stay part of the routine rather than being one-off exercises.
saying these in an interview costs you the question
- Adds more values from the same class instead of splitting it
- Treats the specification as a complete map of code paths
- Reads an intermittent duplicate as flakiness, not a second path
- Asserts only the response, never the count of effects
- Chases a coverage percentage instead of unreached branches
- Assumes a class that passed for months is therefore uniform