Why does a traceability-derived impact set both overstate and understate what a change actually affects?
answer
- The set describes the record, not the system
- Hubs pull in everyone who reaches them
- Reachability is not the same as effect
- The dangerous coupling was never linked
- Weight a hop by how it is used
basics
~20 sIt overstates because shared components pull in every requirement that reaches them, changed behaviour or not. It understates because real coupling — shared data, ordering, implicit contracts — was never written down as a link at all.
solid answer
~50 sA derived impact set describes the record, not the system. **Overstatement** comes from hubs: a shared sign-in path, a common data model or a shared setup fixture is linked from dozens of requirements, so any change reaching it drags all of them in and the set stops discriminating. **Understatement** comes from coupling nobody recorded — two features writing the same rows, an ordering dependency between scheduled work, a shared configuration value, behaviour a downstream consumer relies on but never asked for. Neither error is fixed by walking harder; more hops worsen the first and do nothing for the second. Weight a hop by how the shared thing is used rather than by whether it is reached, close blind spots from outside the record, and label each item by where it came from. Then say which of the two errors you think dominates this time.
code
pseudocode · 12 linesfor edge in linksFrom(changedRequirement):
reach = countRequirementsReaching(edge.target)
if reach > HUB_THRESHOLD and not dependsOnChangedBehaviour(edge.target):
include(edge.target, source = "hub", note = "verify usage before acting")
else:
include(edge.target, source = "link")
# the record cannot supply these; they come from outside it
for area in sharedDataWrites(change) + orderingDependencies(change) + pastEscapes(area):
include(area, source = "inspection", note = "no link exists for this")
report(impactSet, groupedBy = source)go deeper
Know that an impact set built from links can be too wide and too narrow at the same time, and that the two errors have different causes: shared components on one side, dependencies nobody recorded on the other.
Explain the hub effect — a component many requirements reach makes any change to it look as though it affects all of them — and give an example of coupling that is real but was never written down as a traceability link.
Demonstrate that you look outside the record: the change's own diff, what this area has let escape before, and a conversation with whoever built it. Expect to be asked which error dominates in your own team's record and what you did about it.
Own the trade-off between a record precise enough to discriminate and one cheap enough that people keep it current. Over-linking produces sets nobody trusts; under-linking produces sets nobody can defend. Decide where your product sits and make that choice explicit rather than emergent.
## An artefact of the record, not of the system A derived impact set is only ever a statement about the links somebody wrote down. The system has real dependencies; the record has recorded ones; the two overlap, and neither contains the other. That single fact explains both failure modes at once, and they run in opposite directions: the record holds relationships that no longer discriminate, and it is missing relationships that would. Neither error means the walk was done badly. Walking harder makes the first one worse and does nothing at all about the second. ## Overstatement: hubs and reachability A **hub** is anything many requirements point at — a shared sign-in path, a common data model, a shared setup fixture, a formatting helper every screen uses. The record stores *reachability*: this requirement is connected to that component. It does not store *usage*: which behaviour of the component the requirement actually depends on. So when a change touches a hub, every requirement that reaches it enters the impact set at equal weight, whether it depends on the changed behaviour or on some unrelated corner of the same component. Thirty requirements arrive, twenty-eight of them for no reason, and the set stops discriminating. A set that names most of the product is not exactly wrong — it is uninformative, which is worse, because it still looks like work. Two repairs help: - **Weight a hop by usage rather than counting it.** Before including everything that reaches a hub, ask which of those requirements depends on the part that changed. That is a few minutes of reading, and it removes most of the inflation. - **Mark hub-derived items as such in the output.** If the reader can see which items arrived through common ground, they can discount those deliberately instead of distrusting the whole set. ## Understatement: coupling nobody wrote down The opposite error is quieter and more expensive. Real systems couple in ways nobody ever expressed as a traceability link, because writing one was never anyone's job: - two features writing and reading the same rows in a shared store; - an ordering dependency between scheduled work, where one job assumes another already ran; - a shared configuration value that changes behaviour in three unrelated places; - behaviour a downstream consumer relies on but never asked for, so no requirement mentions it; - an implicit contract in the shape of a message or a file that a second team parses. None of these appear anywhere in the record, so no amount of walking finds them. They are found from outside it: the change's own diff and the data it writes, the history of what has escaped in this area before, defect records that crossed feature boundaries, and five minutes of conversation with whoever built the neighbouring feature. | Aspect | Overstatement | Understatement | | --- | --- | --- | | Cause | Hubs; reachability stored instead of usage | Coupling that was never worth recording | | Symptom | The set names most of the product | The set looks tidy and misses the breakage | | Found by | Reading how the shared thing is used | Sources outside the record entirely | | Cost | The set gets ignored | The set gets trusted and is wrong | | Repair | Weight by usage; mark hub items | Add items by inspection and label them | Note the asymmetry in the cost row. An inflated set is annoying and gets discounted. A confidently narrow one gets believed. ## Reporting a set that is wrong in both directions Since both errors are always present to some degree, the honest output distinguishes items by where they came from. Three sources, three labels: **derived from an explicit link**, **reached through a hub**, and **added by inspection because the record could not know**. A reader who can see those three groups can act on the set intelligently. A reader handed a flat list of sixty items cannot. It is also worth saying which direction you believe dominates this time, and why. In an area with a thick record and a lot of shared machinery, overstatement dominates and the useful work is pruning. In an area recently rebuilt, or one where several teams contribute without keeping links, understatement dominates and the useful work is asking people. The two situations need opposite effort, and only somebody who has looked at the record can say which one they are in. Finally, resist the temptation to fix understatement by widening. Adding two more hops to compensate for a blind spot you cannot see produces a bigger set with the same hole in it, plus noise. A blind spot is closed with information from outside the record, or it is not closed at all.
- How do you tell a genuine hub from a component that merely looks shared in the record?Check how the change touches it. A component reached by many requirements but changed only in a path one of them uses is not a hub for this change; a component whose changed behaviour every caller depends on is. The distinction lives in the usage, and the record stores reachability rather than usage, so it cannot make that call for you.
- What sources outside the traceability record help you find coupling it never captured?The change's own diff and the data it writes, the history of what has escaped in this area before, defect records that crossed feature boundaries, and five minutes with whoever built the neighbouring feature. All four surface dependencies that were real but were never worth anyone's time to write down as a link.
A contact list built only from phone records: everybody who rang the switchboard looks exposed, and the person who shares a kitchen but never called looks perfectly safe.
saying these in an interview costs you the question
- Treats every requirement reaching a shared component as impacted
- Believes a complete traceability record removes the blind spots
- Adds hops until the set feels safe, then presents it as derived
- Never looks outside the record for coupling nobody linked
- Reports the set without saying which way it is likely wrong