One vulnerable library appears in twelve services' scans — how do you dedupe and route it?
answer
- results are not findings
- collapse on identity, keep per-service instances
- the key includes ecosystem, name, version
- ownership follows the declaration
- one shared parent, one fix, twelve adoptions
basics
~20 sDeduplicate on the finding's identity — the advisory plus the package identity and affected version — not on the scan result. Track one item with one accountable owner, keeping the twelve service occurrences as child records.
solid answer
~40 sScans produce results; triage needs findings. The dedup key is the advisory plus the package identity and version — a package URL such as `pkg:gem/[email protected]` is exactly that identity — so twelve scan results collapse into one tracked item. Under it you keep an occurrence per service, each carrying the thing that differs: which artifact, and whether the dependency is direct or pulled in transitively through something else. Ownership then follows the declaration, not the scan: whoever owns the manifest that declares the direct dependency owns the bump. When all twelve pull it transitively through one internal shared library, the honest owner is that library's team — one fix fans out — and each service's occurrence closes only when it consumes the fixed version. The aggregate stays open until the last occurrence does.
go deeper
Know that the same vulnerable library showing up in many services is one problem, not many, and that the tracking system should reflect that. Be able to say what makes two scan results the same finding.
Explain the finding-versus-occurrence model and the composition of a dedup key: advisory plus ecosystem, package name and version. Describe how ownership is derived from the manifest that declares the dependency.
Show judgment on fan-out: recognise when one internal shared library is the single place to fix, and design closing rules so the aggregate survives until the last deployed occurrence is clean. Mention normalisation and monorepo edge cases.
Own the reporting consequence — raw finding counts across an estate mislead badly without deduplication, and the metric you publish shapes which teams get blamed. Decide where remediation of shared substrate is centrally funded versus pushed to product teams.
## Results, findings and occurrences A scanner emits a **result** every time it runs against every artifact. A twelve-service estate scanned nightly produces thousands of results a week describing a handful of actual problems. If your intake creates a ticket per result you get a queue nobody can work, duplicate effort across teams, and metrics that measure your scan cadence rather than your risk. So the first job of intake is collapsing results into findings. The useful model has three layers: - **Finding** — one advisory affecting one package identity in one affected version range. This is the unit a human reasons about and the unit a triage decision attaches to. - **Occurrence** — one finding present in one service/artifact, with the dependency path that put it there. This is the unit that gets fixed and the unit an exposure clock runs against. - **Result** — one scanner output row. Disposable; it only ever confirms or clears an occurrence. ## Choosing the dedup key The tempting key is the advisory identifier alone. It is wrong for two reasons. One advisory can cover several packages, sometimes across ecosystems, and it usually covers a *range* of versions — so deduping on the identifier alone merges fixes that are genuinely different work, and hides that only some of your occurrences fall inside the affected range. The workable key is the advisory plus a canonical package identity that includes ecosystem, name and version. A package URL is precisely that: ecosystem, namespace, name, version in one string, comparable across tools. Two details bite in practice. First, normalise the identity before comparing — ecosystems differ on case sensitivity and on how a namespace is written, and an un-normalised key silently produces duplicate findings that look identical on screen. Second, decide what happens when a service upgrades from one affected version to another affected version: that is the same finding with a changed occurrence, not a new finding, and treating it as new resets any clock you are keeping. ## Routing the finding to a human Deduplication produces one item; it does not produce an owner. Ownership should be derived, not typed, and it follows the **declaration** rather than the scan. - **Direct dependency**: the team that owns the manifest declaring it owns the bump. Code-ownership metadata on the manifest path is the cleanest source, because it is maintained for other reasons and therefore stays alive. - **Transitive through a third-party parent**: the service team still owns it, because only they can decide between waiting for the parent to release a fix, overriding the resolved version, or replacing the parent. Nobody outside the team can make that call safely. - **Transitive through an internal shared library**: the shared library's team owns the remediation. This is the case that makes deduplication pay for itself — one version bump upstream retires twelve occurrences, and routing it twelve times to twelve product teams both wastes the effort and produces twelve inconsistent answers. So the fan-out question is not "how do I avoid twelve tickets" in the abstract; it is "is there one place where a single change fixes all twelve". If there is, the aggregate's owner is that place, and the product teams' occurrences become an adoption task rather than a remediation task. ## Closing rules, and the metric they protect An aggregate finding closes when the last occurrence closes. This sounds obvious and is violated constantly: the first team to upgrade marks the shared ticket resolved, the dashboard goes green, and eleven services keep running the vulnerable version. Conversely, an occurrence should close on evidence — a subsequent scan of the deployed artifact showing the fixed version — and not on a human ticking a box, because the box gets ticked when the pull request merges rather than when the change reaches production. Keeping occurrences as first-class records also gives you the numbers that actually mean something: how many services are affected, how many are still affected after a week, and how much of your backlog is one problem counted twelve times. Raw finding counts across an estate are close to meaningless without this collapse, because a single popular transitive dependency can dominate the total and make a team look negligent when it has one thing to fix. ## Where the model leaks Monorepos and multi-artifact repositories break the assumption that one repository equals one service — the occurrence key needs the deployable artifact, not the repository. Vendored or shaded code breaks package-identity matching entirely, because the vulnerable code is present under a different name. And a finding whose derived owner resolves to nobody — a service whose owning team no longer exists — must land in an explicit unowned state rather than defaulting to the security team or, worse, to silence.
- Why not just deduplicate on the advisory identifier?Because one advisory can cover several packages and a range of versions. Deduping on the identifier alone merges genuinely different remediations into one item and hides that only some of your occurrences fall inside the affected range — so a partially-fixed estate reads as fully fixed.
- All twelve services pull the library transitively through one internal base library. Who owns the fix?The base library's team owns the version bump, because one change retires all twelve occurrences and produces one consistent answer. The service teams own adoption: their occurrence closes when a deployed build consumes the fixed base version. Routing the same bump to twelve teams multiplies effort and invites twelve different resolutions.
- When should an occurrence be considered closed?When a scan of the deployed artifact shows the fixed version — not when the upgrade pull request merges. Closing at merge measures your intent rather than your exposure, and it flatters the numbers by however long your release train takes.
saying these in an interview costs you the question
- Opens one ticket per scan result
- Deduplicates on the advisory identifier alone across ecosystems
- Closes the aggregate when the first service upgrades
- Assigns the aggregate to the security team as the fixer
- Treats one repository as one deployable service
- Counts raw findings across the estate as a risk metric