skip to content

How do you detect production systems that no asset inventory or account list knows about?

level: seniorimportance: should knowfreq 43%

answer

  1. you cannot query a list for what is absent from it
  2. what does a running workload consume anyway
  3. spend, identity, DNS, egress, certificates
  4. make registration itself a failable control
  5. unmatched item becomes a finding with an owner

basics

~20 s

Not by scanning the inventory: a system missing from it is invisible to it. Enumerate instead from sources every workload touches anyway — billing, identity, DNS, egress — and treat anything present there but absent from your registered estate as a finding.

solid answer

~50 s

You cannot query a list for the things that are not on it, so the only method that works is reconciling two enumerations built from independent sources. On one side, the registered estate: what the platform knows it operates. On the other, sources a running workload cannot avoid consuming — spend records, identity provider sign-ins, DNS zones and records, network egress, code repositories carrying deployment configuration. Anything appearing on the second side and missing from the first is the shadow estate. I then make registration itself a control — every production cluster must be registered with the platform — so a cluster that was never registered is a named failure with an owner rather than an anomaly somebody noticed. Newly found systems enter as unassessed, not as passes, until they can actually be collected. And the paved path has to be cheaper than the shadow path, or the reconciliation just refills every quarter.

go deeper

for a junior

Understand that some production systems are not in any inventory, and that a check can only ever look at systems something told it about. Be able to name one independent signal, such as spend, that reveals them.

for a middle

Explain the reconciliation: two enumerations from independent sources, and the unmatched set as the shadow estate. Say clearly why a newly discovered system is unassessed rather than passing or failing.

for a senior

Demonstrate you have run this. Name the sources you would reconcile, describe how you attribute an owner to an unmatched item, and explain why registration is written as a control rather than chased as a project.

for a principal

Own the incentive design. Discovery grows the denominator and lowers the score, so decide how coverage growth gets reported as progress and how the registered path is made cheaper than the shadow one.

This is the coverage problem in its hardest form. Stale inventory records are recoverable — the record exists and is wrong. A shadow system has no record at all, and no amount of care applied to the inventory will reveal it. ## Why one source can never find it Any enumeration built on a registry answers 'what does this registry contain'. If a team stood up a production Kubernetes cluster in an account nobody added to the platform, that cluster is not a stale record or a failed collection. It is absent from the question. Detection therefore requires a second enumeration whose construction has nothing to do with the first, and then a reconciliation between them. ## Sources a running workload cannot avoid The useful independent sources share a property: the workload consumes them to function, so opting out is not free. - **Spend and billing.** Compute costs money and the money is centrally recorded. Charges attributed to a container or database service in an account with no registered systems is a strong signal. - **Identity.** People and workloads authenticate. Sign-ins into an account or cluster that the platform does not operate show up in the identity provider's logs. - **DNS.** Anything reachable by name has a record somewhere, and zones are usually centrally delegated. - **Network egress and flow records.** Traffic leaving from address ranges nobody claims tells you something is running there. - **Code and deployment configuration.** Repositories containing manifests or plan files pointed at an endpoint the platform does not recognise. - **Certificate issuance.** Certificates issued for names nobody has registered. None is complete alone. The union is much better than any single one, and the reconciliation is a recurring job rather than a one-off sweep. ## Registration as a control in its own right The strongest move is to make the scoping problem a control that can fail. 'Every production cluster is registered with the platform' converts an unknown-unknown into a measurable state: the registry supplies one set, the independent enumeration supplies the population, and the unmatched entries are findings with a due date and a named owner. That is a meta-control — it does not protect a workload directly, it protects every other control's denominator. It is also the honest answer to anyone asking how you know your population is complete: not 'we believe it is', but 'here is the reconciliation, here is its delta, here is who owns each item in it'. ## What a discovered system is, and is not A newly discovered cluster is not a pass and it is not a failure of the encryption control. It is unassessed: it is now known to be in the population and nothing has yet evaluated it. Recording it as failing overstates what you know and buries the real remediation queue; recording it as passing repeats the original sin. It moves into pass or fail only once collection actually reaches it, which usually means deploying credentials and registering it properly. ## Making the reconciliation stick Two forces refill the shadow estate. The first is friction: if the registered path is slow — a ticket, a review, a wait — and the unregistered path takes an afternoon, teams will keep choosing the afternoon. Registration has to carry benefits people want, such as ingress, secrets distribution, log routing and on-call integration, so the cheap path is the visible one. The second is incentives: discovery grows the denominator, which pushes the pass rate down. If a team's scorecard drops because they found things, the rational response is to stop looking. Reporting the growth of the population as progress, separately from the pass rate, is what keeps discovery worth doing. ## Attribution A finding with no owner does not get fixed. Each unmatched item needs an owner derived from something durable — who pays for the account, who created the identity, who owns the repository that deploys into it. Attribution is often the slow part of the work, and it is worth doing before the finding is published rather than after, because an unowned finding tends to be argued about rather than closed.

  • You find eleven unregistered clusters. Do they count as failures of the encryption-at-rest control?
    No. They enter the population as unassessed, because nothing has evaluated them yet. Marking them failed asserts a violation you have not observed and floods the remediation queue with items whose real blocker is access, not configuration. They become pass or fail once collection actually reaches them.
  • What stops the shadow estate refilling next quarter?
    Two things. Make the registered path genuinely cheaper — registration should grant ingress, secrets, log routing and on-call wiring that teams want anyway. And run the reconciliation on a schedule so the delta is a standing metric with an owner, rather than a discovery exercise somebody funds once and never repeats.
  • Why is a meta-control like 'every production cluster is registered' worth writing at all?
    Because it protects every other control's denominator. A rule about encryption only ever describes the systems in scope; the registration control is what makes the scope itself measurable and failable. It is the only control whose failure explains why the others might be reporting a flattering number.

You do not find the guests who never signed the register by reading the register more carefully — you count the coats.

saying these in an interview costs you the question

  • Proposes scanning the existing inventory more thoroughly
  • Relies on teams to self-report the systems they created
  • Treats discovery as a one-off cleanup project
  • Counts a newly discovered system as compliant by default
  • Assumes a single provider listing enumerates everything
  • Publishes unmatched findings with no owner attached

context