Your CMDB's owner field is wrong for a third of production systems - how do you fix triage?
answer
- you do not own the source
- price it in their currency
- observed beats declared
- gate at production onboarding
- no shadow asset database
basics
~20 sYou cannot fix a system you do not own. Price the triage cost in numbers the owning teams answer for, and stop relying on one declared field by deriving ownership from deploy history, on-call rotas and code owners.
solid answer
~50 sMeasure the damage in their currency first: minutes added per alert, escalations sent to the wrong team, alerts closed with ownership never established, and the delay before a human could speak for a system at 04:00. That turns a security complaint into an operational number platform owners recognise. Then design around it: treat the CMDB as one voter among observed signals - deployment actor, on-call rota, code owners, cloud tags - and resolve conflicts in favour of what is observed. Make ownership a gate where you have leverage, such as production onboarding, agreed with the owning teams rather than imposed. Keep a small curated override list for crown-jewel systems and accept an imperfect long tail. Do not build a shadow asset database you cannot keep current, and never let a missing owner become a reason to close alerts.
go deeper
You are not expected to fix this, but know why it matters: a wrong owner costs real minutes on a live alert. Record the stale entry as a finding and keep triaging rather than stalling on it.
Be ready to name ownership signals other than the CMDB - on-call rotas, deploy history, code owners, cloud tags - and explain why continuously regenerated data beats a field last edited two years ago.
Show how you would design triage to survive bad asset data: precedence across sources, disagreement surfaced to the analyst, rehearsed fallback contacts, and never letting missing ownership influence a malicious-or-not verdict.
Own the negotiation. Quantify the cost in operational terms, find the gate where you have leverage, publish per-team coverage, decide explicitly what you will not build, and pre-agree what the SOC may do on a production system nobody claims.
## Why this is a leadership question rather than a tooling one Every enrichment lookup a SOC runs against asset data is a dependency on a system another organisation maintains for a different purpose. A configuration management database exists for change control, licensing and service management; nobody funds it so that a security analyst can find a human at 04:20. When its owner field decays - and it always decays, through reorganisations, leavers, team splits and acquisitions - the SOC feels the cost and has neither the mandate nor the budget to fix the cause. That asymmetry is the whole problem, and it is why the answer is part negotiation, part compensating design, and part accepting a residual. ## Measure it in their currency A security team saying *your data is bad* gets nothing. A security team saying *stale ownership adds a median of eleven minutes to every alert on a production database, sent 38 escalations to teams that no longer own the service last quarter, and in two cases delayed containment past an hour* gets a meeting. The metrics worth collecting are cheap because they fall out of case data you already keep: - Median and 90th-percentile minutes from alert open to a human confirming ownership. - Number of escalations routed to the wrong team, and the time they burned. - Alerts closed with ownership never established. - Coverage: share of production assets whose recorded owner resolves to a current employee in a current role. That last one is the number to publish per owning team, because it puts the cost on the group that can fix it and creates a comparison they care about more than any policy. ## Stop relying on a single declared source The deeper fix is architectural. **Declared data decays; observed data is regenerated continuously.** Ownership can be inferred from signals the estate emits anyway: who deployed to the system in the last ninety days, whose on-call rota covers the service, which repository's code owners build it, which team's service principals authenticate to it, what the cloud resource tags say, who has approved recent changes. None is authoritative alone; together they beat a two-year-old field. So model ownership as a resolution with precedence rather than a lookup: prefer a current on-call rota, then recent deploy or change actors, then code owners, then the CMDB, and surface the disagreement to the analyst instead of hiding it. An analyst who sees *CMDB says A. Mensah (no longer active); on-call rota says Payments Platform, primary reachable* makes a better and faster decision than one handed a single confident wrong answer. ## Use the leverage you actually have You will not get a data-quality programme funded on security's ask alone. What you can often get is a **gate at a moment when someone needs something from you**: production onboarding, telemetry onboarding, a change freeze exception, a compliance attestation. Requiring a resolvable owner before a system is accepted into production monitoring is defensible, cheap for the requesting team, and it fixes new drift at source. It must be agreed with the platform owners rather than announced, or it becomes a queue of exceptions that erodes in a quarter. In regulated environments there is a second lever worth using honestly: incident-notification obligations and audit findings put a clock on identifying affected systems and their owners. That argument moves executives when an operational one does not. Use it where it is true; inflating it burns the credibility you will need next time. ## Decide what you will not do - **Do not build a shadow CMDB.** A security-maintained copy of the whole estate is accurate for one quarter and then becomes a second wrong source with your name on it. Curate a small override table for the crown-jewel systems - the ones where a 04:20 call actually happens - and let the long tail be resolved by the observed signals. - **Do not let missing ownership close alerts.** Ownership is consequence enrichment; its absence says nothing about whether activity was malicious. An unowned asset arguably deserves more scrutiny, not less, because nobody is watching it. - **Do not gate detections on asset data.** Rules that only fire on assets with a valid criticality field create a blind spot exactly where the records are worst. ## Accept and state the residual After all of it, some systems will still resolve to nobody at 04:00. The honest position is to say so, to bound it - which tier of asset, what share, what the fallback path is - and to have the fallback rehearsed rather than improvised: platform escalation line, duty manager, the authority for the SOC to take a defined containment action on an unowned production system without an owner's sign-off. Getting that last authority agreed in advance, in daylight, is often worth more than any amount of data cleaning, because it converts the unresolvable case from a stall into a decision somebody has already approved. ## What an interviewer is listening for The weak answer is *raise it with the CMDB team and add it to the risk register*. The strong answer quantifies the cost, redesigns around declared data rather than waiting for it, finds a gate with real leverage, names what will not be built, and pre-agrees what the SOC may do when ownership simply cannot be established.
- Which single metric would you publish to move the owning teams?Per-team ownership resolution coverage: the share of that team's production assets whose recorded owner maps to a current employee in a current role. It is objective, it is theirs to fix, and it compares across teams, which drives behaviour far better than an aggregate security complaint. Pair it with the triage minutes lost so the number has a consequence attached rather than being scorekeeping.
- An analyst proposes suppressing alerts on assets with no valid owner. What is your answer?No, and it is worth explaining why clearly. Ownership is enrichment about consequence, not evidence about behaviour, so its absence cannot reduce the probability that activity was hostile. An unowned production system is usually less maintained and less watched, which makes it a better target, not a safer one. Fix the routing problem with fallback escalation paths, not by removing visibility.
- How do you keep the derived-ownership approach from becoming its own maintenance burden?Derive at query time from sources that are already maintained for other reasons - rota tooling, deployment records, repository ownership, cloud tags - rather than materialising a copy you have to reconcile. Keep the precedence logic short enough to explain in one screen, show analysts which signal answered and how old it is, and curate a manual override list only for the small set of systems where a wrong answer at 04:00 is unacceptable.
saying these in an interview costs you the question
- Proposes building a security-owned copy of the whole estate
- Suppresses or closes alerts on assets with no resolvable owner
- Raises it as a risk-register item with no measurement
- Treats the CMDB as the only possible ownership source
- Announces a mandate without agreeing it with platform owners