Your DNS zone sits with one team and cloud resources with dozens — who is accountable for names that outlive their targets?
answer
- two owners, neither owns both halves
- creation is gateable, deletion is not observable
- the record dies with the resource
- nothing on a cost report owns a name
- make the residual cheap, not zero
basics
~20 sAccountability belongs with whoever creates the record, enforced by making the record part of the resource's lifecycle. A central zone team can gate and expire records but cannot know when a resource dies, so pure centralisation fails.
solid answer
~50 sThe structural cause is that the record and the resource have different owners, different systems and different lifecycles, and neither owner's job covers the other half. Centralising fixes the wrong end: the zone team can hold the pen but cannot know that a product team released a bucket last Tuesday. The workable model puts accountability on the record's creator and then makes forgetting hard - records created and destroyed by the same automation that creates and destroys the resource, self-service creation demanding a named owner and a renewable expiry, and a gate that refuses a target the requester cannot prove they hold. The hard parts are organisational: no budget line owns a name, there is no defect to cite in a mandate, and teams can refuse. So plan for residual - accept that some names will dangle and spend part of the effort making a hijacked one cheap.
go deeper
Know that the record and the resource are owned by different teams in different systems, and that this split is why stale records happen at all.
Be able to describe binding the record to the resource's lifecycle in provisioning, and why hand-made records for vendor tenants escape it.
Compare centralisation, lifecycle binding and owner-plus-expiry on what each actually controls, and place the creation-time gate that refuses targets the requester cannot prove they hold.
Own the constraints: no budget line for a name, no defect to cite in a mandate, teams that may refuse, third parties publishing under your domain. Say what residual you accept and where you spend to make it cheap.
## Why this is an ownership problem and not a discipline problem Every dangling record has two competent owners who both did their job. The product team's teardown ticket says “delete the resource” and closes correctly when the resource is gone. The platform or network team administers the zone and has no visibility into a resource lifecycle inside another team's cloud account. Nobody is careless. The record outlives the resource because **no single role owns both halves**, and no amount of reminding people to be careful changes a structure. That framing matters in a room, because the reflexive proposals — more training, a stricter checklist, a periodic re-check — all leave the structure untouched and therefore all decay. ## Three models, and what each actually costs **Centralise the zone.** One team holds the pen; every record change is a request. It gives clean control of creation and none of deletion, because the central team cannot observe a resource being released in someone else's account. It also makes the team a bottleneck that product teams route around — typically by acquiring a second domain nobody governs, which is strictly worse. Centralisation is necessary for the *gate*, insufficient for the *lifecycle*. **Bind the record to the resource's lifecycle.** The record is created by the same automation that creates the resource and destroyed by the same teardown. This is the only model where the failure mode is structurally impossible rather than merely discouraged, and it should be the default for anything provisioned as code. Its limit is everything created by hand: vendor tenants, marketing landing pages, third-party services bought on a card. Those are exactly the records that dangle most often, because they are the ones no automation ever touched. **Named owner plus expiry.** Self-service creation that requires a human owner and a renewal date, with unrenewed records repointed rather than deleted first. It reaches the hand-made records the second model misses, at the cost of a recurring renewal burden and the certainty that some renewals will be rubber-stamped. Repointing instead of deleting is the humane version: it removes the exposure while leaving the name recoverable if the renewal lapse was a mistake. Most estates need the second model as the default and the third as the net beneath it, with the first providing the gate that both hang from. ## The gate worth having at creation time Refuse to publish an alias to a provider target that the requester cannot demonstrate they currently hold, and prefer bindings that require an ownership token to stay published. That converts a class of future dangling records into records that are worthless to a claimant even if they do dangle. It is cheap because it applies only at creation, and it is unpopular for the same reason every gate is. ## The constraints a technical answer will miss - **No budget owns a name.** Cloud spend is attributed to accounts; a DNS record costs nothing and appears on nobody's cost report, so it never surfaces in the review where teams prune things. - **There is no defect to point at.** Mandates usually travel on a patch deadline. You are asking teams to change how they tear things down, with no vulnerability identifier to cite, which is a much harder ask and needs a different lever — usually attaching it to the provisioning path everyone already uses rather than to a policy everyone must read. - **A team can refuse.** Product teams that own their accounts can decline extra process, and escalating each refusal costs more than the finding is worth. The durable answer makes the safe path the default in shared tooling instead of asking for compliance. - **Third parties create records too.** Agencies, marketing platforms and acquired subsidiaries publish under your name. Contract language and the creation gate are the only levers, and neither is instant. ## Deciding the residual honestly A principal answer says out loud that this will never reach zero: acquisitions arrive with zones nobody has read, vendors are cancelled by people who never heard of a zone, and hand-made records will keep being made. So split the investment. Spend on the lifecycle binding where provisioning is automated, on the creation gate everywhere, and — crucially — on making a hijacked name *cheap*: host-only cookies, allowlists written as explicit hosts rather than as a domain suffix, no wildcard record standing in front of the whole zone. The second half of that programme needs no cooperation from the teams that create records, which is precisely why it is the part that survives contact with the organisation. ## The sentence to land “We are not going to stop every record outliving its resource. We are going to make the automated path take both, gate the hand-made ones at creation, and make sure that a name a stranger takes inherits almost nothing.”
- Why does centralising all DNS changes fail to solve this on its own?Centralisation controls creation, which is the observable event, but deletion is triggered by something the zone team cannot see — a resource released inside a product team's account. It also creates a bottleneck teams work around, sometimes by buying a domain outside the governed zone. Use it as the gate, not as the lifecycle.
- How do you get an estate of dozens of teams to adopt this without a vulnerability deadline to cite?Attach it to the path they already use rather than to a policy they must read: make the provisioning module that creates the resource create and destroy the record, and make the self-service form require an owner and an expiry. Adoption that requires no decision from a team is the only kind that holds when their quarter gets busy.
- What do you do about names published under your domain by an agency or an acquired subsidiary?Treat them as third-party creation: the levers are the creation gate on your zone and contract language requiring delegated names be returned or repointed at exit. Neither is fast, so acquisitions in particular should be handled as a one-off reconciliation of the acquired zone against live resources, budgeted as part of the integration rather than absorbed.
saying these in an interview costs you the question
- Proposes training or a stricter checklist as the fix
- Centralises the zone and calls the lifecycle solved
- Assumes every record is created by automation
- Aims for zero dangling names with no residual plan
- Puts accountability on a team that cannot observe the resource