skip to content

Your platform bill contains waste nobody can name - how would you build a sweep that proves which resources are dead rather than guessing?

level: seniorimportance: should knowfreq 48%

answer

  1. no single source shows waste
  2. inventory plus billing plus use evidence
  3. every account, every region
  4. evidence differs per resource kind
  5. quarantine before you delete

basics

~20 s

Join three sources: the platform inventory of everything that exists, the billing records that price it, and last-use evidence appropriate to each resource kind. Rank the unreferenced rows by monthly charge, attach an owner and the evidence, and quarantine before deleting.

solid answer

~50 s

Waste is found by reconciliation because no single source shows it. The **inventory** from the management API says what exists, in every region and account, including the ones nobody opens. The **billing records** say what each of those costs, which is how the list gets ranked. **Last-use evidence** is what turns "exists" into "dead", and it is specific to the resource kind: a volume attached to nothing, an address routed to nothing, a store with no connection in a full business cycle, an image nothing has launched from, a provisioned floor whose peak consumption never approached it. Each candidate row carries its charge, its evidence and an owner, so the conversation is about the evidence rather than about deletion. Then quarantine - stop or deny access for a cycle, keep a restorable copy - before anything is actually removed.

code

pseudocode · 22 lines
pseudocode
candidates = empty list

for each account in organization:
  for each region in account.enabledRegions:
    for each resource in inventory(account, region):

      lastUse = usageEvidenceFor(resource.kind, resource.id)
      if lastUse is null or age(lastUse) > oneBusinessCycle:

        charge = monthlyChargeFor(resource.id)   # from billing records
        owner  = ownerLabelFor(resource.id)
        if owner is null:
          owner = "unclaimed"                    # only when the label is absent

        append(candidates, {
          id: resource.id, kind: resource.kind,
          account: account, region: region,
          charge: charge, evidence: lastUse, owner: owner
        })

sortDescending(candidates, by = charge)
return candidates   # a candidate list, never a delete list

go deeper

for a junior

Recall the three inputs a sweep needs - what exists, what it costs, and when it was last used - and that none of them alone identifies waste.

for a middle

Explain why evidence of use is resource-specific, and name the evidence you would trust for a volume, an address, a store and a provisioned floor.

for a senior

Show the operating discipline: a full business cycle of measurement, the evidence published with the row, quarantine before deletion, and a restorable copy for anything holding data.

for a principal

Position the sweep as detective and argue for the preventive half at creation time - expiry, delete-with-parent defaults, retention chosen up front - because ordinary work produces residue continuously.

## Why one source is never enough Each of the obvious places to look is blind in a specific way. - A **cost report** groups by service and dimension. It shows that block storage is large, which is also true in a healthy estate. "Attached to nothing" is not a billing concept, so the report cannot express the thing you are hunting. - An **inventory listing** shows everything that exists but prices none of it, so it cannot be ranked and drowns you in resources that cost almost nothing. - **Intuition** looks where people remember building things. Residue is, by construction, in the places they stopped looking: a region chosen for one experiment, an account opened for a migration, an environment built around a redesign that shipped a quarter ago. Reconciliation is the join of the first two plus a third column that neither contains: evidence of use. ## The three inputs 1. **Inventory**, read from the platform's management API, enumerated across **every account and every region** the organisation owns. Scoping this to "our region" is the single most common way a sweep misses the oldest residue. 2. **Billing records** at resource granularity, so each row carries a monthly charge and the list can be sorted by what is worth arguing about. 3. **Last-use evidence**, which has no single form. This is the part that takes real work, and it is what distinguishes a sweep from a guess. | Resource kind | Evidence that it is dead | What that evidence misses | |---|---|---| | Block volume | Attached to nothing; no read or write in a cycle | A volume detached deliberately, awaiting a rebuild | | Reserved address | Reserved, routed to nothing | An address held for a pending cutover | | Snapshot or image | Source gone; no restore or launch ever recorded | The one copy of something whose source was just deleted | | Managed store | No connection across a full business cycle | A month-end or quarterly consumer | | Provisioned throughput floor | Peak consumption far below the floor for a cycle | Headroom somebody holds on purpose | | Whole environment | No deployment, no login, no job run in a cycle | A rehearsal environment used once a quarter | The right-hand column is why the output of a sweep is a **candidate list**, not a deletion list. Every evidence type has a false positive, and the shared shape of all of them is the same: something used on a longer cycle than the window you measured. Measuring across a full business cycle - a month at minimum, a quarter where the business has quarterly rhythms - removes most of them. ## Owners, and what the sweep is not Each row wants an owner attached so somebody can confirm or dispute it. The scheme that produces that owner - how resources get labelled, what happens to the unlabelled remainder, how shared costs are split - is a substantial subject of its own and not part of the sweep; the sweep is a **consumer** of whatever attribution exists, and it is also the thing that exposes how much of the estate has none. It is worth being precise about what kind of control this is. A sweep is **detective**: it tells you something happened and costs money. It prevents nothing. The preventive half - an expiry set when an environment is created, a rule that a volume is deleted with its machine unless someone opts out, a retention chosen at the moment a schedule is written - lives at creation time. A team that only sweeps is signing up to sweep forever, because ordinary work produces residue continuously. ## Turning a candidate into a deletion safely 1. **Rank by monthly charge** and work the top of the list. The tail is real but it is not worth a conversation with an owner. 2. **Publish the row with its evidence**, not just the verdict, and give the owner a defined window to object. An owner who disputes "no connections in ninety days" usually knows about the quarterly consumer you could not see. 3. **Quarantine rather than delete.** Stop the machine, detach the volume, deny access to the store, disable the schedule. These are reversible in minutes and they surface any hidden dependency as a failure you can undo. 4. **Wait a full business cycle** with the resource quarantined, then delete - keeping one restorable copy of anything that held data, for as long as the cost of that copy is smaller than the cost of being wrong. 5. **Record the deletion and its evidence**, so the next sweep does not re-derive the same judgement from scratch. The discipline this buys is not mainly financial. A sweep that deletes on the strength of a guess breaks something once and is then never allowed to run again, which costs far more over a year than the residue it removed.

  • What is the most common false positive, and how do you avoid acting on it?
    Something used on a longer cycle than the window you measured - a month-end consumer, a quarterly rehearsal, a store read only during an audit. Measure across a full business cycle, publish the evidence to the owner before acting, and quarantine first so the discovery is a reversible failure rather than a lost resource.
  • Why is a sweep not a substitute for controls at creation time?
    Because it is detective, not preventive: it finds residue after it has been billed and does nothing to stop the next batch. Expiry on environments, delete-with-the-machine as the default for volumes, and retention chosen when a schedule is written reduce the production rate, which is what makes the sweep small.
  • How do you handle a candidate whose owner label is missing entirely?
    Treat the absence as its own finding. Escalate to whoever owns the account or the region rather than the resource, use creation records in the audit trail of management API calls to identify who made it, and quarantine with a longer objection window, since nobody will be watching for the failure.

saying these in an interview costs you the question

  • A cost report grouped by service will show the waste.
  • If nobody claims it in a week, delete it.
  • Sweep the regions we use; the others are empty.
  • A resource with no tag is by definition abandoned.
  • Finding the waste once fixes the problem.