skip to content

A nightly cleanup job meant for staging deleted live ledger data because both environments sit in one account — what allowed it, and which fix closes the route?

level: seniorimportance: should knowfreq 46%

answer

  1. what was the last thing standing?
  2. the filter, not the platform, decided
  3. unlabelled meant old, not scratch
  4. preventive against detective control
  5. credential that cannot name production

basics

~20 s

One account held both environments, so the job's filter rather than the platform decided what it could delete, and resources predating the label escaped that filter. Only a credential that cannot name production closes the route; an alert merely reports it.

solid answer

~50 s

The job held credentials valid for the whole account, so every resource in both environments was *nameable*; the only thing narrowing it was its own filter, and the filter was wrong — ledger resources created before the labelling rule existed carried no `environment` label, so a match on missing-or-staging swept them in. The lesson is that the separation was never enforced: it was a property of the script. A tighter pattern moves the failure rather than removing it, and an alert on bulk deletion is **detective** — it shortens recovery without blocking the call. The fix that closes the route is structural: production in its own account, and the cleanup job holding a short-lived credential issued for the non-production account, which cannot name a production resource however its filter behaves. Keep the alert as well; it still catches mistakes made inside production.

code

pseudocode · 9 lines
pseudocode
credentials = credentials_for(account)        # one account holds BOTH environments
items = list_all_items(credentials)           # returns staging AND live ledger items

for each item in items:
    is_staging   = item.name starts_with "stg-"
    is_unlabeled = item.label.environment is missing   # true for pre-convention LIVE items

    if is_staging or is_unlabeled:
        delete(item, credentials)             # this branch fires on the live ledger

go deeper

for a junior

Notice that the job could delete live data because its credentials covered the whole account. The filter inside the script was a wish, not a restriction.

for a middle

Explain why the unlabelled resources were the trap: a convention applies only from the day it is introduced, and anything older is a silent exception no filter anticipates.

for a senior

Separate preventive from detective clearly, sequence the remediation from stopgap to structural, and say which class of mistake still survives the account split.

for a principal

Frame it as a layout decision with a price: what the organisation buys by making the boundary structural, what it pays in duplicated setup and access, and where you would not pay it.

## What actually happened The two environments shared one account, kept apart by a prefix and an `environment` label. The nightly cleanup job authenticated once, listed everything it could see, and deleted whatever its filter matched. Three facts combined: 1. **The credential was account-wide.** Everything in the account — staging and the live ledger — was nameable by that job. 2. **The filter was the only narrowing.** The job matched on a `stg-` prefix *or* on the absence of an `environment` label, on the assumption that unlabelled meant scratch. 3. **The assumption was historically false.** Ledger resources created before the labelling convention existed carried no label at all, so the second clause of the filter selected live data. The platform did exactly what it was asked. There is no authorization step anywhere in that sequence that could have refused, because nothing in the request said which environment the caller was entitled to act on — the caller was entitled to the whole account. ## Why the filter was the only control This is the general shape, not a detail of this job: - A convention is evaluated **by the caller**, so a caller with a bug evaluates it wrongly and the call still succeeds. - Conventions are retroactive only if someone backfills them; anything older than the rule is a silent exception. - Filters are written against the state of the estate on the day they were written, and estates grow new shapes. - A job that runs unattended has no second pair of eyes at the moment that matters — review happened weeks earlier, against a different estate. ## Fixes, and what each one actually closes | change | prevents or detects | what it costs | |---|---|---| | Tighten the pattern | neither — it moves the next failure | an afternoon, and false confidence | | Backfill the missing labels | neither — removes today's instance only | a day, then ongoing discipline | | Alert on bulk deletion | detects, after the call has succeeded | a threshold to tune, a pager to answer | | Require a confirmation step | prevents interactive slips, not unattended runs | friction on every manual run | | Production in its own account, with a per-environment credential | prevents: the job cannot name production | the setup and the access, built twice | Only the last row changes what is *possible*. The others change what is *likely*, which is worth having but is a different claim, and mixing the two up is the most common weak answer to this question. ## Sequencing the fix 1. Stop the job and establish what it deleted from the audit trail of management calls, which records who called what and when. 2. Narrow the credential the job runs with immediately, as a stopgap, even before the accounts are split. 3. Stand up the second account, rebuild the non-production estate in it, and move the job there with a short-lived credential issued for that account only. 4. Point the production-side equivalent of the job — if one is needed at all — at production as a separate deployment with its own credential, so the two can never be confused by a configuration value. 5. Keep the bulk-deletion alert. It now catches the class the boundary cannot: a mistake made inside production by something that legitimately belongs there. ## What the boundary still does not cover - **A mistake inside production.** Production's own automation, correctly credentialled, can still delete production data. The split bounds cross-environment reach, not intra-environment error. - **A human with access to both.** Someone holding credentials in each account can make the same mistake twice; that is an access-review problem. - **A deletion that was authorized and unwanted.** If the job was supposed to run and the instruction was wrong, no boundary helps. ## The interview point What separates a strong answer here is refusing to accept the framing that this was carelessness. The question to ask is *what was the last thing standing between that call and the data* — and if the honest answer is "a string comparison inside the script", the layout is the defect and the script is the trigger. Preventive controls change the set of possible calls; detective controls change how fast you learn. Name which one you are proposing, and do not describe an alert as if it had stopped anything.

  • The team proposes an alert on bulk deletions instead of splitting the accounts — what does that buy, and what does it not?
    It buys time-to-know: you learn within minutes instead of at the morning stand-up, which shortens recovery and limits how much follows the first mistake. It does not buy prevention — the call has already succeeded when the alert fires, and a threshold low enough to catch a slow drip pages constantly. Keep it, because it catches mistakes made inside production that no boundary between environments can, but do not accept it as the answer.
  • Once production has its own account, which credential should the cleanup job hold?
    One issued for the non-production account alone, minted per run and short-lived rather than a long-lived key living in the job's configuration. Then the job's reach is bounded by the account rather than by the correctness of its filter. If production genuinely needs the same cleanup, run it as a separate deployment with its own credential, so no configuration value can point one job at the other environment.
  • What would you check first to bound what was lost?
    The audit trail of management API calls for the account: it records which principal called which operation against which resource and when, so you can reconstruct the exact set the job deleted rather than guessing from the filter. That set is what you then compare against what the job was supposed to touch, and the difference is both the recovery scope and the evidence for the layout change.

saying these in an interview costs you the question

  • Blames the engineer rather than the layout that let one filter decide everything.
  • Says a stricter naming pattern would have prevented it.
  • Offers an alert on bulk deletion as a preventive control.
  • Assumes resources created before the labelling rule carry the label anyway.
  • Thinks running the job under a person's credentials instead would be safer.