skip to content

As the lead, how do you set the deprecation window for a retired payload version, signal it, and decide when it is safe to switch off?

level: principalimportance: should knowfreq 36%

answer

  1. cost falls on two different parties
  2. window from the slowest caller's cycle
  3. signal machine-readably and to humans
  4. count by version and by caller
  5. brownouts find who ignores notices

basics

~20 s

Set the window from the caller population's slowest realistic cycle, not from habit; signal it both machine-readably in responses and through a human channel; and decide the shut-off from per-caller usage measured over a full cycle, escalating through warnings and brownouts rather than a single date.

solid answer

~50 s

There is no correct number of months - the window is a negotiation between your carrying cost and your callers' release cadence, and the only honest input is data. Publish the retirement at the moment the replacement ships, carry it machine-readably in every response of the old version (a deprecation flag and a date the client can act on) as well as in the human channels callers actually read. Then measure: request counts dimensioned by version *and by caller identity*, observed over at least one full business cycle, because the caller that will break is a quarterly batch job nobody remembers. Escalate rather than jump: notices, then scheduled brownouts of the old version that make the deadline felt, then failure. Decide the date from who is left and what you owe them, and say out loud what never removing it would cost every future change.

go deeper

for a junior

Recall that an old payload version is announced as going away on a stated date, and that callers need both a notice and time to move.

for a middle

Explain the mechanics: a machine-readable deprecation signal on the old version's responses, a stated date, and usage measured per version rather than guessed.

for a senior

Show the evidence discipline - counts per caller over a full business cycle, data at rest as well as live traffic, and announced brownouts to flush out silent callers.

for a principal

Own the trade: set the window from the caller population and the commitments you made, put a number on the carrying cost of never retiring, and publish a support policy so the next retirement is routine.

## Why this is a judgment call and not a rule A deprecation window trades two costs that fall on different people. Keeping the old version alive costs your team - every later change ships twice, every incident starts with "which version", every new engineer must learn both. Removing it costs your callers, who must schedule work they did not ask for. There is no formula that resolves that, which is why it is a lead's decision rather than an engineer's. What *can* be made rigorous is the evidence, the signalling and the escalation. ## Setting the window Derive it from the caller population, not from a number you have seen elsewhere: - **Release cadence of the slowest caller class.** Internal services on continuous delivery move in days; embedded or shipped-to-customer clients move in quarters or not at all. - **Business cycles.** A window that does not contain a month-end, a quarter-end and whatever your annual peak is has not tested the callers that only run then. - **Migration difficulty.** A rename costs a caller an afternoon; a change that forces them to store something new costs a project. The window should be proportional to the work you are imposing, and you should have done that work once yourself to know. - **Commitments already made.** Contractual or published support terms bound the window from below regardless of what the data says. ## Signalling Signal in **two registers**, because neither alone reaches everyone: 1. **Machine-readable, in every response of the old version.** A deprecation indication with a date, carried in the response metadata, so that a caller's own tooling can surface it without a human reading anything. Standard response-header fields exist for exactly this purpose. 2. **Human, in the channels callers actually use.** Release notes, direct contact to identified callers, and the documentation page they land on. Directly contacting the top talkers by name is worth more than every broadcast combined. The signal must state a *date* and what happens on it. "Deprecated" with no date is noise, and callers rationally ignore it. ## Proving the last old caller is gone This is the part that separates an answer from a guess. You need a request counter dimensioned by **version and caller identity**, and you need it to have been collecting for long enough: | Evidence | What it proves | |---|---| | Traffic share of the new version | Almost nothing - one rare caller can be the important one | | No support tickets mentioning the old version | Nothing; silence is not absence | | Zero requests on the old version over seven days | That no daily job uses it | | Zero requests per caller over a full business cycle | That nothing periodic is left, which is the real question | Two refinements matter. First, **retention**: if the old version's data is also written to storage or a log that is read later, callers are not the whole story - a reader of that data at rest can outlive every live caller. Second, **brownouts**: schedule short, announced outages of the old version before the final date. They convert callers who ignore notices into callers who file a ticket, which is the cheapest possible way to discover them - and they must be announced, short and off-peak, or they are simply an outage you caused. ## Escalation, not a cliff A workable sequence: 1. Announce with a date when the replacement ships. 2. Machine-readable deprecation on every old-version response from that day. 3. Direct outreach to identified callers, repeated as the date approaches. 4. Short announced brownouts in the final weeks. 5. Switch off, keeping the ability to restore for a short, pre-agreed period. 6. Remove the code once the restore window has passed. ## The cost of never deciding The failure mode is not switching off too early; it is never switching off. A version that nobody dares retire is paid for by every future change, forever, and the bill is invisible because it arrives as "everything here is a bit slower to do". Naming that cost - in engineer-weeks per year, out loud, to the people who will feel the migration - is part of the job. Publishing a support policy *before* the first retirement is what makes the second one routine.

  • Why is a seven-day window of zero traffic weak evidence for switching off?
    Because periodic callers are invisible in it. A month-end reconciliation, a quarterly report or an annual renewal job can be the only remaining user of a version and will never appear in a week. Evidence has to cover at least one full business cycle, per caller rather than in aggregate.
  • What does a brownout achieve that another notice does not?
    It reaches callers who do not read notices. A short announced outage of the old version produces an error somebody has to triage, converting an unknown caller into a ticket while the deadline can still be moved. It has to be announced, short and off-peak, or it is just an outage you caused.
  • How do you argue for a retirement date to stakeholders who only see the migration cost?
    Quantify the other side. Put a number on what the version costs per year - changes shipped twice, incidents that start with a version question, onboarding time - and set it against the one-off caller migration. Also name who currently pays: the team, invisibly and permanently, on every unrelated change.

saying these in an interview costs you the question

  • Picks a window by habit rather than from caller cadence
  • Treats a published notice as proof that nobody is left
  • Measures only the aggregate traffic share of the new version
  • Announces a deprecation with no date and no consequence
  • Assumes live callers are the only readers of that version
  • Never retires anything and calls that customer focus