skip to content

The dedicated circuit is finally delivered and the ledger must move off the tunnel — how do you sequence a cutover whose rollback is real?

level: principalimportance: nice to knowfreq 32%

answer

  1. reverse a preference, never a rebuild
  2. both paths live through the whole move
  3. one flow first, then the ledger
  4. a rollback nobody sized is fiction
  5. soak through a full business cycle

basics

~20 s

Run both paths live and move traffic by changing preference, never by tearing the old one down. Shift a low-risk flow first, soak it against agreed signals, then the ledger — and keep the tunnel configured, sized and up, because that is what rollback consists of.

solid answer

~50 s

A cutover between hybrid paths should be a preference change, because a preference change is the only step that reverses in seconds. So: bring the circuit up alongside the tunnel, prove it with probe traffic in **both** directions, then raise its preference for one low-risk flow and soak it long enough to see a peak. Agree the rollback trigger before the window — which signal, which threshold, how long you wait — because nobody negotiates that at two in the morning. Then move the ledger the same way and hold it through a full business cycle, month-end included, before anyone touches the tunnel. The rollback is real only if the tunnel is still configured, still up, and still sized for today's traffic rather than the traffic it was built for. The usual end state is not one path: it is the circuit preferred and the tunnel retained as the standby.

code

pseudocode · 21 lines
pseudocode
plan cutover(ledgerFlow):
    require circuit.isUp and tunnel.isUp
    require tunnel.capacity >= currentPeakLoad        // else rollback is fiction
    require filters.allow(circuit) and filters.allow(tunnel)

    step 1: raise preference of circuit for lowRiskFlow
    observe for one full peak: tailRoundTrip, replicationLag, errorRate

    if any observed signal outside agreedBand:
        lower preference of circuit for lowRiskFlow    // tunnel still configured
        stop and re-plan

    step 2: raise preference of circuit for ledgerFlow
    observe for one full business cycle

    if any observed signal outside agreedBand:
        lower preference of circuit for ledgerFlow
        stop and re-plan

    keep tunnel configured and sized as the standby path
    // decommissioning it is a separate, deliberate decision

go deeper

for a junior

Know that moving between two network paths is safest when both are live and the change is which one is preferred, so going back is one small step rather than a rebuild.

for a middle

Explain why preference is the reversible lever, why the return direction has to prefer the same path, and what has to remain configured for a rollback to work at all.

for a senior

Run it: probe traffic first, a low-risk flow across a peak, then the ledger, with agreed signals and a trigger armed in advance. Name the specific ways a rollback quietly stops existing.

for a principal

Decide how long the organisation runs two paths and what that time buys, who calls rollback, and whether the old path becomes the permanent standby — with an owner for its capacity and its exercise schedule.

## The cutover is a preference change, not a rebuild The temptation with a newly delivered circuit is to move everything at once and decommission the tunnel that has been "temporary" for three months. That design has no rollback: reversing it means rebuilding something during an incident, which is the worst possible moment to configure anything. The alternative is to run both paths simultaneously and let a preference decide which one carries traffic. Reversal is then a single, already-tested change, and it takes as long as the routing takes to settle. This is also why the cutover and the redundancy design are the same piece of work. If you keep the tunnel to make the rollback real, you have already built the two-path steady state, and the decision at the end of the exercise is simply not to dismantle it. ## Preconditions before anything moves 1. Both paths are up and both carry probe traffic, in both directions, with results from each side and not just from the platform's view. 2. The tunnel's capacity has been re-measured against **current** peak, not against the load it was sized for when it was stood up. Traffic grows quietly. 3. The signals that will judge the cutover are agreed and already being collected: tail round-trip time, replication lag, application error rate. A signal you start collecting during the window has no baseline. 4. The rollback trigger is written down as a threshold and a time box, with a named person who calls it. 5. Filtering rules on both sides admit both paths. Tightening them to the circuit is a later, separate change, precisely because it is the change that kills the rollback. ## The ways a rollback turns out not to exist - The tunnel endpoint was reused, repurposed or torn down "since the circuit is live now". - On-premises filtering was narrowed to the circuit's addresses in the same window, so falling back restores a path the firewall no longer trusts. - Name resolution was pointed at the new path and the answers were cached far longer than the rollback needs. - The tunnel was left up but never re-sized, so it exists and cannot carry the load it has to inherit. - The old path's addresses were handed to something else the day after. ## Both directions, or neither A hybrid cutover fails asymmetrically more often than it fails outright. Traffic leaves over the circuit while the return traffic still prefers the tunnel; anything doing stateful filtering in the middle sees a reply for a session it never saw opened, and drops it. So preference has to be set and verified on both ends, and the verification has to be a flow that completes, not a reachability test in one direction. ## Sequencing 1. **Probe only.** Both paths live, traffic still on the tunnel, the circuit carrying synthetic traffic and measured for a full day including the busy hour. 2. **One low-risk flow.** Something whose failure is visible but survivable. Soak it across a peak, against the agreed signals. 3. **The ledger.** Same mechanism, one preference change, with the rollback trigger armed and someone watching the agreed signals rather than the dashboard they like. 4. **A full business cycle.** For a ledger that means month-end, because month-end is when the traffic shape and the deadlines are different from every other day. 5. **Decide the end state deliberately.** Keep the tunnel as the standby — the usual and usually correct answer — or retire it as an explicit decision with its own risk acceptance, not as housekeeping. ## The judgment a lead actually owns Everything above is mechanics. The part that has no single right answer is how long you are willing to run and pay for two paths, and what you are buying with that time. Running dual-live for a month buys a rollback that works and a redundancy posture you would otherwise have to build later; running it for a quarter mostly buys inertia. The other open call is what happens to the tunnel afterwards: retaining it turns an unplanned survival story into a designed degraded mode, at the cost of a second link nobody looks at until it is needed — which is exactly why its capacity and its exercise schedule belong to someone by name.

  • Why move a low-risk flow before the ledger, rather than cutting everything at once?
    Because it separates two questions that a big-bang cutover asks at the same time: does the new path work at all, and does it work for this workload. The low-risk flow answers the first cheaply, over a real peak, with a rollback nobody will hesitate to use. What it cannot answer is whether the path meets the ledger's latency band, which is why the second step still gets its own soak.
  • The circuit is live and stable for two weeks. Why not decommission the tunnel?
    Because the tunnel is now the second path, and removing it returns the estate to a single physical chain — the exact posture that made the first outage possible. Retiring it should be an explicit decision with a named alternative standby, not the tidy-up at the end of a project. If it stays, someone owns re-sizing it and exercising it.
  • What makes a rollback trigger useful rather than decorative?
    A threshold on a signal that was already being collected before the window, plus a time box and a person who calls it. Without a baseline the threshold is guesswork; without the time box the team waits for improvement that never comes; without a named owner the decision escalates into the incident itself instead of ending it.

You do not close the old road on the day the new one opens. You reroute the signs, watch the traffic for a season, and only then decide whether the old road comes up.

saying these in an interview costs you the question

  • Decommissions the tunnel in the same window the circuit goes live
  • Treats the cutover as one big-bang move with no low-risk step
  • Sets path preference in one direction and tests only outbound reachability
  • Narrows firewall rules to the circuit before the soak period ends
  • Assumes the standby still fits traffic that grew since it was sized
  • Agrees the rollback trigger during the incident rather than before