You own moving a shipment-tracking API from six VMs onto a shared Kubernetes cluster; how would you phase the cutover so every step stays reversible?
answer
- same artifact, two platforms
- state outside both
- weights above both pools
- singletons move once
- soak before delete
basics
~20 sRun the same build on both platforms with state kept outside both, shift traffic in weighted steps from a layer above them while the VMs stay warm for rollback, hand over singletons like scheduled jobs exactly once, and decommission the VMs only after a soak.
solid answer
~50 sFirst I remove host assumptions on the VMs themselves, so both platforms run one artifact against the same external database, object store and session store. Then I deploy to the cluster with zero traffic and check capacity, probes, egress allowlists and dashboard parity, including a deliberate node drain. Traffic moves through a layer above both pools, DNS weights or an external load balancer, in steps such as 1%, 10%, 50%, 100%, each with an SLO gate, so rollback is a weight change. Singletons, the nightly archive `CronJob` and any consumer that must not double-process, move in one explicit handover: disable on the VMs, then enable in the cluster. I never migrate data in the same step as compute. After a full business cycle at 100%, the VMs scale to zero, and they are deleted only later.
go deeper
Know that a migration can move traffic gradually while the old servers keep running, so a problem is fixed by moving traffic back.
Explain where the traffic split lives, above both platforms, and why DNS caching slows both the shift and the rollback.
Show the operational detail: singleton handovers, failure rehearsal with real node drains, and gates defined before each step.
Weigh dual-running cost against risk, keep data and compute migrations separate, and decide when a big-bang cutover is the better call.
## Principles A platform migration is not a service redesign: the **shipment-tracking API** keeps its code, its API and its database. What changes is where it runs. That makes a reversible plan possible, provided three rules hold: - **One artifact on both sides.** The VMs and the pods run the same build and configuration, so a traffic shift tests the platform and nothing else. - **State lives outside both.** Database, object storage and session store stay where both platforms can reach them. Moving data in the same step as compute removes the rollback path. - **Every step is undone by the step before it.** If going back requires a restore or a rebuild, the step was too big. ## The phases 1. **Make the VMs cluster-shaped.** Move sessions to the shared store, uploads to object storage, logs to stdout, and deploy that version to the six VMs. This is the riskiest change and it happens on the known platform. 2. **Deploy dark.** Create the Deployment, Service and CronJob (suspended) on the **12-node cluster**, which also hosts GPU model serving. Confirm the API pods land on general nodes (GPU nodes typically carry a taint), that requests and limits match measured VM usage, and that egress to partners works from the new addresses. 3. **Rehearse failure.** Drain a node carrying API pods; this cluster takes about **13 minutes** per drain. Watch error rate, latency and PodDisruptionBudget behaviour before any user depends on it. 4. **Shift traffic in steps.** Move weight gradually, holding each step long enough to see a peak. 5. **Hand over singletons.** Disable the VM crontab entry, then unsuspend the CronJob. Queue consumers that must not run twice move the same way. 6. **Soak, then retire.** Stay at 100% through a full business cycle (month-end, a carrier peak), then scale VMs to zero, keep their images for a few weeks, and only then delete them and their firewall rules. ## Where the traffic split lives The split must sit **above** both platforms, because a Kubernetes Service only balances across pods. | Mechanism | Strength | Weakness | |---|---|---| | Weighted DNS records | no new infrastructure | caching delays every change; some clients ignore TTL | | External load balancer with both pools | instant, precise weights | the VM pool and cluster ingress must both be registered targets | | API gateway in front of both | per-route or per-customer weights | another component to operate during the migration | Lower DNS TTLs days before the first step if DNS is the mechanism; otherwise a rollback takes as long as the old TTL. ## Singletons and state Traffic weights do not apply to scheduled work. A job that runs on the VMs and in the cluster runs **twice**. Treat each singleton as its own tiny cutover with an owner and a time. The same applies to anything that holds a lease or lock: confirm the old holder has stopped before the new one starts. Keep the database migration, if there is one, as a separate project. When compute and data move together, a latency problem and a data problem look identical, and neither can be rolled back alone. ## Gates and rollback Define the gate before the step: - error rate and p99 latency no worse than the VM pool at the same weight; - no rise in partner rejections (a sign of a missed IP allowlist); - pod restarts and evictions at zero outside planned drains. Rollback is a weight change back to the VMs. That only works if the VM pool is still sized for full traffic, so do not shrink it during the shift. ## When not to phase Phasing costs weeks of dual running. For an internal tool with a maintenance window and no singletons, a scheduled big-bang cutover with a tested rollback can be the better call. The judgment is proportional: the more external callers, allowlists and scheduled jobs a service has, the more the phased plan pays for itself.
- When would you skip the phased plan and do a single scheduled cutover instead?When the service has few callers, a maintenance window is acceptable, there are no external allowlists or singletons, and rollback is still cheap. Dual running costs time, capacity and attention; for a small internal tool that cost can exceed the risk it removes.
- What would make you roll back after reaching 50% of traffic on the cluster?A breach of the gates agreed before the step: error rate or p99 latency worse than the VM pool at the same load, partner rejections rising, or unplanned restarts and evictions. Rollback is a weight change to the VMs, which is why the VM pool stays sized for full traffic until the soak ends.
- How do you stop the nightly archive job from running on both platforms during the transition?Create the CronJob suspended, then do a named handover: disable the VM crontab entry, confirm the last VM run finished, and unsuspend the CronJob before the next schedule point. If the job is not idempotent, add a lock or run marker so an accidental double run is harmless.
saying these in an interview costs you the question
- Moving the database into the cluster in the same step as the API is fine
- Enabling the CronJob before disabling the VM crontab is the safe order
- DNS weight changes reach every client immediately
- The VMs can be deleted the moment traffic reaches 100%
- A phased migration requires splitting the service into microservices first