A team books its move off a managed engine as a few weeks of engineering work - why does the dataset's size set the schedule instead?
answer
- effort is elastic, bytes are not
- the schedule is a copy, not a task
- measure what changes while you copy
- volume over copy rate equals hours
- book the tail, not the transfer
basics
~20 sBecause the schedule is set by how long the bytes take to copy and how much of the data changes while they are copying. Engineering effort can be parallelised across people; the copy cannot, it grows with the dataset, and its tail decides the cutover window.
solid answer
~50 sEffort is the wrong unit because it is the one part of the move you can add people to. Three measured numbers set the plan instead: the volume that must move, the sustained copy rate you actually achieve on the real dataset, and the rate at which the source changes while the copy runs. The first two give the initial copy duration; the third turns that duration into a **tail** of accumulated changes that must be drained before you switch. What the business has to approve is the drain plus verification - not the whole copy - and if the tail accrues faster than it drains, no cutover window exists at that copy rate and the shape of the plan must change rather than its dates. This is why the size of the stored data, not the difficulty of the work, is the number to measure first.
code
pseudocode · 16 lines// assumptions, all measured on a rehearsal against the real dataset
datasetGiB = 40000
copyRateGiBPerHour = 250
changeGiBPerDay = 300
initialCopyHours = datasetGiB / copyRateGiBPerHour // 160 h
accruedDuringCopy = changeGiBPerDay * (initialCopyHours / 24) // 2000 GiB
tailDrainHours = accruedDuringCopy / copyRateGiBPerHour // 8 h
changeGiBPerHour = changeGiBPerDay / 24 // 12.5
if changeGiBPerHour >= copyRateGiBPerHour:
report "tail never drains at this copy rate - change the shape, not the date"
else:
report "cutover window = tailDrainHours + verification + switch"
// 12.5 < 250, so this branch fires: window is about 8 h plus verificationgo deeper
Recall that moving a large dataset takes time proportional to its size, and that the data keeps changing while the copy runs.
Explain the three measured inputs - volume, sustained copy rate, change rate - and how they produce an initial copy duration and a tail to drain.
Show that you rehearse on the real dataset, book the tail rather than the transfer, and recognise the arithmetic that says no cutover window exists at all.
Decide what the organisation will accept: how much data moves at all, whether the move goes in slices, and what outage the business is genuinely willing to underwrite.
## Why effort is the wrong unit An estimate in engineer-weeks describes the part of the move that is elastic. Writing the new configuration, describing the new machines in whatever form your team already uses, adapting the application's connection handling, writing the verification queries - all of that can be split across people and compressed. The part that cannot be compressed is the **copy**, and the copy is set by physics and volume rather than by headcount. A migration off a managed tier is therefore scheduled like a shipment, not like a feature. This is what **data gravity** means in practice. It is not a mystical property of large datasets and it is not a technical impossibility: it is the plain fact that the bigger the stored data, the longer any move takes and the narrower the set of ways it can be done. Note the direction - data gravity here is a **time** constraint on the plan. The per-gigabyte charge for moving bytes appears separately on the bill and is a different conversation. ## The numbers that actually set the plan 1. **Volume to move.** Not the whole store - the part that must be live on the other side on day one. Retention you can trim and cold data you can leave behind never enter the arithmetic. 2. **Sustained copy rate, measured on the real dataset.** Not a headline figure, not a laptop benchmark. Measure it, because every other number derives from it. 3. **Change rate during the copy.** Whatever is written to the source while the initial copy runs becomes the tail you must drain afterwards. 4. **The window the business will accept.** This is what you are really negotiating, and it covers the drain plus verification plus the switch, not the whole copy. | Unit of estimate | What it predicts | What it misses | |---|---|---| | Engineer-weeks | how long the building takes | the copy, which no extra person shortens | | Total copy duration | when the bulk transfer ends | that writes continued the whole time | | Tail drain plus verification | the real cutover window | nothing - this is the number to book | ## Working it through Take invented but consistent figures, and state them as assumptions. Suppose the live portion is about 40,000 GiB, the rehearsal achieves a sustained 250 GiB per hour, and the source accumulates around 300 GiB of change per day. - Initial copy: 40,000 / 250 = **160 hours**, near enough seven days. - Change accrued during those 160 hours: 300 x (160/24) = about **2,000 GiB**. - Draining that tail at the same rate: 2,000 / 250 = about **8 hours**, and the tail accrues at 12.5 GiB per hour against a drain rate of 250, so it converges comfortably. - So the window to negotiate is roughly those 8 hours plus verification and the switch - not the seven days. Booking seven days of downtime because the copy takes seven days is the classic mistake in the opposite direction. Now change one assumption: if the source accrued 8,000 GiB per day, the tail would grow at over 330 GiB per hour against a 250 GiB per hour drain, and **no cutover window exists at that copy rate**. That is not a scheduling problem. It is a signal that the shape must change. ## What actually shrinks the window - **Raise the sustained rate** - more parallel streams, a shorter path, a copy that does less work per byte. - **Move less** - trim retention first, leave cold data behind and move it afterwards, drop what nobody reads. - **Move in slices** - migrate a tenant, a shard or a table group at a time, so each slice carries its own small tail and a bad slice is recoverable. - **Keep the source intact and read-only until verification passes**, so the rollback is "point back" rather than "restore". A standby replica does not give you this: replication copies a bad write faithfully, so what protects the cutover is the untouched source or a retained backup. - **Verify by reading back**, not by trusting the copy tool's exit code: row counts, checksums over samples, and the business queries that matter. ## What the rehearsal must prove Rehearse on the real dataset, because the whole plan rests on one measured number. The rehearsal should produce the sustained copy rate actually achieved, the tail it accrued, the drain time, and the verification duration. Everything else - the window you ask for, the go/no-go threshold, the rollback point - is derived from those. A plan whose copy rate was assumed rather than measured is not a plan; it is an intention with dates attached. ## How to answer this in an interview Say plainly that the work is elastic and the copy is not, then name the three measured numbers and show that the window is the tail, not the transfer. The strong signal is knowing when the arithmetic says no window exists at all, and that the response is to change the shape - move less, move in slices, or raise the rate - rather than to ask for a longer outage.
- What do you do when the tail accrues faster than it can be drained?No cutover window exists at that copy rate, so the shape has to change rather than the dates. Raise sustained throughput, reduce what must move by trimming retention or leaving cold data behind, move in slices so each carries its own small tail, or pause the writers for a bounded period the business has explicitly agreed. Asking for a longer outage instead is how these migrations fail publicly.
- Which single number should the rehearsal produce?The sustained copy rate actually achieved on the real dataset. Everything else in the plan derives from it: the initial copy duration, the size of the tail that accrues behind it, the drain time, and whether the drain fits the window the business will approve. A rate that was assumed rather than measured invalidates every date downstream of it.
Moving house: the packing can be split across as many helpers as you like, but the lorry still has one journey, and everything you keep using until the last minute has to go on a second trip you schedule around.
saying these in an interview costs you the question
- Estimates the move in engineer-weeks and never measures a copy rate
- Books the whole transfer as downtime instead of the tail
- Plans a single dump and load for a dataset written to hourly
- Believes more engineers shorten a transfer bounded by throughput
- Never rehearses the copy on the real dataset before booking a window
- Treats a standby replica as protection against a bad cutover