skip to content

An exit plan allots one week to copy a 400 TB media archive off a platform — what makes that estimate wrong?

level: seniorimportance: should knowfreq 52%

answer

  1. the schedule, not just the bill
  2. sustained throughput, not the link rating
  3. gigabytes times eight, over the line
  4. cold data must be retrieved before it moves
  5. the copy window floors the dual-run months

basics

~20 s

Sustained throughput sets the copy time: 400 TB at a steady 1 Gbit/s takes about 37 days. Archived objects must first be retrieved to a readable tier, and both the retrieval and the outbound gigabytes are charged.

solid answer

~40 s

Two things break the week. First, arithmetic: 400 TB is 400,000 GB, which is 3,200,000 gigabits; at a genuinely sustained 1 Gbit/s that is 3.2 million seconds, roughly **37 days** — and the figure to use is throughput the copy holds on real data, not the link's rating. Second, the archive is probably not readable where it sits: objects in a cold storage tier normally have to be retrieved before anything can read them out, which is a separate charge and often a separate wait, on top of the per-gigabyte outbound charge for the bytes themselves. The copy window then sets the cutover date, and the cutover date sets how many months both platforms are billed — so an underestimate here is paid twice.

code

pseudocode · 14 lines
pseudocode
gigabytes    = 400000                  // a 400 TB archive
gigabits     = gigabytes * 8           // 3,200,000

// the middle term is measured, not quoted from the link's rating
sustainedGbps = measureOnSample(representativeObjects)   // say 1.0

seconds = gigabits / sustainedGbps     // 3,200,000
days    = seconds / 86400              // about 37

if days > plannedCopyDays then
    report "the copy, not the rewrite, sets the cutover date"
    overlapMonths = ceil(days / 30) + cutoverMonths
else
    overlapMonths = cutoverMonths

go deeper

for a junior

Remember that moving a large data set is measured in days or weeks, and that the number comes from throughput the copy actually holds rather than the speed printed on the link.

for a middle

Do the arithmetic aloud: gigabytes times eight gives gigabits, divided by sustained gigabits per second gives seconds. Then say which data must be retrieved from a cold tier before any of it can move.

for a senior

Show that the copy window floors the dual-run window, and say how you would verify the copy, handle writes arriving during it, and stop it from starving the production service.

for a principal

Treat the window as what decides whether the move is one quarter or three, and say what evidence you would require before letting anyone commit the business to a cutover date.

## The copy is a schedule before it is a charge Exit estimates treat the outbound copy as a charge — gigabytes times a per-gigabyte rate — and then assume it happens in the background. For a long media archive the more dangerous half is the calendar. A copy that takes five weeks rather than one week does not change the per-gigabyte total at all, but it adds a month to the window in which **both** platforms are billed in full, and it pushes every dependent date behind it. ## The arithmetic, done out loud Take the acquisition case: a media transcoding service and its archive must leave one platform so two estates can be merged onto one. The archive is 400 TB. - 400 TB is **400,000 GB**, which is **3,200,000 gigabits** (eight bits to the byte). - At a genuinely sustained **1 Gbit/s**, that is **3,200,000 seconds** — about **37 days**. - At a sustained **10 Gbit/s** it is about **3.7 days**; at **200 Mbit/s** it is about **185 days**. Every one of those numbers is an assumption, and the assumption that matters is the middle term. A link is rated at a speed; a copy sustains a different one, because it shares the egress path with production traffic, because the platform's bulk transfer path applies its own pacing, and because small objects cost a request round trip each regardless of how fat the pipe is. An archive of many small objects can run an order of magnitude below the same link's rate on large ones. **Measure it on a representative sample before promising a date.** ## Cold data is not readable data Archive tiers are cheap per gigabyte-month precisely because the object is not sitting ready to be read. Providers differ in how this is presented — some make retrieval an explicit request with a wait before the object becomes readable, some price fast and slow retrieval differently, some apply a minimum storage duration so that removing an object early still bills the remaining days — but the shape is common: **before an archived byte can leave, somebody pays to make it readable.** So the data line has two components, not one: | Component | Charged on | Applies to | |---|---|---| | Retrieval | gigabytes restored to a readable state | only what sits in a cold tier | | Outbound transfer | gigabytes leaving the platform | everything that moves, hot or cold | And the retrieval is also a schedule item: if retrieval runs in batches, the copy cannot outrun the rate at which objects become readable. ## What else eats the window 1. **Verification.** A copy nobody checked is not a migration. Comparing the source and destination inventories, and re-copying what does not match, is real time and re-charged bytes. 2. **The delta.** The archive keeps receiving new material while the bulk copy runs. The normal pattern is one bulk pass followed by repeated incremental passes over what changed, each pass charged again for what it moves, until the remaining delta is small enough to close in a short freeze. 3. **Failures.** Long transfers fail partway. Only the re-sent bytes are charged again — not the whole copy — but they are charged, and they extend the window. 4. **Contention.** The copy competes with the production service on the same egress path, so either the service degrades or the copy is throttled to protect it. Almost always it is throttled, which is a decision that should be in the plan rather than discovered. ## Why the underestimate is paid twice The copy window is the floor under the overlap, and the overlap is two monthly bills. A four-week slip on the copy is four weeks of the current platform plus four weeks of the target platform, plus whatever the delayed consolidation was worth. That is why the honest version of this question is not *what does the transfer cost* but *how long does it take, and what is running while it does*. The corollary is a planning rule: for a volume large enough that the copy is measured in weeks, the copy is the critical path and the rewrite is not. Start the bulk pass as early as the data will allow — it can begin long before the replacement code exists, since the bytes are not what is being rewritten — and let the incremental passes carry it forward to the cutover. Teams that sequence the copy after the rewrite discover the constraint on the week they wanted to switch, which is the most expensive week to discover it.

  • What do you do about writes that arrive while the bulk copy is running?
    Run the bulk pass first, then repeated incremental passes over what changed since the last one, until the remaining delta is small enough to close inside a short write freeze at cutover. Each pass is charged again for what it moves, so the plan should say how many passes you expect. If the source changes faster than a pass completes, the copy never converges and you need a freeze rather than another pass.
  • Why does the copy window show up twice in the exit estimate?
    Once as outbound gigabytes charged per gigabyte, and again as calendar time, because both platforms are billed until the copy finishes and the last consumer moves. A slower link therefore costs money twice: it leaves the per-gigabyte total unchanged but lengthens an overlap that is billed by the month.

A self-storage warehouse rents space cheaply by the month, then charges a handling fee and a loading slot when you take the pallets back out. The monthly rent was never the whole price of keeping things there.

saying these in an interview costs you the question

  • Plans the copy from the link's rated speed rather than measured sustained throughput.
  • Assumes archived objects can be read out without being retrieved first.
  • Treats the bulk copy as background work with no effect on the cutover date.
  • Believes a retried transfer is free because the first attempt was already paid for.
  • Ignores that the copy shares an egress path with the production service.