skip to content

Why is stored volume alone a poor measure of data gravity when you price the data dimension of lock-in?

level: middleimportance: should knowfreq 47%

answer

  1. volume is one factor, not the measure
  2. multiply by the unit move cost
  3. charged on the way out
  4. elapsed copy time, not disk size
  5. a live dataset changes under the copy

basics

~20 s

Data gravity is stored volume multiplied by what moving one unit costs — the per-gigabyte charge leaving the platform, the elapsed copy time, the effort of extracting it, and reconciling a dataset that keeps changing. Volume prices none of those.

solid answer

~50 s

Volume is only half of the product. Data gravity is roughly `volume × unit move cost`, and the unit cost is where two estates of identical size stop resembling each other. It contains the transfer charge, which is levied on traffic leaving and normally not on traffic arriving; the elapsed time of the copy, set by the throughput you can actually sustain out rather than by how much sits on disk; the extraction effort, which is engineer-months rather than gigabytes when the only reader is a proprietary interface; and the drift you have to reconcile because a live dataset keeps changing while the copy runs. A large cold archive that is mostly expired and readable directly can be lighter than a much smaller live dataset that only one proprietary interface can read. That is also why the dimension is a *priced* cost, not a technical impossibility.

code

pseudocode · 23 lines
pseudocode
function dataGravity(dataset):
    movableGiB = dataset.storedGiB - dataset.deletableGiB
    transferCharge = movableGiB * ratePerGiBOut
    copyDays = movableGiB / sustainedGiBPerDayOut

    if dataset.readableByOpenInterface:
        extractionEffort = 0
    else:
        extractionEffort = engineerMonthsToBuildExport(dataset)

    driftCost = copyDays * dataset.dailyChangeGiB * reconcilePerGiB

    return transferCharge
         + valueOfTime(copyDays)
         + valueOfEffort(extractionEffort)
         + driftCost

// same storedGiB, different answers
coldArchive.deletableGiB        = most of storedGiB
coldArchive.readableByOpenInterface = true      // takes the zero branch

liveLedger.deletableGiB         = 0
liveLedger.readableByOpenInterface  = false     // takes the engineer-months branch

go deeper

for a junior

Remember that the data dimension of lock-in is priced on the way out, not on the way in, and that the amount stored is only part of what makes a move expensive.

for a middle

Be able to decompose the unit move cost out loud: per-gigabyte charge, sustained throughput and therefore elapsed time, extraction effort where the data is only readable one way, and reconciling the change that lands during the copy.

for a senior

Show that you estimate it on a real estate — which portion is actually movable after expiry, what throughput you have measured rather than assumed, and how you would validate that what arrived equals what left.

for a principal

Treat gravity as an exposure that compounds silently. Set retention so volume stops growing unattended, and decide deliberately whether a proprietary interface may become the only path to a dataset.

## Data gravity in one line **Data gravity** is the tendency of a large stored dataset to keep everything else near it — because moving it is expensive. As a number it is close to `volume × unit move cost`. Volume is the easy half and the half everyone quotes. The unit cost is where the real variance lives, and it is why two datasets of the same size can differ by an order of magnitude in how hard they are to take elsewhere. ## What the unit cost is actually made of - **The transfer charge.** Traffic leaving a platform is normally metered and charged per gigabyte, while traffic arriving normally is not. That asymmetry is the whole reason the exit is the expensive direction even though filling the store was nearly free. The rate is published by the provider; what you need for an estimate is the rate times the volume that genuinely has to move. - **Elapsed time.** How long the copy runs is set by the throughput you can sustain on the way out — the export path, the link, how many parallel streams the source tolerates — not by how large the dataset looks on disk. A copy measured in weeks is a schedule risk and an extra window in which two systems must both be kept correct. - **Extraction effort.** If the only way to read the data is an interface specific to this platform, someone has to write and test an export before a single byte moves. That is engineer-months, and it is the part of the data dimension that most resembles the code dimension. - **Drift.** A live dataset keeps changing while the copy is running. You pay for that either with a catch-up stream, a reconciliation pass over the window, or a freeze that the business has to agree to. - **Validation.** Proving that what arrived equals what left — counts, checksums, re-running a sample of the reporting and comparing — is real work, and it is the line most often missing from an estimate. ## Same size, very different gravity | Dataset | Volume | Unit move cost | Gravity | |---|---|---|---| | Cold archive of old events, mostly past its retention, readable in a format anything can parse | large | low: most of it can be deleted rather than moved, and what remains streams out directly | low | | Live transactional history readable only through a proprietary interface, with a residency rule attached | far smaller | high: an export has to be built, the copy must be reconciled against continuing writes, and the destination is constrained | high | The table is the answer to the question: the first row has more bytes and less gravity. ## What lowers gravity, in the order that pays 1. **Delete or expire what nothing reads.** The cheapest gigabyte to move is the one you do not move. In an old estate this single step often removes most of the volume. 2. **Keep at least one copy in a form that is not platform-specific.** If the bytes can be read without the interface that wrote them, the extraction effort collapses toward zero. 3. **Do not let one proprietary interface become the only path to the data.** That is the choice that converts a data problem into an engineering project. 4. **Re-measure it periodically.** Gravity is the one dimension that grows as a side effect of the system working, so a figure from last year is optimistic by construction. ## The directions that get stated backwards - **Charged leaving, not arriving.** If you claim a transfer bill, say which boundary the bytes crossed and in which direction; the exit direction is the charged one, which is precisely why exits are expensive. - **Priced, not impossible.** Data gravity is a bill and a calendar, not a technical limit. Describing it as "we physically cannot move it" is the failure mode that stops anyone from costing it. - **Staying against leaving.** The recurring storage bill is the cost of staying. Gravity is the cost of leaving. They are different numbers with different owners, and the recurring bill being small says nothing about the exit bill. ## How this sits inside the lock-in inventory When you size the four dimensions of lock-in, this is the row that is most often quoted wrongly, because stored volume is the one figure everybody already has on a dashboard. Quoting it alone overstates the gravity of a bloated archive and badly understates the gravity of a small dataset that only one interface can read.

  • Can a small dataset ever have higher gravity than a much larger one?
    Yes, and it is common. A modest live dataset readable only through a proprietary interface needs an export built and tested before anything moves, needs reconciling against writes that continue during the copy, and may be pinned by a residency rule. A large archive that is mostly past its retention and readable directly is mostly deleted rather than moved.
  • What is the cheapest way to reduce data gravity without moving anything?
    Delete or expire what nothing reads. It removes volume from every term of the product at once — transfer charge, copy time, reconciliation and validation — and it is a retention decision rather than an engineering project. Keeping one copy in a form that is not platform-specific is the second cheapest, because it removes the extraction effort.
  • Why does elapsed copy time belong in the estimate rather than just the charge?
    Because time is what turns a line item into a project. A copy that runs for weeks forces a window in which both sides must be kept correct, extends the period the organisation is paying attention, and creates the drift that has to be reconciled. Teams that price only the per-gigabyte charge routinely underestimate the move by the largest factor.

A warehouse charges nothing to fill and something for every box you carry back out. What makes moving expensive is how many boxes you kept and how awkward they are to lift, not how big the warehouse was.

saying these in an interview costs you the question

  • Equates data gravity with the number of terabytes stored
  • Assumes copying out is cheap because storing was cheap
  • Believes transfer is charged arriving rather than leaving
  • Ignores that a live dataset changes during the copy
  • Calls expired data that nothing reads heavy to move
  • Describes data gravity as a technical impossibility