Two components one hop away serve the same 2 MB value, one from RAM and one from a warm disk engine - why is the latency gap small?
answer
- access time versus transfer time
- the ratios are access-time ratios
- same bytes over the same link
- crossover near tens of kilobytes
basics
~10 sThe famous ratios are ratios of access time, roughly constant per operation. Both paths still move two megabytes over the same link - about 1.6 ms - which dwarfs any microsecond difference in starting.
solid answer
~40 sA read's latency is access time plus transfer time. Access time is what it costs to begin getting the bytes - a memory access, a device access, a lookup - and it is roughly constant per operation. Transfer time is proportional to the payload and is bounded by memory bandwidth, device throughput and, once a network is involved, the link rate. The media ratios everyone quotes are access-time ratios. On a two-megabyte value, serialising the payload onto a ten-gigabit link costs on the order of 1.6 milliseconds on both paths, while the access-time difference is measured in microseconds - a few percent of the total. Even a cold device access, hundreds of microseconds, is a minority of it. The medium has not changed; the question has stopped being about the medium.
go deeper
Remember that reading something big and reading something small are not the same cost. Speed comparisons between media are about how fast a read starts, not how fast the bytes move.
Explain the split: access time is per operation and transfer time is per byte, so a large payload pushes the total into a term both components pay at the same rate.
Work the numbers in the room - megabytes against a link rate, microseconds of access against milliseconds of transfer - and say which regime the workload is actually in before accepting any ratio.
Use it to redirect a review: on large payloads the medium is not the deciding factor, so the argument should be about bandwidth, ownership of the authoritative copy and what happens when a component is unavailable.
## Two different costs wearing the same word The latency of a read is really two costs added together. 1. **Access time** - what it costs to begin getting the bytes: a memory access, a device access, a lookup on a server, the first packet of a round trip. It is roughly a constant per operation and barely depends on how much is being read. 2. **Transfer time** - what it costs to move the bytes once they are moving: bounded by memory bandwidth, device throughput and, when a network is in the path, the link rate. It is proportional to the payload. Every memorable memory-against-disk ratio is a ratio of **access time**. That is why it shows through clearly on a small read and quietly stops mattering on a large one: for a large payload, both components spend most of their time on a cost that is nearly identical for both. ## The arithmetic on a two-megabyte value Take the same two-megabyte value served two ways, each one network hop from the caller. | cost | in-memory component | warm disk engine | |---|---|---| | access on the server | around a hundred nanoseconds | resident page: comparable; cold page: tens to hundreds of microseconds | | reading the payload locally | memory bandwidth, gigabytes per second | memory or device bandwidth, gigabytes per second | | moving it across the network | about 1.6 ms on a ten-gigabit link | about 1.6 ms on the same link | Serialising two megabytes onto a ten-gigabit link takes on the order of 1.6 milliseconds before anything else is counted. Against that, an access-time difference measured in microseconds is a few percent of the total. Even a **cold** device access of several hundred microseconds is a minority share. The ratio the slogan promises has collapsed toward one - not because the media became equal, but because the question stopped being about them. ## Where the crossover sits A rough way to hold it: the medium's advantage shows through while the payload is small enough that **moving** the bytes costs less than **beginning** to move them. With access-time differences in the tens to hundreds of microseconds and link rates in gigabits per second, that crossover lands somewhere around tens of kilobytes on typical hardware. Below it, per-operation costs dominate and the medium argument is live. Above it, the bytes dominate and it is not. The exact number moves with hardware; the structure does not. This has a practical edge for anyone comparing components: - a benchmark that reads small entries and a production workload that reads large values are measuring different regimes; - a ratio quoted from the first will not survive in the second; - when payload sizes vary widely, aggregate latency is dominated by the large reads even if most requests are small, so an average ratio describes neither regime; - headroom on the link, not the medium, is what a large-value workload actually consumes. ## What this does and does not license It does not say large values belong on one component rather than the other, and it says nothing about what a large value does to the other callers sharing a component - that is a separate subject with its own answer. It says one thing precisely: **on large payloads, the medium is not the reason to prefer one component over the other**. If a design is choosing between them for large objects, the honest grounds lie elsewhere - who holds the authoritative copy, what happens when the component is unavailable, what the bandwidth and the storage actually cost. The corollary is useful in reviews. When someone justifies a component on the media ratio, ask how large the values are. If they are kilobytes, the argument is at least in the right regime and the remaining question is the baseline. If they are megabytes, the ratio has been quoted from a regime the workload is not in, and the discussion should move to bandwidth and ownership. ## What varies - **Where the components sit.** The collapse argument assumes both pay the same network. If one is inside the calling process and the other is across a link, the large-payload comparison is memory bandwidth against link rate, and that is not close. - **How a change to a large value is made.** Stores in this class differ in whether part of a value can be altered on the server or whether the whole value must be moved out and back. Where it is the latter, every update to a large value is a large-payload transfer in both directions, which puts the workload firmly in the regime described here. - **The engine's residency.** A warm engine's access is a memory access, a cold one's is a device access. Both are small against a multi-megabyte transfer, which is what makes this argument unusually robust: it holds whether the engine is warm or not.
- At what payload size does the medium's advantage still show through?While per-operation cost dominates - small entries, roughly up to tens of kilobytes on typical hardware. There the access-time difference is a real share of the total. Above that the bytes dominate and both paths converge, and because payload sizes usually vary, the honest answer names the regime the workload actually lives in rather than one number.
- Does a cold source change the large-payload conclusion?Barely. A cold device access of several hundred microseconds is still a minority of the roughly 1.6 milliseconds it takes to move two megabytes over a ten-gigabit link. That is why this argument is robust: it holds whether the source is warm or cold, unlike the small-read comparison, which turns entirely on that question.
saying these in an interview costs you the question
- Applies the small-read latency ratio to multi-megabyte payloads
- Forgets both paths move the same bytes over the same link
- Assumes bandwidth is free once the first byte has arrived
- Confuses random-access latency with sustained sequential throughput
- Quotes an average ratio across wildly different payload sizes