skip to content

In a tiered log store, what differs between a hot, a warm and a cold copy of the same data?

level: middleimportance: should knowfreq 47%

answer

  1. Tiers trade query speed for storage price
  2. Ask what stays directly searchable
  3. Restoring a cold copy costs time and money
  4. A transition copies bytes, then deletes them
  5. Tiering lowers the rate, not the volume

basics

~20 s

Tiers differ in the media the bytes sit on, whether the data is still directly searchable and how fast, and how many copies exist. Moving down a tier lowers the price per gigabyte; only deletion removes the bytes.

solid answer

~40 s

The hot tier accepts writes, sits on fast local storage with its search structures resident and usually several replicas, and answers interactively. A warm copy is read-only on cheaper media with fewer replicas, still directly searchable but in seconds rather than milliseconds. A cold or archive copy is compressed bytes in object storage, often without resident search structures, so it answers only via a slow scan or after a restore measured in minutes to hours and billed as retrieval plus temporary capacity. A transition between tiers is a copy followed by a delete, not a pointer flip: both copies are billed while it runs and the I/O competes with live ingest. Two traps follow. Tiering discounts the rate but removes no bytes, and a query spanning a boundary runs at the slowest tier's speed.

code

json · 9 lines
json
{
  "stream": "catalogue-access",
  "tiers": [
    { "name": "hot",  "from_age_days": 0,  "answer_within": "1s",   "replicas": 2 },
    { "name": "warm", "from_age_days": 7,  "answer_within": "60s",  "replicas": 1 },
    { "name": "cold", "from_age_days": 45, "answer_within": "4h",   "replicas": 1 }
  ],
  "delete_after_days": 395
}

go deeper

for a junior

Know the vocabulary: recent logs sit on fast storage and are searchable immediately, older logs are moved to cheaper storage and answer more slowly or must be restored first. Do not confuse moving data with deleting it.

for a middle

Explain what actually differs across tiers — media, whether search structures stay resident, replica count and query latency — and describe a transition as a copy followed by a delete, with both copies billed while it runs.

for a senior

Show that you set tiers from how fast an answer is needed rather than from age alone, and that you have checked what a query spanning a tier boundary does to latency and to the bill.

for a principal

Own the policy shape: per-stream classes with an owner, a stated time-to-answer at each age band and a deletion date, rather than one global age rule. Be ready to say what you accept losing when a restore takes hours.

Tiering is the practice of keeping the same records on progressively cheaper storage as they age, accepting progressively worse access in exchange. The names differ between platforms; the gradient does not. ## What actually differs across a tier | Property | Hot | Warm | Cold / archive | |---|---|---|---| | Accepts new writes | yes | no | no | | Media | fast local storage | cheaper, slower storage | object storage | | Search structures resident | yes | usually | often not | | Directly searchable | interactively | seconds to minutes | slow scan, or only after a restore | | Copies typically kept | two or more | one or two | one, plus the provider's durability | | Cost per gigabyte-month | highest | middle | lowest | | Time to first answer | sub-second | seconds | minutes to hours | Three of those rows carry most of the weight. **Whether the tier still accepts writes** decides whether its data can be compacted and made read-only. **Whether the search structures are resident** decides whether a query can touch it at all without preparation. **Time to first answer** decides whether the tier is any use for the question you are actually going to ask of the data. ## A transition is a copy, not a pointer flip Moving a day of records down a tier is real work: 1. read the bytes out of the source tier; 2. re-encode them for the destination — usually compact and compress harder, sometimes discard the resident search structures; 3. write them to the destination and verify; 4. update the catalogue so a query knows where that time range now lives; 5. delete the source copy. Consequences follow directly from that list. Both copies exist and are billed for the duration of the transition. The I/O competes with live ingest on the same hardware, so a large backlog of transitions is felt as ingest latency rather than as a storage event. A transition that fails partway can leave orphaned bytes that nothing queries and everything bills. And a small amount of catalogue metadata stays in the fast tier permanently, because the platform must still know what exists and where it went. ## Tiering discounts the rate; it does not reduce the volume This is the most common misunderstanding on the subject. An 18.4 GB-per-day stream kept for thirteen months is roughly 7.3 TB whether it sits hot, warm or cold. Tiering changes what each of those terabytes costs per month; it does not remove a single byte. **Only deletion removes bytes.** If the underlying problem is that too much is being kept, or that too much was accepted in the first place, tiering buys time rather than a solution. The corollary matters for planning: a tier plan and a deletion date are two different decisions, and a policy that states only the first will grow forever at a slowly improving unit price. ## Tier by required answer speed, not by age alone Age is a proxy for how urgently data will be wanted, and it is often a poor one. - A high-volume access stream is queried heavily for a day and almost never afterwards. It can move down quickly and aggressively. - A low-volume administrative stream might be queried once a quarter, but when it is queried the answer is wanted in minutes. Its volume makes fast storage trivially cheap, so tiering it saves almost nothing and costs restore latency exactly when that hurts. - A stream feeding week-over-week comparisons must stay directly searchable for at least the comparison window, or every one of those queries pays a restore. So the policy is per stream class: state the required time-to-answer at each age band, then pick the tier that meets it. Writing the requirement down also gives you something to hold the platform to later. ## The cross-boundary query trap A query whose range spans a tier boundary runs at the speed of the slowest tier it touches, and on some platforms it silently triggers a restore or a full scan of the older span. That turns a harmless-looking change — widening a dashboard's default range, or an alert rule that looks back further after a tuning session — into a slow, expensive query that then repeats on a schedule forever. Two defences: keep dashboard and alert ranges inside the fast tier by default, and make crossing a boundary an explicit, visible act rather than something a default time picker can do by accident. ## What restoring actually costs Restoring older data is not a single number. It has a latency (how long until the first result), a monetary cost (retrieval charges plus temporary capacity in the fast tier to hold what you brought back), and a scope decision (restore the whole range, or only the streams you need from it). A tier whose restore takes hours is entirely adequate for an audit request that arrives with a week's notice and useless in the first thirty minutes of an outage. Decide which of those you are buying when you set the age threshold, not while you are waiting for the restore.

  • Your dashboards get slower after a tiering policy ships, although nothing else changed. What happened?
    Their default time ranges now cross a tier boundary, so each query runs at the speed of the slowest tier it touches, and on some platforms it triggers a scan or a restore of the older span. Fix it by keeping default ranges inside the fast tier and making a cross-boundary query an explicit, visible choice.
  • A stream is only 40 MB a day but must be answerable within minutes at any age. Where does it live?
    In the directly searchable tier for its whole retention. Its volume makes that trivially cheap, while tiering it buys almost nothing and costs restore latency precisely when the answer is wanted. Tier by the required time-to-answer for each stream class, not by age applied uniformly across the platform.

Hot storage is the filing cabinet beside the desk, warm is the box in the basement you can still open, and cold is the off-site vault that answers tomorrow morning.

saying these in an interview costs you the question

  • Thinks moving data to a cold tier deletes it or reduces retained volume
  • Assumes a cold copy answers at the same latency as recent data
  • Treats a tier transition as a metadata pointer flip with no I/O cost
  • Applies one age-based tier rule to every stream regardless of how it is queried
  • Forgets that replicas are counted separately within each tier