Weekly-read application logs are moved to the coldest archive storage tier to save money — what goes wrong?
answer
- cheap rent, expensive reads
- the discount is funded somewhere
- retrieval is charged per gigabyte read
- restore first, then read
- weekly reads multiply the retrieval charge
basics
~20 sA colder tier discounts rent by charging separately for reads and by making them slow. Data read every week pays a retrieval charge every week and waits for a restore each time, so the bill usually goes up, not down.
solid answer
~50 sA storage tier is not simply cheaper storage; it is a different trade. Going colder cuts the per-gigabyte rent and pays for that cut on three other prices: a **retrieval charge** per gigabyte read back out, a **restore latency** before the bytes are readable at all, and a **minimum storage duration** billed per object. Data read weekly hits the first two roughly fifty times a year, so the retrieval charges quickly swamp the rent saving, and every read now has to issue a restore and wait instead of just reading. The tier is built for data that is written, kept because someone may one day ask, and almost never actually read. Age is only a good proxy for that when reads genuinely stop with age — which is true of a log archive and false of anything a dashboard still touches.
go deeper
Recall that a colder tier trades cheap rent for charged, slower reads. If anything reads the data regularly, it does not belong there.
Explain the four prices a tier carries — rent, retrieval charge, restore latency, minimum duration — and show with rough arithmetic how weekly reads cancel the rent saving.
Show the operational half: a read from the coldest tier becomes request-wait-read, so callers break before the invoice arrives. Say what you would measure on the prefix first.
Frame it as a promise being priced: the tier discount buys the provider a guarantee about your read pattern. Decide where that promise is structurally true and make that the default, rather than tiering by retention alone.
## What a colder tier is actually selling An object store normally offers the same bytes at several **storage tiers**. They do not differ on one price; they differ on four, and only the first one goes down: - **Rent** — charged per gigabyte per month. This is the headline number, and it drops sharply as you go colder. - **Retrieval charge** — charged per gigabyte read back out of the tier. The warm tier normally has none. Colder tiers do, and the colder the tier, the larger it is relative to a month's rent. - **Restore latency** — how long after you ask before the bytes are readable at all. Warm tiers are immediate. Intermediate cooler tiers are usually still immediate but charge for the read. The coldest tiers make you request a restore and wait, typically minutes to hours; some sell a faster restore at a higher charge. - **Minimum storage duration** — the tier bills a floor number of days per object even if the object is deleted or moved sooner. So the discount is real and it is funded out of the other three. "Colder" means *you are promising to read this rarely*. The provider prices the promise, and charges you when you break it. ## Why "read weekly" is the disqualifying fact | Access pattern | Warm tier | Coldest archive tier | |---|---|---| | Written, kept for a retention window, never read | rent only | cheapest by a wide margin | | Read a handful of times a year | rent only | rent plus a few retrieval charges | | Read every week | rent only | rent, ~52 retrieval charges a year, and a wait on every read | The arithmetic is the whole answer. Retrieval is priced per gigabyte *retrieved*, and providers set it at a meaningful fraction of a month's rent per gigabyte — that is exactly how the rent discount is funded. Read the data set once a month and a large part of the saving is gone; read it weekly and you are paying several times what the warm tier would have cost. Nothing about the data changed; only the price schedule you opted into did. ## The half that is not about money Even if the money worked, the **shape of a read changes**. In a warm tier, reading an object is one request that returns bytes. In the coldest tier, it is: issue a restore request for that object, poll or wait for it to complete, then read the temporary readable copy that the restore produced. Code written against the warm tier gets an error or a status, not data. A weekly report with a fixed window, a dashboard, an on-call query, a support tool — all of them break on this before anyone looks at the invoice. This is why "it is only a bit slower" is the wrong mental model: for the coldest tiers it is not slower, it is *not readable until you ask and wait*. ## The condition that makes a colder tier right A colder tier earns its place when three things hold at once: 1. Reads are genuinely rare — not "we think they are rare", but measured, or structurally guaranteed by what the data is for. 2. A read is allowed to be slow. An auditor who gets the answer tomorrow is fine; a page that renders now is not. 3. The object will stay there longer than the tier's minimum storage duration, so the rent saving is not cancelled by a floor charge. A log archive kept because a regulator may one day ask satisfies all three. A month of logs that an operations team still greps does not — it satisfies none of them. ## Doing it properly - **Tier on read frequency, expressed as an age.** An age-based transition rule is the only lever the store gives you, so it only works where age really does predict "nobody reads this any more". Confirm that before writing the rule. - **Split the data, not the tier.** Keep the recent window that is still queried in the warm tier and transition only the tail behind it. One rule with one age over a whole path prefix is what produces the weekly-read-from-archive situation. - **Keep a cheap warm index.** If something must be searchable weekly, keep a small summary or index warm and put only the bulk bytes cold. The retrieval charge then applies to the few objects the search actually points at. - **Measure before and after.** Record how many objects and gigabytes the rule will move, and what reads that path prefix currently serves per week. If that number is not near zero, the transition is a cost increase dressed as a saving. ## Where this is got wrong in review The recurring failure is treating the tiers as a single axis labelled "cheap to expensive" and moving everything as far down it as retention allows. They are not one axis: each step down trades away read economics and read availability for rent. Reading the tier's terms — retrieval charge, restore latency, minimum duration — before writing the rule takes minutes and is the entire skill here.
- What would you measure before moving a path prefix to a colder tier?How often objects under that prefix are read, and by what. Reads per object per month, the age of the newest object anything still touches, and whether any reader has a latency deadline. If reads do not fall away with age, the age-based rule is moving live data into a tier that charges for every read.
- The data must be both cheap and quick to read. What options are left?Not a colder tier. Reduce what you store rather than where: compress, aggregate, drop fields, or shorten the retention so the expiration rule removes it sooner. Alternatively split it — a small warm summary or index that serves the frequent reads, with the bulk bytes cold behind it.
- Is an intermediate cooler tier a safe middle ground for weekly reads?Safer on latency, not on price. Intermediate tiers usually keep reads immediate, so nothing breaks functionally, but they still levy a retrieval charge per gigabyte and still carry a minimum storage duration. Weekly reads of the whole set can still cost more than warm rent.
Self-storage in the next county is cheaper by the month, but you pay a fee and wait for a van every time you want to look inside a box. Fine for papers you may never open; terrible for the toolbox you use on Saturdays.
saying these in an interview costs you the question
- Treats a colder tier as simply cheaper storage with no other change
- Assumes reading from the coldest archive tier is immediate
- Believes retrieval is free because the rent went down
- Writes one age-based rule over data that is still queried
- Thinks restore latency is a network problem rather than a tier property