skip to content

Retention and Cost Control

How teams keep log bills sane: tiering, retention policies, and deciding which logs never get stored. Interviewers ask because uncontrolled log volume is a universal pain and compliance sets hard floors.

on this pageshow

questions

4

What drives the cost of a centralized log platform, and which of those does shortening retention actually reduce?

level: middleimportance: must knowfreq 64%

answer

  1. The bill has more than one term
  2. Write-time work is paid once, up front
  3. Retention only touches stored bytes
  4. Replicas and search structures multiply volume
  5. Only pre-ingest levers move every term

basics

~20 s

Four things drive a log bill: bytes ingested, the parsing and index-building done on them at write time, stored bytes multiplied by replicas and retained days, and query load. Shortening retention shrinks only the third.

solid answer

~40 s

Split the bill into four terms. **Ingest** is charged on bytes accepted, plus the network and shipper CPU that move them. **Processing** is the write-time work — parsing, enrichment, and building whatever search structures the store keeps, which are themselves stored bytes. **Storage** is retained bytes times the replication factor, summed over every tier and every copy. **Query** is the scan and compute burned mostly by dashboards and alert rules re-running on a schedule, whether or not a human is watching. Cutting the retention window touches only the storage term, and only for the days removed: the ingest and processing for those bytes were paid once, at write time, and are not refunded. It also arrives slowly, as existing days age out. Moving the other terms means reducing volume before it is accepted.

code

pseudocode · 7 lines
pseudocode
monthly_bill =
      gb_per_day * 30            * ingest_rate
    + gb_per_day * 30            * expansion * process_rate
    + gb_per_day * retention_days * expansion * replicas * storage_rate
    + (alert_rules * evals_per_day + human_searches) * bytes_scanned * query_rate

# retention_days appears in exactly one of the four terms

go deeper

for a junior

Be ready to say that a log bill is charged on more than stored bytes: the platform also charges for accepting each record and for the work done on it when it arrives. Know that logs cost money continuously, not once.

for a middle

Explain the four terms and say which of them are sunk at write time. An interviewer expects you to estimate stored capacity as volume times expansion times replicas times days, and to state that shrinking the window leaves ingest and processing untouched.

for a senior

Show that you measure before cutting: rank streams by bytes and by whether anything has queried them, then attack volume at the source. Be ready to explain why a retention change takes weeks to show up on an invoice.

for a principal

Own the question of who pays. A per-team budget with visible ingested volume changes behaviour that a central cost-cutting exercise never does. Be ready to name the term you deliberately allow to grow, and what that buys.

A log platform's bill is not one number with one dial on it. It is four largely independent terms, and cost conversations go wrong when a team assumes the dial it can reach is attached to the term that dominates. ## The four terms 1. **Ingest** — the bytes accepted at the front door. A hosted platform usually charges this directly per gigabyte. Self-hosted, you pay it as network transfer, as the CPU the node-level log shipper burns batching and compressing, and as the capacity of the collection tier itself. This term is a function of raw volume and of nothing else. 2. **Processing** — the work done once per record at write time: parsing an unstructured line into fields, enrichment lookups that attach orchestrator or geographic metadata, and building whatever search structures the store maintains. Those structures cost twice: CPU to build, and then they are themselves stored bytes. For verbose records they can rival the size of the original text. 3. **Storage** — retained bytes multiplied by the replication factor, summed over every tier and every copy. The multipliers surprise people: eighteen gigabytes a day of raw text, expanded by write-time structures and held in two replicas, is a far larger billable figure than the number anyone quoted, and that is before snapshots. 4. **Query** — the scan, decompression and compute a search burns. The human part is small and bursty. The part that dominates is automated: alert rules and dashboard panels re-evaluating on a fixed interval, forever, each one a real query over real data whether or not anybody is looking at the answer. ## Why shortening the retention window disappoints Shortening how long data is kept removes the oldest days from term 3, and only from term 3. Those bytes were already ingested and already processed; that charge was paid once, at write time, and nothing refunds it. The saving is therefore bounded by the storage share of the bill multiplied by the fraction of retained days removed — and on an indexed, hot-heavy platform the storage share is frequently the smaller half. It disappoints twice, because it also arrives late: nothing is deleted when the setting changes, so the saving accrues only as existing days age past the new horizon. Teams reach for it first anyway, because it is a single setting needing no code change and no negotiation with a service owner. Every lever that genuinely moves the bill requires somebody to change what a service emits. ## What moves what | Lever | Ingest | Processing | Storage | Query | |---|---|---|---|---| | Shorten the retention window | no | no | removed days only | slightly | | Drop a class of record before it is shipped | yes | yes | yes | yes | | Sample a high-volume class at collection | yes | yes | yes | yes | | Pre-aggregate a class into a counter at the edge | yes | yes | yes | yes | | Index fewer fields of each record | no | yes | yes | queries cost more | | Move older data to a cheaper tier | no | no | rate only, not bytes | slower | | Reduce replicas on older data | no | no | yes | availability risk | | Lengthen an alert rule's evaluation interval | no | no | no | yes | The table states the point plainly: **only levers applied before ingest move all four terms.** Everything applied after the front door is a discount on one of them. ## Measure before you cut - Rank producers by bytes per day. The distribution is nearly always extreme — a handful of streams carry most of the volume. - Rank the same producers by demonstrated readers: which have been touched by a human search, or named by an alert rule, in the last quarter? - Join the two rankings. High volume with no readers is the free money. High volume with many readers is where you sample rather than drop. - Measure the expansion factor your store applies, so that a capacity forecast starts from billable bytes rather than from raw text. - Look at the automated query load separately from the human load. It is a fixed cost you chose, and it is usually invisible on the very dashboard that is causing it. ## A worked example A vinyl-record marketplace ships 18.4 GB of application and access logs a day into a fully indexed hosted platform with a 30-day window. Someone proposes cutting the window to 14 days. If ingest and write-time processing are 62% of the invoice, stored bytes 31%, and query the remainder, then removing 16 of the 30 days can take at most about 16% off the total, and only once the old days have aged out. Meanwhile, ranking by producer shows that 7.5 GB of the daily 18.4 GB is access lines for a readiness endpoint polled twice a second by 109 replicas, roughly 400 bytes each, and that nothing has queried them in ninety days. Discarding that one class before it is shipped removes 41% of the volume from all four terms at once, and it takes effect the day the rule ships. The retention change is not wrong; it is simply the smaller of the two, and it was reached for first because it was easier.

  • Your platform reports three times more stored bytes than the raw text you ship. Where did the extra come from?
    From replicas, from the search structures built at write time, and from additional copies such as snapshots and downstream exports. Compression pushes the other way, so the net expansion factor is a property of your store and your record shape. Measure it before forecasting capacity; never plan from raw volume.
  • A team halves its retention window and the invoice barely moves for a month. Why?
    Two reasons compound. Nothing is deleted when the setting changes — the saving arrives gradually as existing days age past the new horizon. And the storage term was the smaller share of the bill to begin with, so even the full effect is bounded by the storage share times the fraction of days removed.
  • Which log costs keep accruing after everyone has stopped looking at a service?
    Ingest and processing continue for as long as the service emits records, storage continues until the data is deleted, and query cost continues because alert rules and dashboards keep evaluating on their schedule. Only stopping emission or discarding at collection stops the first two; only deletion stops the third.

Retention is the size of the warehouse; ingest and indexing are the delivery crew you already paid. Renting a smaller warehouse does not get their wages back.

saying these in an interview costs you the question

  • Says a log bill is just storage, so retention is the only lever
  • Believes deleting old data refunds the ingest and indexing already paid
  • Forecasts capacity from raw volume, ignoring replicas and index expansion
  • Treats query cost as human-driven only, forgetting scheduled alert queries
  • Proposes cutting volume without first measuring which streams anyone reads
open as a page

In a tiered log store, what differs between a hot, a warm and a cold copy of the same data?

level: middleimportance: should knowfreq 47%

basics

~20 s

Tiers differ in the media the bytes sit on, whether the data is still directly searchable and how fast, and how many copies exist. Moving down a tier lowers the price per gigabyte; only deletion removes the bytes.

open as a page

In a log collection pipeline, which records do you discard or sample before ingest, and what makes that irreversible?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Discard or sample the classes with high volume and no demonstrated readers: repetitive health-endpoint access lines, duplicated stack traces, oversized fields. It is irreversible because those records never reach storage, so version the rules and count what you discard.

open as a page