skip to content

How would you design a data stream and ILM policy to keep 400 GB of logs a day for 13 months on a fixed budget?

level: principalimportance: should knowfreq 38%

answer

  1. Access profile decides everything else
  2. Retention granularity equals rollover cadence
  3. Cheap tiers trade latency for storage
  4. Object storage carries the long tail
  5. Never delete without a verified backup

basics

~20 s

Work backwards from query patterns: keep a few days hot on fast disk, weeks warm and force-merged, months as snapshot-backed cold or frozen indices where object storage carries the bulk, and delete on schedule with a snapshot guard.

solid answer

~50 s

Start from the access profile, not the volume. Ask who queries data at each age and how fast they need an answer — dashboards over 24 hours, incident investigation over a week, compliance retrieval over a year — because that dictates which tier each age band can tolerate. Then size the rollover so each backing index is a manageable unit: roll on largest-primary-shard size with an age ceiling, so retention granularity stays roughly daily. Hot keeps a few days on fast disk with replicas for search throughput. Warm holds several weeks: read-only, force-merged, often fewer replicas. Cold and frozen carry the long tail as searchable snapshots, so the thirteen-month obligation is mostly an object-storage bill rather than a cluster of disks — frozen keeps only a local cache and accepts multi-second queries. Delete at 13 months, guarded by `wait_for_snapshot` against an SLM policy so nothing expires unbacked-up. Then verify the assumptions with real query telemetry rather than guesswork.

code

json · 43 lines
json
// PUT _ilm/policy/logs-13m
{
  "policy": {
    "phases": {
      "hot": {
        "actions": {
          "rollover": {
            "max_primary_shard_size": "50gb",
            "max_age": "1d",
            "min_docs": 1
          }
        }
      },
      "warm": {
        "min_age": "3d",
        "actions": {
          "readonly": {},
          "forcemerge": { "max_num_segments": 1 },
          "set_priority": { "priority": 50 }
        }
      },
      "cold": {
        "min_age": "30d",
        "actions": {
          "searchable_snapshot": { "snapshot_repository": "logs-repo" }
        }
      },
      "frozen": {
        "min_age": "90d",
        "actions": {
          "searchable_snapshot": { "snapshot_repository": "logs-repo" }
        }
      },
      "delete": {
        "min_age": "395d",
        "actions": {
          "wait_for_snapshot": { "policy": "nightly-logs" },
          "delete": {}
        }
      }
    }
  }
}

go deeper

for a junior

Recall the pieces this design is built from: a data stream with rollover, phases that move data to cheaper storage as it ages, and a delete phase at the retention limit.

for a middle

Explain why retention granularity equals rollover cadence, why shard count must be right in the template, and what force merging and dropping replicas actually save at each phase.

for a senior

Be ready to defend concrete boundaries against a real access profile, wire the delete phase to a snapshot guard, and explain the interaction between SLM retention and mounted searchable snapshots.

for a principal

Own the argument connecting access patterns to spend: what each transition buys, what the thirteen-month promise obliges operationally, how you validate boundaries with telemetry, and what changes when volume or retention requirements shift.

## Start with access, not volume 400 GB a day is a distraction until you know who reads it. The design question is: for each age band, what latency is acceptable and how often is it queried? A workable interrogation is to ask for the last thirty days of query telemetry and bucket it by the time range requested. In almost every logging system the answer is heavily skewed — the overwhelming majority of queries cover the last day or two, a thin tail covers the last month, and access beyond that is a handful of investigations and audits per quarter. That shape is what makes tiering pay. If instead the audit team runs heavy aggregations over the full thirteen months every week, frozen is the wrong answer and the budget conversation has to happen up front rather than after the first slow query. The second question is what "keep for 13 months" actually obliges. Retention for regulatory retrieval, where a request may take hours to answer, is a completely different engineering problem from retention for interactive search. It is worth establishing whether the requirement can be met by snapshots in a repository at all, since a snapshot that can be restored on demand satisfies many retention obligations at a fraction of the cost of anything mounted. ## Shape the stream Separate streams per meaningful class of data, not one stream for everything. Different retention obligations, different mappings and different query patterns should not be forced to share a lifecycle — and separation also isolates a mapping explosion in one noisy service from everything else. Conversely, resist a stream per microservice: thousands of tiny streams multiply shards, cluster state and policy evaluation cost for no benefit. Size the rollover so each backing index is a comfortable unit of work: a size trigger on the largest primary shard, with an age ceiling so a quiet stream still rolls on a predictable cadence, and a `min_docs` guard so quiet periods do not manufacture empty generations. Roughly daily generations are a good default because retention granularity can never be finer than the rollover cadence — you delete whole backing indices, so a weekly index means expiring a week at a time. Get the shard count right at the template level: it is fixed for each generation's life, and the only cheap way to change it is to edit the template and let the next rollover apply it. ## The phases **Hot** carries the write index and the last few days. Fast storage, replicas for both redundancy and query throughput, no merging or read-only games — this data is still changing. **Warm** takes over once writes stop. Mark read-only, force merge to few segments, drop the recovery priority so hot indices restore first after a restart, and consider `best_compression` since this data is now written once and read occasionally. Denser, cheaper nodes; still local disk, so latency stays low. **Cold** for the month-plus band: mount as a fully mounted searchable snapshot so the repository copy provides redundancy and local replicas can go, roughly halving local storage for that band while keeping query speed close to warm. **Frozen** for the long tail: a partially mounted searchable snapshot, where only a bounded shared cache lives on local disk and blocks are fetched from object storage on demand. This is what makes thirteen months affordable — the tail becomes an object-storage line item rather than provisioned disk — and the price is first-query latency measured in seconds. Confirm that whoever queries year-old data can live with that before committing. **Delete** at the retention boundary, with `wait_for_snapshot` naming an SLM policy so an index is never deleted before a snapshot taken after it entered the delete phase exists. Also decide deliberately whether deleting a mounted index should delete its underlying snapshot; keeping the snapshot beyond the searchable retention is a common and cheap way to satisfy an archival obligation. ## The snapshot side An SLM policy defines schedule, repository, which indices, a name pattern and a retention block with expiry and minimum and maximum counts. Two things are worth stating explicitly at design time. First, retention in SLM and retention in ILM are separate systems that must not contradict each other: an SLM retention that expires snapshots still underpinning mounted cold or frozen indices is a data-loss event, not a cleanup. Second, snapshot policies deserve the same monitoring as the ingest path, because a `wait_for_snapshot` guard turns a silently failing SLM policy into an ever-growing set of undeletable indices — the failure is safe but not free. ## Verify, then defend the numbers Write down what each transition buys and costs — hardware saved in warm, replicas removed in cold, disk avoided in frozen, object-storage spend added — and validate the boundaries against real query telemetry after a month. Expect to move the frozen boundary at least once. The design is not the policy JSON; it is the argument connecting the access profile to the spend, and the ability to say what changes when the daily volume doubles or the retention requirement moves to three years.

  • Why must SLM retention be longer than the ILM delete age when cold and frozen indices are mounted from snapshots?
    Because a mounted searchable snapshot is not a copy — the snapshot in the repository is the data. If SLM expires a snapshot that still backs a mounted index, that index loses its storage, which is data loss rather than cleanup. Retention windows in the two systems must be reasoned about together, with the snapshot side deliberately outliving anything mounted from it.
  • What breaks if you choose weekly rollover instead of daily for this volume?
    Retention granularity collapses to a week, because expiry means deleting whole backing indices. Every backing index also becomes seven times larger, so shards blow past a sensible size unless you multiply the shard count, and force merges and tier migrations become long, lumpy operations. Daily generations keep each unit of lifecycle work small and make the retention boundary precise.
  • How would you validate that the frozen boundary is set correctly after a month in production?
    Bucket real query telemetry by the age of data requested and compare it against the phase boundaries. If a meaningful share of queries reaches into frozen, users are paying multi-second latency you did not intend and the boundary should move later. If almost nothing touches cold, that band is buying query speed nobody uses and can move to frozen earlier, freeing local disk.
  • The daily volume doubles to 800 GB. What in this design has to change?
    Rollover fires roughly twice as often at the same size threshold, so generation count and lifecycle work double — usually fine. What needs review is shard count per generation in the template, hot-tier disk for the same number of days, and the object-storage bill for the long tail. The phase ages themselves are driven by the access profile and should not change just because volume did.

saying these in an interview costs you the question

  • Picks tier boundaries before looking at query patterns
  • Puts all services in one stream with one retention
  • Deletes indices without verifying a snapshot exists
  • Lets SLM expire snapshots that still back mounted indices
  • Assumes frozen queries are as fast as warm ones

context