skip to content

A daily Elasticsearch log index has 20 primary shards holding 400 MB each — how do you fix the sizing?

level: seniorimportance: should knowfreq 50%

answer

  1. Split the problem into existing and future indices
  2. The existing count cannot be edited
  3. Retention may solve half of it for free
  4. Stop letting the calendar decide index size
  5. Verify with shard store sizes afterwards

basics

~20 s

You cannot change the primary count of existing indices, so fix it in two directions: shrink and force-merge the already-created ones down to one primary, and change the index template plus the rollover trigger so future indices are created with far fewer shards and roll on primary shard size rather than on the calendar.

solid answer

~50 s

Each shard is a full Lucene index, so 20 shards of 400 MB per day is roughly 8 GB spread over 20 shards where one would do — multiplying heap, file handles, cluster-state entries and per-search tasks for no benefit. Existing indices cannot be resized in place, so handle them with `_shrink` to a single primary followed by a force merge, and let retention delete the rest. For new data, change the index template so the pattern creates one or two primaries, and stop rolling strictly by calendar day: rolling over on a primary-shard-size condition means each index reaches a healthy size regardless of how traffic fluctuates, so a quiet weekend does not produce another set of near-empty shards. Verify afterwards with `_cat/shards` sorted by store size and by re-checking total shards against node heap.

code

json · 9 lines
json
{
  "index_patterns": ["app-logs-*"],
  "template": {
    "settings": {
      "index.number_of_shards": 1,
      "index.number_of_replicas": 1
    }
  }
}

go deeper

for a junior

Recall the hard constraint that starts the answer: an existing index's primary shard count cannot be changed, so any fix means creating differently-sized indices.

for a middle

Explain why 400 MB shards are wasteful — fixed per-shard overhead and one search task per shard — and name shrink plus a template change as the two levers.

for a senior

Give the full runbook and the judgment call: shrink and force-merge closed indices where it pays, change the template, switch the rollover trigger from calendar to primary shard size, then verify with shard sizes and thread-pool rejections.

for a principal

Treat it as a standard rather than an incident: shard sizing and rollover conditions belong in shared component templates, with fleet-wide monitoring of shards per GB of heap so no team can quietly create the problem again.

## Reading the symptom Twenty primaries at 400 MB each means about 8 GB of primary data per day. With one replica that is 40 shard copies per day for 8 GB of data. Over 30 days' retention that is 1,200 shard copies for 240 GB — a volume a handful of shards could hold comfortably. This is the textbook oversharded log cluster, and it usually arises one of two ways: someone copied a shard count from a much larger deployment into an index template, or the number was sized for peak volume that never materialised. The costs are the standard ones: per-shard heap and file handles, a bloated cluster state that the master must publish on every change, and — most visible to users — a search over a week of data fanning out to hundreds of shard tasks, each doing a millisecond of real work. ## What you cannot do You cannot update `index.number_of_shards` on an existing index; it is a final setting, because the primary count is an input to the document routing formula. So there is no in-place fix. Anything that changes the shard count creates a new index. ## Fixing what already exists For indices no longer being written to — which, for daily logs, is everything but today's — the cheap operation is `_shrink`: 1. Block writes (`index.blocks.write: true`) and relocate a copy of every shard onto one node with allocation filtering. 2. Wait for green, then `POST /<index>/_shrink/<index>-shrunk` with `index.number_of_shards: 1`. 3. Force merge the shrunken index down to a small number of segments — it is read-only now, so this is safe and improves compression and search speed. 4. Swap aliases atomically, verify counts, delete the source. Whether this is worth doing for every historical index is a judgment call. If retention is 30 days, the pragmatic answer is often: shrink nothing, fix the template today, and let the oversharded indices age out over the next month. Shrinking is worth the effort when retention is long, when the shard budget is genuinely at risk, or when historical searches are slow. ## Fixing the future — the real fix Change the index template that governs the pattern so new indices are created with a shard count matched to the volume. At roughly 8 GB per day and a target in the tens of gigabytes per shard, one primary per index is plenty — and it means a single index can then cover several days. That leads to the second and more important change: **stop creating an index per calendar day**. A daily index forces the shard count to be right for the average day, and it is wrong on every quiet day and every spike. Roll over on a size condition instead — a primary shard-size trigger creates a new index once the largest primary reaches the target, so each index lands in the healthy band regardless of traffic. Combine it with a max-age condition as a backstop so slow-moving data streams still roll periodically, and keep retention driven by age. One caution: shard count is set when the index is created, so a template change only affects indices created after it. Force a rollover, or wait for the next one, before expecting the new setting to take effect. ## Verifying After the change, re-check: - `GET _cat/shards?v&s=store` — the distribution should no longer be a long tail of sub-gigabyte shards. - `GET _cluster/stats` — total shard count against total heap, aiming well under roughly 20 shards per gigabyte of heap per node. - Search latency on a typical multi-day query, and search thread-pool rejections, which should fall as the fan-out shrinks. ## What a weak answer looks like The common wrong answers are: "update the shard count on the existing indices" (impossible — it is a final setting), "add more replicas so the shards get bigger" (replicas are copies; they change nothing about primary size and make the shard count worse), and "reindex everything into a new index" (correct but far more expensive than shrink, and unnecessary for data that will be deleted by retention anyway). The strong answer separates the two timelines — existing indices versus future indices — and puts the emphasis on the template and the rollover trigger, because that is what stops the problem recurring.

  • Why is rolling over on primary shard size better than rolling over daily for this workload?
    A calendar trigger fixes the time window and lets the data volume vary, so shard size swings with traffic and every quiet day produces undersized shards. A size trigger fixes the shard size and lets the time window vary, which is the property you actually want. Keeping a max-age condition alongside it ensures low-volume streams still roll eventually.
  • Would you shrink all 30 days of existing indices, or just fix the template?
    Usually just fix the template if retention is short — the oversharded indices delete themselves within the retention window, and shrinking each one is a multi-step operation per index. Shrink when retention is long, when you are close to per-node shard limits, or when historical searches are visibly suffering from the fan-out.
  • After updating the index template, why do new indices still have 20 shards?
    Shard count is fixed at index creation, so a template change applies only to indices created after it. The currently-written index keeps its old settings until the next rollover. Force a rollover if you want the new sizing to take effect immediately rather than at the next scheduled roll.

saying these in an interview costs you the question

  • Proposes updating number_of_shards on the existing indices
  • Adds replicas expecting the primary shards to grow
  • Reindexes 30 days of logs that retention will delete anyway
  • Changes the template and expects existing indices to change
  • Keeps one index per calendar day and calls the problem solved

context