skip to content

In a large Splunk estate, when is precomputing an expensive search into a summary index or an accelerated report the wrong call?

level: principalimportance: nice to knowfreq 36%

answer

  1. Pay once, on a schedule
  2. Freshness and disk buy speed
  3. Some statistics do not recombine
  4. A gap nobody is told about
  5. Precomputation can hide a defect

basics

~20 s

A summary index stores a scheduled aggregation's results as small events you search instead of raw data; report acceleration has Splunk maintain equivalent precomputed data for a qualifying saved search. Both buy speed with freshness, storage and recurring load.

solid answer

~50 s

Precomputation moves cost from the reader to a recurring background job. A **summary index** is the manual form: a scheduled search aggregates a slice of raw events and writes the rows into a separate, much smaller index that later searches read. You own its schedule, its granularity and its backfill — and its gaps, because a failed run or a batch of late data leaves a hole nothing announces. **Report acceleration** is the managed form: for a saved search whose expensive part is an aggregation over events straight out of an index, Splunk maintains precomputed data and uses it transparently, reading raw events for whatever it does not cover. Precompute when the same aggregation is read far more often than the underlying question changes. Do not precompute exploratory work, anything that needs the raw events for drill-down, or a search that is only expensive because its scope or its ingest hygiene is wrong.

go deeper

for a junior

Recall that Splunk can store the results of a scheduled aggregation so later searches read a small precomputed set instead of raw events, and that those results are only ever as fresh as the last run.

for a middle

Explain the two standing costs: a summary is stale between builds and holds only what the aggregation kept. Be ready to say which statistics recombine across buckets and which quietly do not.

for a senior

Show that you try the cheaper levers first — narrowing the retrieving stage, filtering before the aggregation, fixing what is ingested — and that you monitor the build job, because a silent gap in a summary is worse than a slow search.

for a principal

Treat precomputation as a shared budget rather than a per-team switch. Own who may accelerate, how build cost is weighed against reads saved, when the honest answer is to reduce ingest instead, and what stays raw so investigations remain possible.

## What precomputation is, and the two forms it takes An expensive search is expensive because it reads a great many events and then reduces them to a few numbers. Precomputation does the reducing ahead of time, on a schedule, so the reader pays for a small result set instead of for the raw data. The cost does not disappear; it moves from the person running the search to a recurring background job, and it changes shape on the way. **A summary index** is the manual form. A scheduled search runs over a slice of raw data — the last hour, say — computes an aggregation, and writes the resulting rows into a separate Splunk index as events in their own right. That index is tiny compared with its source, so a search across months of it is cheap. You own everything about it: the schedule, the granularity of each bucket, which fields are kept, and the backfill when something goes wrong. **Report acceleration** is the managed form. For a saved search whose expensive part is an aggregation applied to events coming straight out of an index, Splunk maintains precomputed data alongside the indexes over a chosen range and uses it automatically when the search runs, reading raw events only for the portion the precomputed data does not cover. A related mechanism accelerates a data model instead of a single search, precomputing across the fields the model defines so that many searches can be served from one body of precomputed data. ## What precomputation costs | Cost | Summary index | Report acceleration | |---|---|---| | Freshness | stale by up to one build interval | same, with the uncovered edge read from raw data | | Storage | a second index you size and retain | precomputed data held alongside the indexes | | Recurring load | a scheduled search you own and monitor | maintenance work the platform schedules for you | | Failure mode | a silent gap where a run failed or data was late | falls back to raw data, so it gets slow rather than wrong | | Flexibility | any shape you can express | only searches that qualify | | Maintenance | a definition change breaks comparability with old rows | changing the search rebuilds what was precomputed | Two rows deserve spelling out. **Late data** is the quiet one: an event indexed after the window containing its timestamp has already been summarised will not appear in that summary, and nothing announces it. **Non-decomposable statistics** are the sharp one: counts, sums, minima and maxima recombine correctly across buckets, but a distinct count does not and neither does a percentile. Averaging hourly averages weights every hour equally regardless of traffic, and the median of hourly medians is not the daily median. When you need those, store something that recombines — the components of the calculation, the set of keys, an approximating structure — rather than the answer itself. ## When precomputation is the wrong answer 1. **The question is still moving.** Exploratory work whose shape changes from week to week produces summaries nobody reuses. 2. **The freshness requirement is tighter than the build interval.** An alert that must fire within minutes cannot be served from something rebuilt hourly. 3. **The expense is really scope or ingest hygiene.** A search that is slow because it reads an unbounded range, or because the data was onboarded with no useful metadata to filter on, has a defect. Precomputing it conceals the defect and pays for it every interval, indefinitely. 4. **You need the events.** A summary row is a number; an investigation needs the underlying events. Either keep enough keys to re-query the raw data, or accept that the trail stops at the summary. 5. **It is read rarely.** A build every fifteen minutes serving a report somebody opens twice a month is a straight loss. 6. **There is no capacity left to build with.** Search concurrency is finite, and every acceleration competes with the humans using the platform. ## Deciding for an estate At platform scale this stops being a per-search question. Treat precomputation as a shared budget with three rules worth defending: - **Cheaper levers first.** Narrow the retrieving stage, filter before the aggregation, and fix what is being ingested. Most requests for acceleration turn out to be one of those three. - **Every acceleration has an owner and a review date.** The build cost recurs and is invisible; the saving on reads is not automatically larger than it. - **Measure the ratio.** Recurring build cost against reads actually served is a number you can compute, and it is the only honest basis for saying no to a team. On a vinyl-record marketplace running 41 services, one team accounted for roughly seventy per cent of the indexed volume and also owned 23 of the platform's 31 accelerated reports. Their build jobs consumed enough of the shared search capacity that interactive searches queued during business hours. The fix was not more hardware: retiring 14 reports nobody had opened in ninety days, and switching off the debug logging that produced most of the volume, restored the platform without adding a node. That is the judgement this question is asking for — precomputation is a way of spending capacity, and somebody has to own how much of it gets spent.

  • Which statistics cannot simply be combined across the buckets of a summary?
    Distinct counts and exact percentiles. Counts, sums, minima and maxima recombine cleanly, and an average recombines if you keep the sum and the count rather than the average itself. For distinct counts keep the key set or an approximating structure; for percentiles keep the distribution rather than the answer, because a median of medians is not a median.
  • A precomputed summary has a two-hour hole in it. What are the likely causes?
    A scheduled build that failed or was skipped because the search tier was saturated; data indexed after the window containing its timestamp had already been summarised; a definition change that left old and new rows incomparable; or an outage at the source. The remedy is a backfill over the affected range, plus alerting on the build job itself so the next gap is not found by a reader.
  • How would you decide whether a particular acceleration is paying for itself?
    Compare recurring build cost against reads actually served. A summary rebuilt every fifteen minutes and read twice a week is a loss, and one rebuilt hourly behind a dashboard forty people open daily is a clear win. Reviewing that ratio on a schedule is what keeps a platform's precomputation budget from growing without anyone deciding to grow it.

saying these in an interview costs you the question

  • Accelerates a search whose real problem is an unbounded time range
  • Averages summarised averages and calls it the fleet average
  • Assumes late-arriving events appear in an already-built summary
  • Forgets that build jobs compete with people for search capacity
  • Expects to drill down to raw events from a summary row