skip to content

In Splunk, how is an SPL search composed as a pipeline, and what separates a stage that streams from one that must see every result?

level: middleimportance: must knowfreq 68%

answer

  1. Left to right, stage by stage
  2. Something retrieves, the rest reshapes
  3. One event at a time, or all of them?
  4. A barrier forces central gathering
  5. Filter before the aggregation, not after

basics

~20 s

An SPL search starts with a retrieving stage that pulls events from Splunk indexes over a time range, then pipes them onward. Streaming stages act on one event at a time; transforming stages must gather the whole result set first.

solid answer

~50 s

Think of SPL as a shell pipeline over events. The first stage retrieves: it names the indexes and the time range and does the heavy filtering, and it is the only stage that decides how much data leaves disk. Every later stage receives the previous stage's output. **Streaming** stages examine one event at a time and can be pushed down to the indexers to run in parallel — extracting a field, evaluating an expression, dropping events by a condition. **Transforming** stages must see the whole result set before they emit anything — counting per group, ordering rows, keeping a top-N — so they run centrally and act as a barrier: everything after them runs on one node against a small table rather than a stream of events. Put filters before the first barrier, never after it.

code

pseudocode · 5 lines
pseudocode
retrieve  : events from the order-service index, last 24 hours
  | per event : pull the response code out of the raw text   -> streaming
  | per event : keep only codes of 500 and above             -> streaming
  | whole set : count the survivors per service              -> BARRIER
  | whole set : order the resulting rows, highest first      -> after the barrier

go deeper

for a junior

Be ready to read an SPL search out loud from left to right and say what each stage does to the rows it was handed. Knowing that the pipe passes one stage's output into the next is the recall being tested here.

for a middle

Explain the difference between a stage that decides per event and one that needs the entire result set, and give an example of each. An interviewer expects you to point at the first barrier in a search you are shown.

for a senior

Show that you position filters relative to that barrier deliberately, and can explain why two searches differing only in stage order behave completely differently. Talk about work pushed down to the indexers versus work done centrally.

for a principal

Own the guidance the estate follows: what a reviewable search looks like, where teams are told to filter, and how you stop a shared search tier being consumed by pipelines that gather everything centrally before narrowing it.

## A search is a chain of stages In SPL a search is written as a chain of stages separated by pipes, and it means exactly what it looks like: the first stage produces rows, and each later stage receives whatever the stage before it emitted. There is no planner quietly rewriting your intent into a different order, so reading the search from left to right tells you the order the work actually happens in. That is the single most useful thing to know about the language, because it makes the cost of a search something you can see rather than something you have to measure. Every search opens with a **retrieving stage**. It is the part that names which Splunk indexes are read, over which time range, and any conditions that can be evaluated against the raw text of an event or against the metadata stamped on it when it was written — which host it came from, which input it was read from, and which sourcetype it was classified as. This stage alone decides how many events leave disk. Nothing downstream can go back for more; every later stage can only discard, reshape or summarise what it was handed. A smaller family of searches begins instead with a **generating stage** that produces rows from something other than raw events — a lookup table, a catalogue of what is stored, a precomputed summary — but the model is identical: something produces rows, and the pipeline transforms them. ## Streaming stages and stages that need the whole set The distinction interviewers are really testing is whether a stage can decide what to emit **from one event in isolation**, or whether it must **see the entire result set** before it can emit anything at all. | | Streaming stage | Transforming stage | |---|---|---| | Input it needs | one event at a time | every row produced upstream | | Typical work | extract a field from raw text, compute a value, drop events failing a condition, rename or remove fields | count or average by a grouping field, order rows, keep the top or rarest values, pivot into a table | | What it emits | events, still individually identifiable | result rows; the original events are gone | | Where it can run | pushed down to the indexers, in parallel where the data lives | centrally, once results have been gathered | | Effect on the pipeline | data keeps flowing, partial results can appear | a barrier: nothing after it starts until everything before it finishes | Two readings follow from that table. First, a search is not one homogeneous unit of work: it has a **parallel section** executed across the indexing tier and a **serial section** executed on a single node, and the boundary between them is the first transforming stage. Second, the events themselves stop existing at that boundary. Anything a later stage needs must have been carried through the aggregation, either as a grouping field or as a computed statistic. ## Why stage order is the main cost lever Because the pipeline runs as written, where you place a filter decides how much work is done. 1. **A condition before the barrier is cheap.** It is evaluated on every indexer in parallel, against data that is already local, and it shrinks what has to cross the network to the coordinating node. 2. **The same condition after the barrier is expensive.** You have already paid to read, transport and aggregate rows that you then throw away. 3. **Cosmetic work belongs after the barrier.** Renaming, formatting and reordering columns applied to a few hundred result rows costs nothing; applied to millions of events it is millions of operations. 4. **Perceived speed changes too.** A pipeline of only streaming stages can show results as they arrive. Insert one aggregation and the search shows nothing until the last event has been counted, which is why an apparently trivial edit can make a search feel like it has hung. A worked case: on a vinyl-record marketplace running 41 services, a report retrieved nine days of one index, aggregated pressing-plant errors per service, and then filtered the resulting table down to the four services the report is actually about. Moving that filter into the retrieving stage took the events read from 214 million to 6.9 million and the runtime from about eleven minutes to under forty seconds. The shape of the pipeline never changed; only the position of the filter relative to the barrier did. ## Reading an unfamiliar search - Find the retrieving stage and ask what scope it sets: which indexes, what time range, what metadata conditions. - Find the first stage that must see everything. Everything before it is per-event work; everything after it operates on a small table. - Ask whether a filter sits after that boundary and could be moved before it. - Ask what the aggregation kept, and whether a later stage depends on a field that no longer exists at that point. The two-way split is a model, not a complete taxonomy. Some stages can only run once results have been gathered centrally even though they emit one row per input; a few restructure the dataset wholesale. Reason with the model, then check how a specific stage behaves before you build a large search around it.

  • Why can a search that ends in an aggregation show nothing for a long time and then everything at once?
    Because the aggregation cannot emit a row until it has counted the last event. A pipeline of per-event stages finishes each event as it passes, so partial results appear immediately. Insert a stage that needs the whole set and the search must complete its upstream work before anything is displayed.
  • If a later stage needs a field, does it matter whether it was extracted before or after the aggregation?
    Yes. An aggregation emits only what it was told to produce — the grouping fields and the computed statistics. The original events, and every field on them that was not carried through, are gone from that point on. Either extract and carry the field before the barrier, or derive it from what the aggregation kept.
  • What does the retrieving stage give you that no later stage can?
    Scope. It is the only stage that decides which indexes and which time range are read, and which conditions can be applied against ingest-time metadata before events are loaded. Everything downstream can only discard or reshape what that stage produced, so a search that retrieves too much can never be made cheap further along the pipeline.

It behaves like a shell pipeline: some stages are a per-line filter that can start emitting immediately, and some behave like a sort, which cannot print a thing until it has read the last line.

saying these in an interview costs you the question

  • Thinks SPL is declarative and a planner reorders the stages for you
  • Places filters after the aggregation, assuming the optimiser sorts it out
  • Believes every stage of a search runs on the indexers
  • Cannot name a stage that needs the entire result set
  • Assumes the original events are still available after an aggregation