skip to content

Splunk

The enterprise machine-data platform: index anything, then search it with SPL, a pipeline language where you chain commands rather than write one query. Interviews focus on SPL fluency and on indexing volume, since licensing is priced on ingested gigabytes.

on this pageshow

questions

4

In Splunk, what is decided about an event at index time versus search time, and which of those decisions are permanent?

level: middleimportance: must knowfreq 74%

answer

  1. Two moments: writing and reading
  2. Schema is applied on read
  3. Which index, which timestamp, which stamps
  4. Fixing the written half means re-ingesting

basics

~20 s

At index time Splunk fixes an event's index, timestamp, host, source and sourcetype. Fields, aliases and lookups are worked out at search time and can be changed at any moment. The index-time decisions stick: correcting one means re-ingesting the data.

solid answer

~40 s

Splunk splits the work in two. At index time the stream is broken into events, a timestamp is parsed from each one, the event is stamped with the host it came from, the source it was read from and the sourcetype saying what kind of data it is, and it is routed to a particular index — which in turn decides its retention and who may search it. At search time the schema is applied: fields are pulled out of the raw text, aliases and calculated fields evaluated, lookups joined on, tags and event types applied. Search-time definitions are retroactive; add one this afternoon and it works against years of stored data. The index-time decisions are baked into what was written, so fixing a wrong timestamp, sourcetype or index means ingesting the data again.

go deeper

for a junior

Recall that Splunk keeps the raw event and works fields out when you search rather than when data arrives. Be able to name what is stamped at ingest: the index, the timestamp, the host, the source and the sourcetype.

for a middle

Explain which half of the work is reversible. An interviewer expects you to say that a new field extraction applies to years of stored data immediately, while a wrong sourcetype or timestamp stays with the events.

for a senior

Show that you onboard defensively: timestamp and timezone checked on a sample, classification agreed, index chosen for retention and access rather than convenience. Be ready to describe recovering from a bad onboarding without re-ingesting everything.

for a principal

Own the policy behind it: who may create an index, what the naming and classification conventions are, and where you accept ingest-time cost to make searches cheap. Permanence at write time makes this governance, not a per-team preference.

## Two moments, two kinds of decision Splunk stores the raw event and works most of the meaning out later. That single design choice is what the index-time versus search-time split describes, and it is why the platform can answer questions nobody anticipated when the data was onboarded. **At index time** — as data arrives at whichever tier parses it — a fixed set of decisions is made and written down with the event: - the incoming stream is broken into discrete events; - a **timestamp** is parsed out of each event and becomes its position on the timeline; - the event is stamped with the **host** it came from, the **source** it was read from, and the **sourcetype** that says what kind of data it is; - it is routed to a particular **index**; - any index-time processing configured for it is applied — masking a value out of the raw text, adding a field into the stored record, or sending a copy to another destination. **At search time** — every time somebody runs a search — the schema is applied on top of what was stored: fields are pulled out of the raw text, aliases rename them, calculated fields derive new ones, lookups join external tables onto rows, and tags and event types classify them. ## What is reversible and what is not | Decision | Made at | Changing it afterwards | |---|---|---| | Which index holds the event | index time | requires re-ingesting the data | | The event's timestamp | index time | requires re-ingesting the data | | The host, source and sourcetype stamps | index time | the stored values stay; you work around them | | The raw text as stored, including masking | index time | irreversible in both directions | | Field extraction | search time | edit it and it applies to everything already stored | | Aliases, calculated fields, tags, event types | search time | retroactive and reversible at any time | | Lookup enrichment | search time | change the table and every later search sees the new values | The asymmetry is the whole point. A field extraction invented this afternoon works against events indexed years ago, because the raw text was kept and the extraction runs while reading. A wrong timestamp cannot be fixed that way, because the event's place on the timeline is part of what was written. ## Why the permanent half is where the damage happens Three ingest mistakes account for most of the pain. 1. **A misparsed timestamp.** A locale-formatted date read the wrong way round, or a missing timezone, puts events at a point on the timeline nobody is looking at. The data is present and effectively invisible: a search over the last week returns nothing, and an alert over that window never fires. 2. **The wrong sourcetype.** Search-time extraction is keyed to how an event was classified, so a misclassified stream yields no fields. You can point extractions at the value that was actually stored and recover most of the value, but every dashboard, alert and lookup written against the intended classification has to be told about it. 3. **The wrong index.** Retention and search access in Splunk are administered per index. An event in the wrong one may be kept for the wrong length of time and be visible to the wrong set of people, which is a compliance problem rather than an inconvenience. On a vinyl-record marketplace running 41 services, a pressing-plant service was onboarded with a day-first date format read as month-first. About 3.4 million events landed eleven days away from where they belonged and nobody noticed until a weekly report came back empty; the fix was re-ingesting from the retained source files. In the same week the same team added two new field extractions covering a year of stored data, and those took effect immediately at no cost. That contrast — one correction requiring the data to be written again, the other applying retroactively for free — is the lesson the question is testing. ## When you would deliberately move work to index time Doing more at index time is not free. It costs CPU in the parsing tier, it can grow what is stored, and it is permanent. Two cases justify it: - **Removing data that must never be stored.** Masking has to happen before the event is written, because search time cannot un-store something. - **A field used constantly as a filter over very high volumes.** Making it part of what is written can be worth the ingest cost when the alternative is pulling it out of raw text on every search across an enormous dataset. Everything else should default to search time, where a mistake costs a correction rather than a re-ingest. The rule of thumb worth saying out loud: **decide as little as possible at write time, and decide that little carefully**, because it is the half of the schema you cannot take back. A practical corollary is that onboarding a new source deserves a sample check — timestamps, timezone, event boundaries, classification and destination index verified on real data — before the tap is opened, since every one of those five is in the permanent column.

  • You discover a whole source was stamped with the wrong sourcetype last month. What are your options?
    Fix the classification at the parsing tier so new data is right, then point the field extractions and any lookups at the value that was actually stored so the existing month becomes searchable. Re-ingesting from the original files is the only way to change the stored stamp, and it is worth it only when the volume is small or the data is central. If it also went to the wrong index, retention and access have to be checked as well.
  • When is it worth doing extraction work at index time instead of search time?
    Rarely. Two cases justify it: data that must never be stored at all, which has to be masked before writing, and a field used constantly as a filter over enormous volumes, where making it part of what is written repays the ingest cost. Everything else defaults to search time, because a mistake there costs an edit rather than a re-ingest.
  • Why does the choice of index matter beyond deciding where the event is stored?
    Retention and search access are administered per index, so the index an event lands in decides how long it survives and which roles can see it. Sending data to a convenient index rather than the right one can quietly shorten its life or expose it to people who should not have it.

Index time is the label on the box and the shelf it goes on; search time is how you read the contents once you take it down. Relabelling means unpacking and repacking everything.

saying these in an interview costs you the question

  • Describes Splunk as parsing fields once at ingest, like a relational schema
  • Believes a wrong sourcetype can be corrected in place on stored events
  • Thinks changing a field extraction requires re-ingesting old data
  • Ignores that the index chosen governs retention and who can search it
  • Treats a misparsed timestamp as cosmetic rather than as data you cannot find
open as a page

In Splunk, how is an SPL search composed as a pipeline, and what separates a stage that streams from one that must see every result?

level: middleimportance: must knowfreq 68%

basics

~20 s

An SPL search starts with a retrieving stage that pulls events from Splunk indexes over a time range, then pipes them onward. Streaming stages act on one event at a time; transforming stages must gather the whole result set first.

open as a page

In Splunk, how do you choose between a universal forwarder, a heavy forwarder and a network input to get data in?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A universal forwarder is a light agent shipping a host's raw data. A heavy forwarder parses events first, so it can mask, drop and route them before indexing. A network input accepts pushed data with no agent and no source-side buffer.

open as a page

In a large Splunk estate, when is precomputing an expensive search into a summary index or an accelerated report the wrong call?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

A summary index stores a scheduled aggregation's results as small events you search instead of raw data; report acceleration has Splunk maintain equivalent precomputed data for a qualifying saved search. Both buy speed with freshness, storage and recurring load.

open as a page