skip to content

In a Logstash pipeline config, what do input, filter and output do, and what does a failed grok match cost?

level: middleimportance: should knowfreq 58%

answer

  1. Three blocks, one direction of travel
  2. A miss is tagged, not dropped
  3. Failure costs more CPU than success
  4. Conditionals live in the output section too
  5. Split on delimiters before reaching for regex

basics

~20 s

Inputs receive events, filters transform them, outputs write them, in that order. A grok pattern that fails to match does not drop the event: Logstash tags it _grokparsefailure and ships it unparsed, after burning more CPU than a successful match.

solid answer

~40 s

A Logstash configuration has three blocks. `input` plugins receive events (`beats`, `file`, `kafka`); `filter` plugins transform them (`grok` for regular-expression extraction, `dissect` for fixed delimiters, `date`, `mutate`, `json`); `output` plugins write them (`elasticsearch`, `file`, `s3`). Plugins run in written order, per event. `grok` matches named patterns against `message` and promotes the captures to fields, with a `:int` or `:float` suffix casting the value. When it misses, the event is **not** dropped — it gets a `_grokparsefailure` tag and continues to the destination unparsed, so a broken pattern shows up as healthy indexing of useless documents. It is also the more expensive path, since the engine backtracks through every alternative before failing. Conditionals in the `output` section route by field or tag, which is how quarantine and multi-destination fan-out are built.

code

text · 19 lines
text
input {
  beats { port => 5044 }
}

filter {
  grok {
    match => { "message" => "^%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:log_level} meter=%{DATA:meter_id} ms=%{NUMBER:duration_ms:int}" }
    tag_on_failure => ["_grokparsefailure"]
  }
  date { match => [ "ts", "ISO8601" ] }
}

output {
  if "_grokparsefailure" in [tags] {
    file { path => "/var/log/logstash/unparsed-%{+YYYY.MM.dd}.log" }
  } else {
    elasticsearch { hosts => ["https://es:9200"] index => "heating-billing-%{+YYYY.MM.dd}" }
  }
}

go deeper

for a junior

Recall the shape of the file: three blocks named input, filter and output, and events flowing through them in that order. Know that grok is the plugin that turns a raw log line into named fields.

for a middle

Explain the mechanics — plugins run per event in written order, grok expands to regular expressions, a miss adds a failure tag rather than dropping the event, and a type suffix is what stops every capture being a string.

for a senior

Show that you have operated one: alerting on the parse-failure rate, spotting backtracking as the cause of a saturated tier, and using conditional outputs to quarantine unparsed data instead of mixing it with good documents.

for a principal

Argue about who owns log formats. A pipeline full of hand-tuned patterns is a standing tax on a platform team; decide whether to keep absorbing format drift centrally or push structured output back onto the producing services.

## The three sections, and what each one owns A Logstash pipeline is one configuration file with three blocks, and an event moves through them strictly in that order. - **`input`** — plugins that pull or receive events. The `beats` input listens for shippers, `file` tails a path, `kafka` consumes a topic, `http` accepts posts. An input's job ends when it has produced an event with a `message` field and some metadata. - **`filter`** — the transformation stage, and where nearly all the CPU goes. `grok` applies named regular-expression patterns, `dissect` splits on literal delimiters, `date` parses a timestamp string into the event's real time, `mutate` renames, converts and removes fields, `json` decodes an embedded document, `kv` splits `key=value` runs. - **`output`** — where the event is written. `elasticsearch` indexes it, `file` writes it, `s3` archives it, `stdout` prints it for debugging. Several outputs may appear, and they may be wrapped in conditionals. Filters and outputs both run *per event*, so the order of plugins inside a block is the order of work. Two `grok` filters mean two regular-expression passes on every line. ## What grok actually does, and what a miss costs `grok` matches a named pattern against a field — almost always `message` — and promotes the named captures to top-level fields. A pattern such as `%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:log_level} meter=%{DATA:meter_id} ms=%{NUMBER:duration_ms:int}` turns one string into four typed fields. The `:int` suffix casts the capture, without which every extracted value is a string. Underneath, each `%{...}` expands to a regular expression, and that is where the cost lives: | | Cost when it matches | Cost when it does not | |---|---|---| | `dissect` on a fixed-delimiter line | very cheap: a scan for literal separators | fails fast, nothing to backtrack | | `grok` with anchored, specific patterns | moderate: one regex pass per event | fails after a bounded search | | `grok` leaning on `%{GREEDYDATA}` or `%{DATA}` | already expensive | catastrophic: the engine backtracks over the whole line | A failed match is **not** a dropped event. Logstash adds a tag — `_grokparsefailure` by default, configurable through `tag_on_failure` — and the event carries on through the rest of the pipeline and out to the destination, unparsed. That is the trap. The pipeline looks healthy, indexing succeeds, and what you have stored is a pile of documents whose only real field is the raw `message`. Nobody notices until a query over the extracted fields returns a fraction of what it should. The second cost is throughput. A pattern that fails has usually done *more* work than one that succeeds, because the regular-expression engine explored every alternative before giving up. On a district-heating billing platform pushing 18 GB a day, one release that changed a log prefix so that roughly 31% of lines stopped matching was enough to push a two-node Logstash tier from about 40% CPU to saturation, at which point the pipeline stopped keeping up and backpressure reached the shippers. Practical defences: 1. **Anchor patterns.** Start with `^` and prefer specific patterns over `%{DATA}` and `%{GREEDYDATA}` anywhere but the final capture. 2. **Use `dissect` when the format is positional** and keep `grok` for genuinely irregular text. 3. **Alert on the failure tag.** Count documents carrying `_grokparsefailure` and treat a rise as a production incident, because it is one. 4. **Give grok a timeout** so a pathological line cannot pin a worker indefinitely. ## Conditional routing in the output section Conditionals are ordinary `if` / `else if` / `else` blocks that test event fields with the `[field][subfield]` reference syntax, and they can appear in `filter` as well as in `output`. Their most valuable use is separating traffic: - unparsed events (tagged `_grokparsefailure`) to a quarantine destination where someone will actually look at them, rather than into the same index as good data; - one service's events to a dedicated index while everything else goes to a shared one; - a copy of an audited subset to long-term archive storage alongside the searchable copy — the one thing no single-output tier can do. The mechanism matters for scale as much as for tidiness. If a regulator asks for six months of billing evidence, keeping that subset in its own destination from the moment it is parsed is far cheaper than trying to reconstruct it later from a mixed index. Conditional output is how you make that split at write time, and it is the strongest argument for running a Logstash tier at all. One caveat: conditionals evaluate per event on every worker, so a long `else if` chain on a hot pipeline is real CPU. Test fields you already have rather than computing new ones just to branch on them.

  • Your Logstash tier is at full CPU and lagging. How do you find out whether grok is the cause?
    Compare the share of documents carrying the failure tag before and after the lag started — a jump means patterns stopped matching and the engine is backtracking. Then look at the pipeline's per-filter timings, which attribute duration to individual plugins, and thread stacks on a saturated worker. Replacing a `%{GREEDYDATA}`-heavy pattern with `dissect` or an anchored pattern is usually the fix.
  • Why does adding a second grok filter to the same pipeline hurt more than it looks?
    Because filters run per event, in order. A second `grok` is a second regular-expression pass over every line, including the lines the first one already parsed and the lines neither will ever match. On a high-volume pipeline that is a straight doubling of the most expensive stage, and it applies whether or not the second pattern is relevant to the event in front of it.

saying these in an interview costs you the question

  • Believing Logstash drops an event when its grok pattern fails
  • Assuming a failed match is cheaper than a successful one
  • Using GREEDYDATA in the middle of a pattern by default
  • Never alerting on the parse-failure tag
  • Thinking filters can only appear once each in a pipeline
  • Reaching for grok when dissect would split the line