skip to content

Elastic Stack

The classic self-hosted logging pipeline: Beats or Logstash ship and parse, Elasticsearch indexes, Kibana searches. Interviews centre on index and shard design, mappings, and ILM retention — the things that decide whether the cluster survives your log volume.

on this pageshow

questions

4

In an Elastic Stack, how do you choose between parsing logs in Filebeat, in Logstash, or in an Elasticsearch ingest pipeline?

level: middleimportance: must knowfreq 71%

answer

  1. Three tiers can do the parsing
  2. Whose CPU pays for the regex?
  3. Only one tier fans out to several outputs
  4. Filebeat modules install pipelines into the cluster
  5. Shipper, Logstash tier, ingest pipeline

basics

~20 s

Parse where you can afford the CPU. Filebeat ships cheaply on the host; a Logstash tier adds heavy enrichment and multi-destination routing but costs a tier to run; an Elasticsearch ingest pipeline parses inside the cluster.

solid answer

~50 s

Three tiers can do the work. **Filebeat** tails files, tracks its offset in a registry and runs a light processor chain on the application's own host — cheap, but it burns the application's CPU and can send to only one output. **Logstash** is a separate tier with `input`/`filter`/`output` stages, the full plugin catalogue, an on-disk queue and conditional fan-out to several destinations; it is the only tier with real durability, and the only one that can write the same event to two unrelated places. An **Elasticsearch ingest pipeline** runs processors inside the cluster just before indexing, which deletes a whole tier but spends cluster CPU that search also wants. Note that a Filebeat module usually installs an ingest pipeline rather than parsing locally. Prefer emitting JSON and parsing nothing; failing that, split on delimiters before reaching for a regular expression.

code

json · 14 lines
json
{
  "processors": [
    { "grok": {
        "field": "message",
        "patterns": ["%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:log.level} meter=%{DATA:meter.id} ms=%{NUMBER:duration_ms}"]
    } },
    { "convert": { "field": "duration_ms", "type": "long" } },
    { "date":    { "field": "ts", "formats": ["ISO8601"], "target_field": "@timestamp" } },
    { "remove":  { "field": "ts" } }
  ],
  "on_failure": [
    { "set": { "field": "error.message", "value": "{{ _ingest.on_failure_message }}" } }
  ]
}

go deeper

for a junior

Be ready to name the pieces and their order: Filebeat or Logstash ships, Elasticsearch stores and searches, Kibana displays. Knowing that parsing can happen in more than one of those places is enough at this level.

for a middle

Explain what each tier can actually do — a light processor chain in the shipper, the full filter catalogue in Logstash, processors inside the cluster in an ingest pipeline — and say which of them can route one stream to several outputs.

for a senior

Show the production judgement: whose CPU each option spends, what happens to a document when a processor throws in each tier, and why you would delete the Logstash tier rather than keep it out of habit.

for a principal

Own the tradeoff across the estate. Decide whether parsing is a platform service or a team responsibility, what you standardise so producers can be moved between tiers later, and what you are willing to spend on a tier that exists only for durability and routing.

## Three places a log line can become fields An Elastic logging deployment gives you three points where an unstructured line turns into structured fields, and they sit at very different distances from the data. 1. **On the host, inside the shipper.** Filebeat tails files, remembers its read offset in a registry so a restart does not re-send everything, and ships. It carries a short processor chain — `dissect`, `decode_json_fields`, `drop_fields`, `add_host_metadata` — that runs in the shipper's own process, on the machine the application is running on. 2. **In a separate Logstash tier.** A dedicated JVM process built as `input` / `filter` / `output` stages, with the full filter catalogue (`grok`, `dissect`, `mutate`, `kv`, `translate`, `geoip`, `ruby`), a queue of its own, and the ability to fan one incoming stream out to several destinations under conditionals. 3. **In an Elasticsearch ingest pipeline.** A named, ordered list of processors (`grok`, `dissect`, `date`, `convert`, `set`, `rename`, `script`) run by nodes holding the ingest role immediately before the document is indexed. An index can name a default pipeline, so producers need not know it exists. A detail that trips people up: a Filebeat **module** does not usually parse on the host. It ships a thin Filebeat configuration *and* installs an ingest pipeline into the cluster, so "the module parses it" really means "Elasticsearch parses it." ## What each tier buys and what it costs | | Filebeat processors | Logstash tier | Ingest pipeline | |---|---|---|---| | Where the CPU burns | on the application host | on dedicated Logstash nodes | on Elasticsearch ingest nodes | | Parsing power | light: fixed-delimiter split, JSON decode, field drops | full: regex grok, lookups, external enrichment | most processors, plus enrichment the cluster can resolve itself | | Routing | a single output only | conditional fan-out to many outputs | one destination per document | | Buffering on failure | the log file on disk is the buffer | on-disk queue and a dead-letter queue | none — a failure is an indexing failure | | What it costs to run | nothing beyond the shipper you already have | a tier to size, patch, scale and monitor | contends with search on the same cluster | ## How the choice is actually made The rule of thumb is *parse as late as you can afford to, and as cheaply as the format allows*: 1. If the application can emit JSON, do that and parse almost nothing. This beats every option below. 2. If the line has fixed delimiters, split it rather than running a regular expression — that is `dissect` in either Logstash or an ingest pipeline, and it is dramatically cheaper. 3. If parsing genuinely needs a regular expression, decide whose CPU pays. Application hosts are usually the worst place: the budget belongs to the application, and a bad pattern degrades the service you are trying to observe. 4. If one stream has to reach several destinations — a hot index, a long-term archive, a security sink — that is Logstash's conditional output, and it is the strongest single reason to run the tier. 5. If there is exactly one destination and the cluster has ingest headroom, an ingest pipeline removes a whole tier from the diagram, along with its queue, its patching and its on-call. Take a district-heating billing platform emitting about 18 GB of meter-reading logs a day across 47 collector hosts. Those hosts are sized for the billing workload, so heavy grok on them is out. Two candidate designs survive: Filebeat straight into Elasticsearch with an ingest pipeline doing the parse, or Filebeat to a two-node Logstash tier that parses, enriches meter IDs against a tariff lookup, and writes both to the search cluster and to object storage. The lookup and the second destination decide it — an ingest pipeline can enrich from the cluster but cannot write the same event to two unrelated places, so the Logstash tier earns its keep. If the requirement had been parse-and-index only, the tier would be pure cost. ## Failure modes, which is where the answer gets senior - **Parsing on the host** couples log-format changes to application deployments and puts a runaway regular expression next to production traffic. - **A Logstash tier** is the only one of the three with real durability, but it is also a stateful hop you now have to capacity-plan, and its queue is a disk you can fill. - **An ingest pipeline** is the cheapest to operate and the least forgiving: there is nowhere for a document to wait. A processor that throws sends an indexing failure back to the client, and whether that becomes a retry, a dead letter, or a silently lost line depends entirely on what is in front of it. The senior version of the answer is that the three are not exclusive. Real deployments do structural work at the shipper (JSON decode, host metadata), expensive or multi-destination work in Logstash, and last-mile normalisation in an ingest pipeline attached to the index — and the design question is which of those three you can delete.

  • You move parsing off the application hosts into a Logstash tier. What new failure mode have you bought?
    A stateful hop. Logstash now has a queue that can fill, workers that can saturate, and a version to patch, and every log line in the estate flows through it — so it becomes a single point of failure for observability precisely when you need observability. It buys durability and routing; it costs capacity planning and on-call.
  • Why is emitting JSON from the application usually better than any of the three parsing options?
    Because the producer already knows the field boundaries and the types, so nothing downstream has to guess them with a regular expression that silently breaks the day someone adds a field. It removes an entire class of parse failures, cuts CPU everywhere, and makes the log format a contract the application owns rather than something a pattern reverse-engineers.
  • Can a Filebeat installation send the same log lines to two different destinations?
    Not from one Filebeat process — it accepts a single output. Reaching two destinations means running two shipper processes over the same files, or shipping to a tier that can fan out, which in this stack means Logstash and its conditional output section.

It is the same call as validating a form in the browser, in an API gateway, or in the database: the further from the source you push it, the harder it is to get around, and the more of the shared machine it costs.

saying these in an interview costs you the question

  • Assuming a Filebeat module parses on the host rather than in the cluster
  • Treating a Logstash tier as mandatory in every Elastic deployment
  • Ignoring that ingest parsing competes with search on the same cluster
  • Claiming Filebeat can be configured with several outputs at once
  • Reaching for grok when the line has fixed delimiters
  • Putting expensive regular expressions on the application's own hosts
open as a page

In a Logstash pipeline config, what do input, filter and output do, and what does a failed grok match cost?

level: middleimportance: should knowfreq 58%

basics

~20 s

Inputs receive events, filters transform them, outputs write them, in that order. A grok pattern that fails to match does not drop the event: Logstash tags it _grokparsefailure and ships it unparsed, after burning more CPU than a successful match.

open as a page

Several teams ship logs into one Elastic Stack cluster. How do you make a common field schema stick, and what breaks when two services send one field name with different types?

level: principalimportance: should knowfreq 41%

basics

~20 s

Agree a common vocabulary such as the Elastic Common Schema and enforce it in the shared shipping tier, not in documentation. When two services disagree on a field's type, one index rejects documents and searches across several go quietly wrong.

open as a page

What do Logstash's persistent queue and dead-letter queue each protect, and how does backpressure reach Filebeat?

level: seniorimportance: nice to knowfreq 27%

basics

~20 s

Logstash's persistent queue writes in-flight events to disk before filtering, so a restart or a downstream outage does not lose them. The dead-letter queue holds events Elasticsearch will never accept. Backpressure fills the queue and stalls Filebeat.

open as a page