In Graylog, how do streams and their connected pipelines decide a message's fate, and what happens if it matches no stream rule?
answer
- Messages are sorted as they arrive
- One message, possibly several categories
- Numbered stages, not pipeline order
- Nothing matched is not nothing kept
basics
~20 sEvery Graylog message is tested against all stream rules and can join several streams, each with its own index set. Connected pipelines then run stage by stage and may enrich, re-route or drop it. Unmatched messages stay in the default stream.
solid answer
~40 sA Graylog **stream** is a named routing category defined by **stream rules** - field comparisons such as exact match, contains, regular expression, numeric comparison or field presence - combined with match-all or match-any. Routing runs in the message filter chain, a message can land in many streams at once, and each stream is bound to an **index set** that decides which indices it is written into. **Pipelines** are connected to streams and hold numbered **stages**: every pipeline attached to the message's streams executes stage 0 before any of them executes stage 1. Rules use `when`/`then` blocks and can `set_field()`, `route_to_stream()`, `remove_from_stream()` or `drop_message()`, and a message routed to a new stream picks up that stream's pipelines only from the next stage. A message matching nothing is not lost: it remains in the default stream.
code
text · 14 linespipeline "checkin telemetry"
stage 0 match either
rule "drop turnstile heartbeats"
stage 1 match all
rule "route rejected scans"
end
rule "route rejected scans"
when
has_field("scan_result") && to_string($message.scan_result) == "rejected"
then
set_field("gym_alert", true);
route_to_stream(name: "checkin-failures");
endgo deeper
Be able to say that a Graylog stream is a saved routing category, that stream rules match on message fields, and that you normally search inside a stream rather than across everything the cluster holds.
Explain that streams are evaluated for every message so one message can join several, that each stream is bound to an index set, and that pipelines are connected to streams and execute in numbered stages.
Demonstrate the ordering rules under real load: stage numbers order work across all connected pipelines, route_to_stream() only reaches later stages, and routing happens before pipeline processing, so a rule-created field cannot drive stream matching.
Own the routing design itself: how many streams a shared cluster should carry, whether teams may create their own, what duplicate storage across index sets costs, and what evaluating every stream rule against every message means at full ingest rate.
## A Graylog stream is a routing decision, not a folder A **stream** is a named set of **stream rules** evaluated against every message during processing. Matching is not exclusive: a message is tested against all streams and joins each one whose rules it satisfies, so one message can be in several streams at once. "Which stream is this message in" is therefore the wrong question; the right one is "which streams claimed it". Routing happens inside the **message filter chain**, the same processing phase that runs extractors. That ordering has a consequence people trip over constantly: a field created later by a pipeline rule does not exist yet when stream rules are evaluated, so it cannot drive routing. If routing must depend on parsed content, either the parsing happens earlier, or the message is moved from inside a rule with `route_to_stream()`. ## What stream rules can compare Each rule names a field and a comparison. The available shapes cover: - **match exactly** and **contains** on a field's value, - **match regular expression**, - **greater than** and **smaller than** for numeric fields, - **field presence**, which asks only whether the field exists at all. A stream combines its rules with either *match all rules* or *match at least one rule*, and a rule can be inverted. Because every rule of every stream is evaluated for every message, rule cost is multiplied by the ingest rate: a regular expression is paid per message per stream, while an exact match on a short field is cheap. ## Every stream is bound to an index set A stream points at an **index set**, and the messages the stream claims are written into that index set's indices. Two things follow immediately. A message claimed by three streams backed by three different index sets is stored three times. And moving a class of messages between streams changes where they physically live, which means storage layout is a routing decision rather than a separate setting. ## Pipelines, stages, and the ordering rule A **pipeline** is connected to one or more streams and contains numbered **stages**. A stage declares a match condition and lists rules; a rule has a `when` block deciding whether it applies and a `then` block of actions. The part that catches people: 1. All pipelines connected to the message's streams execute together, ordered by **stage number** - stage 0 of every connected pipeline runs before stage 1 of any of them. 2. Within one stage, rule execution order is not something to rely on. Two rules in the same stage must not assume one has already run. 3. If a rule calls `route_to_stream()`, pipelines connected to the newly added stream are picked up from the **next** stage onward, never retroactively for stages that already ran. So "do this after that" is expressed by stage number, and only by stage number. ## Re-routing, removing and dropping Three actions look similar and are not variations of one thing: | Action | Effect on the message | Effect on storage | |---|---|---| | `route_to_stream()` | Adds it to another stream | Also written through that stream's index set | | `remove_from_stream()` | Takes it out of one stream | No longer written through that stream | | `drop_message()` | Ends processing for it entirely | Not indexed anywhere | Only `drop_message()` discards data, and it is the right tool for high-volume noise you have decided never to keep - the saving is real because the decision lands before indexing rather than after. ## The message that matches nothing Every message enters the **default stream**, so a message matching no other stream's rules simply stays there. It is indexed, searchable and retained under that stream's index set: not lost, and not free. Two practical consequences follow: - Unrouted volume is invisible until someone looks for it. A sender whose format changed can stop matching its team's stream and quietly pile up in the default stream instead. - Read access to the default stream is read access to everything, because everything passes through it. A stream can be configured to remove its matches from the default stream, which is what keeps the default stream meaningful as "things nobody claimed" rather than a second copy of the entire cluster. ## A worked routing design On a 7-node cluster taking over from a hosted vendor mid-quarter, a climbing-gym platform routed its services into 11 streams: one per owning team, plus a `checkin-failures` stream fed only by `route_to_stream()` from a pipeline rule, because the field it keys on is produced by a Grok rule and so cannot be matched by a stream rule. Door-controller heartbeats - 3,140 a second of them - are discarded with `drop_message()` in stage 0, before they cost an index write. Every team stream removes its matches from the default stream, which turns the default stream's own throughput, normally under 40 messages a second, into the signal for "someone is sending us something nobody owns".
- Two Graylog pipelines are connected to the same stream. In what order do their rules run?Stage number wins over pipeline identity: stage 0 of both pipelines runs before stage 1 of either. Within a single stage, rule order is not guaranteed, so a rule must not assume another in the same stage has already set a field. If one rule must see another's output, put them in different stages.
- What does connecting a Graylog pipeline to the default stream do?It makes that pipeline run for effectively every message, because every message enters the default stream unless a rule or a stream setting removes it. That is right for normalisation meant to apply to everything, and expensive for narrow parsing: the cost is paid at full ingest volume rather than on the subset the work is actually for.
- How would you diagnose a message that never reached the stream you expected?Replay a real message through Graylog's processing simulation and see which stream rules matched. The usual cause is that the field being compared did not exist yet: stream routing happens in the message filter chain, ahead of the pipeline processor, so a field a pipeline rule creates cannot drive stream matching. Route with `route_to_stream()` from the rule instead.
Stages are rounds in a relay: every connected pipeline finishes round zero before anyone starts round one, so work that must happen first has to sit in an earlier round rather than earlier in the file.
saying these in an interview costs you the question
- Says a message can belong to only one Graylog stream
- Thinks messages matching no stream rule are discarded
- Believes rules inside one stage run in a guaranteed order
- Expects a stream rule to match a field a pipeline rule created
- Confuses drop_message with removing a message from one stream
- Ignores that each stream writes into its own index set