In Graylog, how does a raw log line reaching an input become a structured message, and where should field extraction happen?
answer
- Something has to be listening
- A protocol bound to a port
- The codec decides what arrives structured
- Two places parsing can live
basics
~20 sA Graylog input binds a protocol to a port and its codec decodes each payload into a message with source, message and timestamp fields. Extractors on that input parse further; pipeline rules do so later, across inputs.
solid answer
~50 sA Graylog **input** is a listener launched on one node or globally on every node, bound to a protocol and port - Syslog UDP/TCP, GELF UDP/TCP/HTTP, Raw/Plaintext, or an input pulling from Kafka or AMQP. The input's codec turns each payload into a Graylog message: GELF arrives as JSON and already carries `short_message`, `host`, `timestamp`, `level` and any underscore-prefixed custom fields, while a raw line lands mostly as one `message` string. Further structure comes either from **extractors** attached to that input - Grok pattern, regular expression, JSON, split & index, substring, copy input, lookup table - or from **pipeline rules** running later, in the pipeline processor. Prefer extractors for a format only that input receives; prefer pipeline rules when the same parsing must cover several inputs, needs conditions, or has to be reviewable in source.
code
json · 10 lines{
"version": "1.1",
"host": "checkin-api-03",
"short_message": "membership scan rejected",
"timestamp": 1757068213.417,
"level": 4,
"_gym_id": "bristol-north",
"_member_id": 41827,
"_scan_latency_ms": 312
}go deeper
Be ready to say that a Graylog input is a listener bound to a protocol and a port, and that a message arrives carrying fields such as source, message and timestamp. Knowing that GELF is JSON while syslog is a line of text is enough here.
Explain how the codec, an extractor and a pipeline rule divide the parsing work, name the common extractor types, and describe what changes when an input is launched globally rather than on one node.
Show judgement about where parsing belongs once many senders and formats share a cluster: per-input extractors that drift as senders move, versus pipeline rules that apply uniformly, plus what UDP loss and GELF chunking do under load.
Own the policy rather than the mechanism: who may create inputs, whether parsing is declared in reviewable source or clicked into a UI, and how a sender-format contract stops a shared cluster becoming hundreds of bespoke extractors nobody dares change.
## What a Graylog input actually is A **Graylog input** is a listener that the cluster starts and supervises. It is defined by a type, a bind address, a port and type-specific options, and it is launched either on a single node or **globally**, meaning every node runs its own copy of it. A global input is what lets a load balancer spread senders across nodes; a node-local input is for a source that must be read by exactly one process. The common types fall into three shapes: - **Push over a network socket** - Syslog UDP, Syslog TCP, GELF UDP, GELF TCP, Raw/Plaintext UDP and TCP. The sender opens a connection or fires a datagram at the port. - **Push over HTTP** - GELF HTTP, where each request body is one message. - **Pull from a broker** - inputs that consume from Kafka or AMQP. Here Graylog is the consumer, so backpressure and replay come from the broker rather than from the wire. That choice is not cosmetic. UDP has no handshake, no retransmission and no backpressure: when the receive buffer fills, the kernel discards datagrams and neither side records it. GELF over UDP additionally splits a large payload into chunks, and losing one chunk discards the whole message. TCP and broker-backed inputs turn loss into a visible stalled queue instead of a silent gap. ## From bytes to fields: what the codec already does Every input pairs with a codec that turns a payload into a Graylog message - a map of fields around a small reserved core of `message`, `source` and `timestamp`. How much structure you get for free depends entirely on the format on the wire. | Format on the wire | What the codec produces | What is left to do | |---|---|---| | Raw plaintext line | One `message` string plus `source` | Everything | | Syslog line | Host, timestamp, severity and the text | Application fields inside the text | | GELF (JSON) | `short_message`, `host`, `timestamp`, `level` and every underscore-prefixed custom field | Usually nothing | GELF exists for exactly this reason: the sender does the structuring, custom keys are prefixed with an underscore so they cannot collide with the format's own keys, and a numeric `level` carries severity. A service you control should emit GELF. A network appliance you do not control will send syslog, and syslog needs parsing. ## Extractors: parsing attached to one input An **extractor** is configured on a specific input and runs on messages that arrive through it. Graylog ships several types - **Grok pattern**, **regular expression**, **regex replace**, **JSON**, **split & index**, **substring**, **copy input** and **lookup table** - each reading one source field and writing one or more target fields. Two properties matter far more than the type list: 1. An extractor's scope is the input, not the cluster. A sender that is moved from the Syslog TCP input to a GELF HTTP input silently stops being parsed, and the messages still arrive, so nothing looks broken. 2. Extractors are built in the UI against a live input. That makes them quick to iterate on and awkward to diff, review or promote between environments. ## Pipeline rules: parsing attached to streams The alternative is a **pipeline rule**, written in Graylog's rule language and reaching messages through a pipeline connected to a **stream**. Rules execute in the pipeline processor, which is ordered after the message filter chain, so by the time a rule runs the message has already been routed. A rule tests with functions such as `has_field()`, parses with `grok()`, `key_value()` or `parse_json()`, and writes results with `set_field()`. Because a rule hangs off a stream rather than an input, one definition covers every sender routed there, and the rule is text that can live in version control and be reviewed like any other change. ## Choosing where extraction lives | Question | Extractor on the input | Pipeline rule | |---|---|---| | Applies to | Messages from one input | Every message in the connected streams | | Conditional logic | Limited to the extractor's own condition | Full `when`/`then` expressions | | Review and promotion | Clicked in the UI | Rule source, portable and reviewable | | Runs | In the message filter chain | After routing, in the pipeline processor | The working rule: **extractors for a format only one input ever receives; pipeline rules for anything that must be uniform across senders, conditional, or auditable.** ## What this looks like on a real cutover A climbing-gym membership platform moving off a hosted logging vendor mid-quarter stood up a 7-node cluster and pointed its services at it in three waves. The first wave sent GELF over TCP from services the team owned and needed almost no parsing at all. The second wave was door-controller firmware that speaks only syslog over UDP, so it got a dedicated input with a Grok extractor for the appliance's fixed line format - acceptable precisely because nothing else will ever send that shape. The third wave, a payments partner posting JSON over HTTP, was deliberately given no extractors: its normalisation went into a pipeline rule, so the same parsing still applied when the partner's second region came online three weeks later on a different input.
- What does launching a Graylog input globally change about where it runs?A global input starts on every node in the cluster, so any node can receive traffic and a load balancer can spread senders across them; the configuration and its extractors exist identically everywhere. A node-local input runs on one node only, which is what you want for a source that must not be read by two consumers at once.
- Why can two identical log lines end up with different fields in Graylog?Extractors belong to a single input. If one sender points at the Syslog TCP input and another at the GELF HTTP input, only the first gets that input's extractors, so the same text yields different fields. Parsing that must be uniform across senders belongs in a pipeline rule connected to the stream both messages reach, not in per-input extractors.
- What breaks when a high-volume sender is pointed at a UDP input?UDP offers no delivery guarantee and no backpressure: once the receive buffer fills, the kernel discards datagrams and Graylog never learns they existed, so loss appears only as a gap. Large GELF payloads are also split into chunks, and a missing chunk discards the whole message. A TCP or broker-backed input converts that silent loss into visible queueing instead.
The input is the letterbox and its codec the person who opens the envelope; an extractor is a note taped to that one letterbox, while a pipeline rule is the mailroom's standing instruction for everything reaching a department.
saying these in an interview costs you the question
- Says the input itself parses arbitrary log formats automatically
- Believes extractors on one input apply to every input
- Assumes UDP syslog delivery is reliable because messages usually arrive
- Cannot say where extractors run relative to pipeline rules
- Treats GELF as just another name for plain syslog
- Puts all parsing in UI extractors that nothing reviews