In Splunk, how do you choose between a universal forwarder, a heavy forwarder and a network input to get data in?
answer
- An agent on the host, usually
- Light shipper or full parser?
- Only a parser can drop by content
- No agent means no buffer
- The meter counts what gets indexed
basics
~20 sA universal forwarder is a light agent shipping a host's raw data. A heavy forwarder parses events first, so it can mask, drop and route them before indexing. A network input accepts pushed data with no agent and no source-side buffer.
solid answer
~50 sPrefer an agent on the host. A **universal forwarder** tails files and local inputs and ships their contents on; it queues while the far end is down, but it does not parse events, so it cannot judge their content. A **heavy forwarder** is a full instance in forwarding mode: it breaks the stream into events and parses them, which is what lets it mask values, drop whole classes of events and route data to different destinations before anything is indexed — at real CPU cost. A **network input** is a port the platform listens on; no agent is needed, which makes it the answer for appliances that can only push and the worst choice otherwise, since nothing buffers at the source and a receiver restart loses what was in flight. Because the classic licence counts indexed volume, only a tier that parses can drop data before you pay for it.
go deeper
Recall the three routes in: a light agent on the host, a heavier parsing instance in the middle, or a port the platform listens on. Know that the light agent is the normal choice and why an agent is preferred at all.
Explain what the parsing step buys you — timestamp extraction, masking, dropping, routing — and why an agent that never parses events cannot filter them by content.
Show that you can place a parsing tier for a real estate: which sources need one, what a listening port costs in reliability and host attribution, and how source-side buffering behaves while the indexing tier is unavailable.
Own the ingest architecture and the bill attached to it: what may be onboarded, who runs the parsing tier, and whether volume is reduced at the application, by choosing what an agent reads, or by paying for a tier that parses data in order to throw it away.
## Three ways data reaches the indexing tier **A universal forwarder** is a small agent installed on the machine that produces the data. It tails files, reads local inputs and ships their contents onward. It is deliberately minimal: it carries no user interface, uses very little CPU and memory, can spread output across receivers, remembers how far through each file it has read, and queues while the far end is unavailable. What it does not do is understand your events. It ships a stream rather than parsed records, so it can choose which files to read but cannot decide anything from what an individual event says. **A heavy forwarder** is a full Splunk instance configured to forward rather than index. Because it runs the parsing work, it turns the stream into events, extracts timestamps and applies index-time processing. That is what unlocks everything the light agent cannot do: masking a value out of the raw text, dropping a whole class of events, routing different data to different destinations including ones outside Splunk, and hosting collection that needs a full instance, such as polling a remote service. It costs what a full instance costs in CPU, memory, disk and patching. **A network input** is a port the platform listens on for data pushed by something that cannot run an agent: a switch, a firewall, an appliance, a managed service that only emits over the network. It needs nothing at the sender, which is both its appeal and its problem. ## What each one buys and what it costs | | Universal forwarder | Heavy forwarder | Network input | |---|---|---|---| | Runs on | every host that produces data | a small dedicated tier | nothing; the sender pushes | | Footprint at the source | very small | that of a full instance | none | | Parses events | no | yes | only once received | | Can drop or mask by content | no | yes | not before arrival | | Buffering when the far end is down | yes, at the source | yes | none; in-flight data is lost | | Knows which host produced the data | yes, it is that host | yes | inferred from the sender or a relay | | Typical use | the default for servers and containers | a parsing and routing tier | appliances that can only push | The reliability row is the one candidates most often miss. An agent on the host owns a position in a file, so a receiver restart or a network partition costs time and nothing else — it resumes where it stopped. A pushed stream has no such memory: whatever was in flight when the receiver went away is gone, and over a connectionless transport nothing will tell you. The standard mitigation is not to abandon the network input but to put a relay in front of it that writes what it receives to disk, and then read those files with an agent. That restores both the buffer and a usable idea of which device each event came from. ## Where you parse decides what you pay for The classic Splunk licence is measured by how much data is indexed per day, which makes the placement of the parsing tier a commercial decision as much as a technical one. Only something that has parsed an event can decide to discard it, so content-based filtering happens at the first tier that parses: a heavy forwarder if the data passes through one, otherwise the indexing tier itself. Filtering later — in a search, in a dashboard, through an index's retention — saves nothing, because the volume has already been counted. That gives a short ladder, cheapest first: 1. **Do not produce it.** Turning off debug logging in the service that is drowning you is free and reduces load everywhere at once. 2. **Do not collect it.** Choosing which files an agent reads costs nothing at run time. 3. **Parse it and drop it.** A parsing tier can throw data away by content, but you pay for the tier doing the throwing. 4. **Store it and search around it.** No saving at all, and the option most estates drift into by default. On a vinyl-record marketplace running 41 services, one team was producing about seventy per cent of the daily indexed volume — 2.6 TB of a 3.7 TB day — almost all of it a debug-level access log left switched on after a launch. A parsing tier could have dropped it, at the price of hardware able to parse 2.6 TB a day; a single change in that team's service removed it at the source. Reaching for step three before step one is the expensive mistake this question is looking for. ## Choosing, in practice - Default to an agent on the host. It is the cheapest reliable option and it attributes data correctly. - Introduce a parsing tier when you need to mask, drop, route or collect from a remote service, and size it for the volume flowing through it rather than for the number of sources behind it. - Accept a network input only where nothing can be installed, and put a relay with disk in front of it. - Do not put a full parsing instance on every host: it is expensive per node and multiplies what must be patched.
- Why is a listening port a poor default for collection even though it needs no agent?Because there is no memory at the source. An agent tracks its position in a file and resumes after a restart or a partition; a pushed stream loses whatever was in flight, and over a connectionless transport nothing reports the loss. Host attribution also degrades, since events appear to come from whatever relayed them. Putting a relay that writes to disk in front, and reading those files with an agent, restores both.
- A team wants to halve its Splunk bill by discarding noisy events. Where does that discarding have to happen?At the first tier that parses the events — a heavy forwarder if the data passes through one, otherwise the indexing tier. A universal forwarder can choose which files to read but cannot judge an event's content. Cheaper still is not producing the data: turning the noisy logging off at the source costs nothing and reduces load everywhere.
- What do you give up by putting a heavy forwarder in front of everything?Cost and operational surface. Each one is a full instance to size, patch and monitor, it consumes CPU proportional to the volume it parses, and it becomes a bottleneck and a failure domain in the path. Running one per host is the usual overreaction; a small shared parsing tier sized for the flow through it is almost always the better shape.
The light agent is a courier who never opens the box; the parsing tier opens it, inspects the contents, and can refuse, redact or redirect it before it ever reaches the warehouse.
saying these in an interview costs you the question
- Thinks a universal forwarder can drop events by matching their content
- Treats a listening port as equivalent to an agent on the host
- Believes filtering after indexing reduces licensed volume
- Puts a full parsing instance on every host by default
- Ignores that a receiver restart loses whatever was pushed in flight