What do Logstash's persistent queue and dead-letter queue each protect, and how does backpressure reach Filebeat?
answer
- Two queues, opposite problems
- One is before filtering, one is after output
- Retryable is retried; permanent is dead-lettered
- A queue nobody reads is still data loss
- The last buffer is the file on disk
basics
~20 sLogstash's persistent queue writes in-flight events to disk before filtering, so a restart or a downstream outage does not lose them. The dead-letter queue holds events Elasticsearch will never accept. Backpressure fills the queue and stalls Filebeat.
solid answer
~40 sThey solve opposite problems. The **persistent queue** (`queue.type: persisted`) writes each batch to disk between the input and the filter stage and acknowledges only after processing, so in-flight events survive a crash and a destination outage becomes a delay bounded by `queue.max_bytes` rather than immediate loss. The **dead-letter queue** (`dead_letter_queue.enable: true`) catches the other case: documents the Elasticsearch output was told are permanently unindexable — a type conflict, a malformed value — which no retry will ever fix. It is worthless unless a second pipeline reads it back with the `dead_letter_queue` input and someone repairs the data. Backpressure then chains upward: the output stops draining, filter workers block, the queue fills, the `beats` input stops acknowledging, and Filebeat stops reading — leaving the log file itself as the final buffer, until rotation deletes it.
code
yaml · 7 linesqueue.type: persisted
path.queue: /var/lib/logstash/queue
queue.max_bytes: 4gb
dead_letter_queue.enable: true
path.dead_letter_queue: /var/lib/logstash/dlq
dead_letter_queue.max_bytes: 1gbgo deeper
Know that Logstash buffers events between receiving them and writing them out, and that this buffer can be kept in memory or on disk. The distinction is what decides whether a restart loses data.
Explain the mechanics: where each queue sits in the pipeline, which configuration setting enables it, and the difference between a failure worth retrying and one that will never succeed.
Demonstrate that you have run this. Recite the backpressure chain to the log file, size a queue against a realistic outage, and treat an unread dead-letter queue as an incident rather than a feature you enabled.
Own the durability budget across the estate. Decide whether local disk queues are enough or a broker belongs in front of the tier, and be explicit about which classes of log you are willing to lose and at what point in the chain.
## Two queues, two entirely different jobs Logstash has two on-disk structures whose names sound similar and whose purposes do not overlap at all. Confusing them is the single most common mistake on this subject. | | Persistent queue | Dead-letter queue | |---|---|---| | Sits | between the input and the filter stage | after an output that rejected an event | | Holds | events accepted but not yet fully processed | events an output was told it can never deliver | | Protects against | process crash, restart, transient output outage | permanently malformed or unindexable documents | | Turned on with | `queue.type: persisted` in `logstash.yml` | `dead_letter_queue.enable: true` | | Drains by | the pipeline catching up | a human, via a pipeline reading the queue | ## The persistent queue: what it changes over an in-memory one By default a pipeline's queue lives in memory. It holds a bounded number of in-flight batches, and if the process is killed — an out-of-memory kill, a node reboot, a bad deploy — every event sitting in it is gone. There is no copy anywhere else, because the input already told the sender it had the data. Setting `queue.type: persisted` writes each incoming batch to an on-disk queue under `path.queue` *before* the filter stage runs, and only acknowledges an event once it has been processed and handed to the outputs. Three consequences follow: 1. **In-flight events survive a restart.** On startup the pipeline replays what was never acknowledged. Duplicates are possible on the replayed batch; loss largely is not. 2. **The queue absorbs an outage.** If the destination is unavailable, events accumulate on disk up to `queue.max_bytes` instead of immediately pushing back on the shipper. That converts a short outage into a delay rather than an incident. 3. **You pay for it.** Every event is now a disk write with periodic checkpointing (tunable through `queue.checkpoint.writes`), so throughput drops and the queue directory becomes a disk you must size and monitor. A full persistent queue does not silently discard — it blocks, which is the correct behaviour and is what makes backpressure work. The queue is per pipeline, not per node, and it lives on that node's local disk. It is durability against a process dying, not against the machine being destroyed. ## The dead-letter queue: what actually lands there The dead-letter queue exists for the opposite problem: an event that will *never* succeed no matter how many times it is retried. In practice it is fed by the Elasticsearch output. When the cluster rejects a document with a response the plugin treats as non-retryable — a type conflict, a malformed value, a rejected field — the event is written to the dead-letter queue with the failure reason attached instead of being retried forever or dropped. Responses that *are* retryable, such as a rejection caused by back-pressure or a server error, are retried rather than dead-lettered. Its size is bounded by `dead_letter_queue.max_bytes` and a policy that decides what happens when it fills, so it is not an unbounded archive of your mistakes. The part interviews actually probe is the last line of the definition: **the dead-letter queue is only useful if someone reads it.** Nothing drains it automatically. You read it back with a pipeline whose input is the `dead_letter_queue` plugin, fix the offending field, and re-index. Teams that enable the setting, never build that pipeline and never alert on the queue's size have converted silent data loss into silent data loss with extra disk usage. On a district-heating billing platform, "we still have the events, they are in a queue nobody opened for four months" is not an answer that survives a regulator asking for six months of evidence. ## How backpressure reaches all the way back to the file The chain is worth being able to recite, because it is what makes the Elastic stack lose data or not: 1. Elasticsearch rejects or slows writes — usually because indexing threads are saturated. 2. The Logstash `elasticsearch` output retries with backoff and stops draining its batch. 3. Filter workers finish what they hold and cannot hand anything on, so the queue stops draining. 4. The queue fills — to its memory bound, or to `queue.max_bytes` on disk — and the input stops accepting. 5. The `beats` input stops acknowledging, so Filebeat's protocol, which waits for acknowledgement, stops sending and stops reading new lines. 6. Filebeat holds its position in its registry, and the log **file on disk** becomes the buffer. Step 6 is where the real risk lives. The file is a good buffer right up until log rotation removes a file Filebeat has not finished reading, at which point those lines are gone and no queue anywhere downstream ever saw them. That is why a persistent queue sized for the realistic outage window, and retention on the application hosts sized for the same window, are the same decision made in two places. On an 18 GB-a-day stream, a two-hour cluster outage is roughly 1.5 GB to hold somewhere — trivial on disk, fatal if the answer is "in memory".
- The persistent queue is enabled and the destination has been down for hours. What breaks first?The queue reaches `queue.max_bytes` and stops accepting, which is by design. From that moment the input stops acknowledging and the pressure moves back to the shippers, so the real limit becomes how long the application hosts retain their log files before rotation deletes them. Sizing the queue and sizing host-side retention are the same decision, and the smaller of the two is the outage you can actually survive.
- How do you get events out of the dead-letter queue once someone has fixed the cause?Run a second pipeline whose input is the `dead_letter_queue` plugin, pointed at the queue path. It reads the stored events along with the recorded failure reason, so you can correct the offending field in filters and write them to the destination again. Nothing drains it automatically, so the queue's size needs an alert or it silently grows until its bound is reached.
- Why does a persistent queue not protect you against losing the Logstash node itself?Because the queue is a local directory on that node's disk. It survives the process dying and restarting; it does not survive the instance being terminated or the volume being lost. Durability beyond the node means replication somewhere upstream — a broker in front of the tier — not a bigger local queue.
The persistent queue is the parcel locker that holds a delivery until you get home; the dead-letter queue is the pile of undeliverable mail at the sorting office, which helps only if somebody opens it.
saying these in an interview costs you the question
- Confusing the persistent queue with the dead-letter queue
- Expecting the dead-letter queue to drain itself
- Believing a persistent queue survives losing the node
- Thinking a full queue silently discards events instead of blocking
- Assuming Filebeat keeps sending when the pipeline stops acknowledging
- Ignoring log rotation as the real end of the buffer chain