skip to content

In Graylog, how would you give one team alerting on and read access to only its own logs?

level: seniorimportance: nice to knowfreq 30%

answer

  1. One boundary serves two purposes
  2. Alerts are searches on a schedule
  3. Read access follows the routing
  4. One stream holds everything by default

basics

~20 s

Route that team's messages into their own Graylog stream, grant read on just that stream through a role or an entity share, and define event definitions whose search is scoped to it. Alerting and access both hang off the stream boundary.

solid answer

~50 s

Graylog hangs both alerting and access off the same object: the **stream**. Scope a team's messages into their own stream with stream rules, then create an **event definition** whose filter is a search query restricted to that stream, evaluated on a schedule over a search window, optionally aggregating - a count or an average, grouped by a field - and firing when the condition holds. A firing event creates a Graylog **event** and triggers a **notification**, an email or an HTTP call, which may carry a backlog of the matching messages. Access uses the same boundary: a role or a per-entity share grants read on that stream, while the saved searches and dashboards built on it are shared separately. The catch is the **default stream** - every message enters it, so broad read access there defeats the whole design.

code

text · 1 line
text
source:checkin-api-* AND scan_result:rejected AND NOT gym_id:"bristol-north"

go deeper

for a junior

Know that Graylog alerts are defined as searches that run on a schedule rather than as code inside the application, and that what a user can see is decided by which streams they are permitted to read.

for a middle

Explain the parts of an event definition - the filter query, the streams it is scoped to, the search window and execution interval, an optional aggregation with a threshold condition, and the notification it triggers.

for a senior

Show that access and alerting share one boundary: routing must be right before either works, the default stream still holds everything, a notification backlog moves log content outside the permission model, and a quiet stream reads as good news.

for a principal

Own the model across teams: who may create streams and event definitions, how sensitive fields are kept out of shared streams at parse time, and whether one shared cluster with per-stream access beats separate deployments per tenant.

## Alerting in Graylog is a scheduled search An **event definition** of the filter-and-aggregation kind is, at heart, a saved search that runs on a timer. You give it a **filter query** in Graylog's search syntax, the **streams** it is scoped to, a **search window** (how far back each run looks), an **execution interval** (how often it runs), and optionally an **aggregation**: a function such as a count or an average, an optional set of grouping fields, and a threshold condition on the result. When the condition holds, Graylog creates an **event** - a stored record in its own right - and the event triggers whatever notifications are attached to the definition. The scheduled shape explains almost every surprise people hit: - Nothing is evaluated per message. A message indexed just after its window was searched is invisible to that execution. - If the search window is shorter than the execution interval, coverage has gaps; if it is longer, executions overlap and the same messages can raise an event twice. - The definition searches the index, so anything delaying indexing - a processing backlog, a slow node - delays detection by the same amount, and clock skew on a sending host can push its messages outside the window entirely. ## The parts of an event definition | Part | What it decides | What goes wrong when it is wrong | |---|---|---| | Filter query | Which messages count | A misspelled field name makes it a silent no-op | | Streams in scope | Whose data is examined | It fires on another team's noise | | Search window | How far back each run looks | Gaps in coverage, or duplicate events | | Execution interval | How often it runs | Detection latency nobody budgeted for | | Aggregation condition | The threshold that fires | An event for every single occurrence | An aggregation can group by a field, so one definition covers many services or many sites rather than needing a copy of itself for each. ## Notifications carry log content A firing event triggers one or more **notifications** - an email, or an HTTP call to a receiving system - and a notification can include a **backlog** of the messages that matched. The backlog is what makes an alert actionable without opening the UI, and it is also the moment log content leaves Graylog's permission model: whatever those messages contained is now in a mailbox or in a third-party receiver. On a stream carrying member records that is a decision to review alongside the stream's read permissions, not a formatting preference. ## Access control rides on the same stream boundary A stream is not only a routing category; it is the unit that read permission is expressed against. Graylog ships built-in **Admin** and **Reader** roles, and beyond those, access is granted per entity: a role or a share on a stream lets a set of users search that stream, while the saved searches and dashboards built on top of it are separate entities shared in their own right. The design that falls out of this: - Decide the stream boundary first. It is simultaneously the storage boundary, the alerting boundary and the access boundary, and you only get to draw it once. - Give each team a role that reads only its own streams. Do not hand out Admin to settle a visibility complaint. - Remember that sharing a dashboard does not grant the underlying stream, and granting a stream does not surface anyone's dashboard. - Keep genuinely sensitive fields out of shared streams at parse time, with a pipeline rule that removes or masks them, rather than hoping to hide them at read time. ## The default stream is where the model leaks Every message enters the default stream. If a team can read it, every per-stream grant elsewhere is decoration, because the same content is searchable there. Two habits keep the model honest: have each team stream remove its matches from the default stream, and treat default-stream read as an administrative privilege rather than a convenience. There is a second, quieter failure. An empty stream is indistinguishable from a healthy one with nothing wrong. If a team's stream stops receiving messages - a renamed field, a sender pointed at the wrong input, a stream rule edited in a hurry - their alerts go quiet and read as good news. The fix is an event definition that fires when the count for that stream over a window falls to zero, so that absence is itself alertable. ## A worked cutover A climbing-gym membership platform migrating off a hosted vendor mid-quarter had two months of overlap on a 7-node cluster. The access model came first: 11 team streams, each removing its matches from the default stream, each with a role granting read on that stream and nothing else, and default-stream read left with the two platform engineers. Alerting was then layered on the same boundary - a definition per team scoped to its own stream, aggregating over a 15-minute window with a grouping field so one definition covered all 34 sites. The definition that earned its keep was not a failure alert at all: it was the zero-volume one on each team stream, which caught two senders still pointed at the old vendor 11 days after they should have moved.

  • Why can a Graylog event definition miss matching messages even when its search query is correct?
    It runs on a schedule over a bounded search window, so anything indexed after that window was searched - a late sender, a processing backlog, clock skew on the source host - falls outside it. Making the window longer than the execution interval buys overlap and therefore fewer misses, at the price of the same messages raising an event twice.
  • What does including a message backlog in a Graylog notification cost you?
    The notification carries the matching messages themselves, which is what makes it actionable without opening the UI, and which puts log content into a mailbox or a third-party receiver. On a stream holding member records that is an access-control decision, so backlog size and destination deserve the same review as the stream's read permissions.
  • How do you stop a per-team Graylog stream from silently going quiet?
    An empty stream looks exactly like a healthy one with no incidents. Add an event definition that fires when the message count for that stream over a window falls to zero, so absence of data is itself alertable, and re-check the stream's throughput after any change to its stream rules or to the inputs feeding it.

The stream is the door: it decides both which alarms watch the room behind it and who is issued a key.

saying these in an interview costs you the question

  • Thinks a Graylog event definition inspects every message as it arrives
  • Grants a team the Admin role so they can see their own logs
  • Forgets the default stream still holds every message
  • Assumes alerting works before messages are routed into a stream
  • Believes stream access automatically grants the dashboards built on it
  • Treats an empty stream as proof that nothing is wrong