skip to content

On a high-traffic Laravel app, how does a Pulse recorder's sample_rate work, and which recorders is it safe to sample?

level: seniorimportance: should knowfreq 22%

answer

  1. per-recorder sample_rate, default 1
  2. Lottery::odds per event
  3. dashboard scales values, shows ~
  4. sampling happens before the threshold
  5. location backtrace on every query

basics

~20 s

Each Pulse recorder's sample_rate is the chance, per event, that it is considered at all; the dashboard scales sampled counts up and marks them with ~. Sample high-volume recorders such as user requests and cache interactions, not rare events like exceptions.

solid answer

~50 s

Every recorder in `config/pulse.php` has a `sample_rate` (env `PULSE_*_SAMPLE_RATE`, default 1). The `Sampling` concern calls `Lottery::odds($rate)->choose()` for each event, so 0.1 keeps roughly one event in ten, and the dashboard scales the numbers back up and prefixes them with `~`. The accuracy of that estimate depends on volume: for `UserRequests`, `CacheInteractions` or `Queues` on a busy app, 10% of millions is still a precise picture. The slow recorders are different: the lottery runs **before** the threshold check, so a slow query that happens five times a day may simply never be recorded. I would sample the high-volume usage recorders, keep `Exceptions` and the slow recorders at 1, and cut other costs elsewhere: disable recorders I never read, add `ignore` patterns, turn off `PULSE_SLOW_QUERIES_LOCATION` to avoid a backtrace per query, and give Pulse its own `PULSE_DB_CONNECTION`.

code

ini · 5 lines
ini
PULSE_USER_REQUESTS_SAMPLE_RATE=0.1
PULSE_CACHE_INTERACTIONS_SAMPLE_RATE=0.05
PULSE_USER_JOBS_SAMPLE_RATE=0.2
PULSE_SLOW_QUERIES_LOCATION=false
PULSE_DB_CONNECTION=pulse

go deeper

for a junior

Recall that each Pulse recorder has a sample_rate that defaults to 1 and that sampled values appear with a ~ on the dashboard.

for a middle

Explain the per-event lottery, how the dashboard scales values back up, and why accuracy depends on how many events a metric has.

for a senior

Show why sampling slow recorders can hide rare slow events, and cut cost with disabled recorders, ignore patterns, no query location and a separate Pulse database.

for a principal

Set a monitoring budget: decide which signals must stay exact, which can be estimated, and when Pulse's overhead justifies moving ingestion off the request path.

## Why Pulse needs sampling at all **Laravel Pulse** records events as they happen: every routed request for the usage card, every cache hit or miss, every queued job, every query that might be slow. The docs note that by default Pulse captures every relevant event, which on a high-traffic app can mean aggregating millions of rows for the longer dashboard periods. **Sampling** is the per-recorder dial that trades precision for cost. ## How sample_rate works Each recorder entry in `config/pulse.php` carries a `sample_rate`, read from an environment variable such as `PULSE_USER_REQUESTS_SAMPLE_RATE` or `PULSE_CACHE_INTERACTIONS_SAMPLE_RATE`, defaulting to `1`. 1. When an event arrives, the recorder's `shouldSample()` runs `Lottery::odds($sampleRate)->choose()`. 2. With `0.1`, roughly one event in ten wins; the rest are dropped immediately. 3. Winning events are recorded normally and aggregated into buckets. 4. On the dashboard, the card scales the values back up by the sample rate and prefixes them with `~` to show they are approximations. The docs' rule of thumb: the more entries a metric has, the lower you can safely set its rate without losing much accuracy. ## Where sampling fits in the slow recorders For `SlowRequests`, `SlowQueries` and `SlowJobs` the order of checks matters: - `SlowRequests` checks `shouldSample()` before it even resolves the route path or compares the duration. - `SlowQueries` and `SlowJobs` defer their work with `Pulse::lazy()`, and inside that closure test `shouldSample()` first, then the `ignore` list, then the threshold. So sampling is applied to **all** events, not only to the slow ones. If a slow query happens five times a day and the rate is 0.1, there is a good chance none of the five is recorded, and the card shows nothing. | Recorder | Typical volume | Safe to sample? | |---|---|---| | `UserRequests` | every authenticated request | yes | | `CacheInteractions` | every cache hit and miss | yes | | `UserJobs`, `Queues` | every job | usually | | `SlowRequests`, `SlowQueries`, `SlowJobs` | only the rare slow ones matter | with care | | `Exceptions` | ideally rare | rarely | ## Other levers that cut cost without losing signal Sampling is one tool among several: - **Disable recorders you never read**: set `enabled` to false (for example `PULSE_CACHE_INTERACTIONS_ENABLED=false`), or remove the entry. - **Ignore noise**: each recorder's `ignore` list takes regexes for paths, job classes, SQL or cache keys; the shipped config already ignores Pulse's and Telescope's own tables and routes. - **Group keys**: `CacheInteractions` and `SlowOutgoingRequests` accept `groups` regexes that collapse IDs, so fewer distinct keys are stored. - **Skip the query location**: with `location` true, `SlowQueries` walks `debug_backtrace()` for **every** executed query, fast or slow, before deciding anything. `PULSE_SLOW_QUERIES_LOCATION=false` removes that per-query cost, at the price of losing the `file:line`. - **Filter by context**: `Pulse::filter()` drops entries based on anything available at ingest time, such as the current user. - **Separate the database**: `PULSE_DB_CONNECTION` moves Pulse's writes and dashboard aggregation queries off the application database. ## A senior answer in practice On a busy app I would start by measuring which recorders produce the most rows, sample only those high-volume usage recorders (0.1 or lower where traffic is heavy), leave `Exceptions` and the slow recorders at 1, turn off recorders nobody looks at, and move Pulse to its own connection. If the write cost during requests is still visible, the next step is the Redis ingest, which moves storage off the request path. ## Where Pulse's own overhead comes from It helps to separate three costs: 1. **Capture** during the request: listeners run for each event, and `SlowQueries` may take a backtrace per query. 2. **Ingest** after the response: buffered entries are written to the database in one transaction per flush, or pushed to Redis with the Redis ingest. 3. **Dashboard queries**: each card polls and aggregates stored rows for the selected period, which is where millions of unsampled rows hurt most. Sampling reduces all three for the recorders it touches; a dedicated database connection isolates the second and third from the application's own queries.

  • Why does the Application Usage card show ~ in front of its numbers after you enable sampling?
    When a recorder's sample_rate is below 1, only a fraction of events are stored, so the card scales the stored values up by the rate to estimate the true total. The `~` prefix tells the reader the number is an approximation rather than an exact count.
  • Does lowering the SlowQueries threshold make Pulse more expensive?
    Somewhat: more queries pass the threshold, so more entries and aggregates are written. The bigger fixed cost is often the `location` option, which takes a backtrace for every executed query regardless of the threshold; turning it off saves that work on every query, not only the slow ones.

Sampling is like a traffic survey that counts one car in ten on a motorway and multiplies by ten: accurate for the busy road, useless for spotting the one lorry that breaks down each week.

saying these in an interview costs you the question

  • sample_rate only applies to events that already crossed the slow threshold.
  • Sampling the Exceptions recorder at 0.1 still shows every distinct exception.
  • A sample_rate of 0.1 makes Pulse store one aggregate bucket in ten.
  • The dashboard shows raw sampled counts, so totals drop tenfold.
  • Raising the SlowQueries threshold removes the per-query backtrace cost.