In Sentry, what actually reduces the event volume a project ingests, and why should an issue alert fire on a new or regressed issue rather than on every event?
answer
- two sides of the wire
- drop classes, not fractions
- before-send versus inbound filters
- new and regressed, not counts
- know what you threw away
basics
~20 sVolume is cut in the SDK -- an error sample rate, ignore lists, a before-send hook that drops the event -- and at ingest by inbound filters, client-key rate limits and spike protection. Alert on new or regressed issues, not event counts.
solid answer
~50 sVolume is shaped in two places. In the SDK you drop events before they leave the process: an error `sampleRate`, a `tracesSampleRate` for transactions, ignore and deny lists, and a before-send hook that returns nothing for a class you have decided is noise. At ingest, Sentry offers inbound filters, rate limits on a project's client key and spike protection, which cost you the round trip but change without a deploy. Prefer dropping **classes** -- bot traffic, third-party script errors, validation failures, a preview environment -- over sampling everything down, because errors are rare and uneven and a uniform sample rate hides exactly the bug you are hunting. For alerts, conditions keyed to an issue being new or regressing fire once per meaningful change; an event-count threshold mostly tracks traffic and re-fires on issues you triaged last week.
code
javascript · 12 linesSentry.init({
dsn: process.env.SENTRY_DSN,
release: process.env.SENTRY_RELEASE,
sampleRate: 1.0, // keep every error
tracesSampleRate: 0.01, // sample transactions instead
ignoreErrors: ["ResizeObserver loop limit exceeded"],
beforeSend(event, hint) {
const err = hint.originalException;
if (err && err.name === "PermitFormValidationError") return null;
return event;
},
});go deeper
Know that not every error a Sentry SDK sees has to be sent, and that noise such as browser-extension failures is normally filtered out deliberately rather than treated as a bug to fix.
Be able to name the levers and say which side of the wire each sits on: sample rate, ignore lists and before-send inside the SDK; inbound filters, key rate limits and spike protection at ingest. Explain what each one actually saves.
Show judgement about what to stop sending and what it costs afterwards -- counts becoming estimates, filtered events being unrecoverable -- and explain why alert conditions keyed to a new or regressed issue beat an event-count threshold.
Own the trade across teams: how the budget is split between projects, which classes of event are filtered by policy rather than by each team's taste, and how you stop a team quietly sampling away the signal that would have caught its own outage.
## Two places volume is decided, and they are not equivalent Once a monitoring quota becomes a topic, the interesting question is not "how do we send less" but "what do we stop sending, and where". Sentry offers knobs on both sides of the wire, and they save different things. | Where | Knobs | What it saves | Cost of changing it | | --- | --- | --- | --- | | In the SDK | error `sampleRate`, `tracesSampleRate` for transactions, ignore and deny lists, a before-send hook returning nothing | the event never leaves the process: no serialization, no round trip, nothing counted | a deploy, repeated across every service | | At ingest | inbound filters, rate limits on a project's client key, spike protection | quota and storage, centrally and immediately | none -- but the client already paid the round trip | The rule that follows is **drop classes, not fractions.** A uniform error sample rate is the crudest lever available and it is usually the wrong first move, because errors are not a dense uniform stream the way requests are. Sampling ten percent of a rare but serious exception means you probably never see it, and it turns every issue's event count into a scaled estimate rather than a fact. Sampling *transactions* is a different matter: they are dense and repetitive, so a fraction is genuinely representative of the whole. ## Deciding what a team stops sending The candidates are nearly always the same, and none of them is a real bug in your code: - browser noise no code path can fix -- failures injected by extensions, errors from third-party scripts served cross-origin, the well-known `ResizeObserver` warnings; - validation errors that represent a user typing something wrong rather than software breaking; - traffic from bots and scanners hitting endpoints that were never meant to exist; - one pathological issue that dwarfs everything else while it waits to be fixed -- better rate-limited at its client key than allowed to consume the whole allowance; - whole environments: local development and short-lived preview deployments reporting into the production project. Consider a municipal parking-permit service running a mid-quarter migration off a hosted vendor. Two stacks are live at once, the legacy one emits connection errors that are expected and understood, and the renewal deadline pushes the front end to a 5,400-request-per-second peak. The right move is not to sample everything down until it fits. It is to drop the known-expected class in the SDK's before-send hook, rate-limit the client key belonging to the legacy stack so it cannot starve the replacement, and leave the error sample rate alone -- the unknown failures are precisely why you are paying for the tool during a migration. Whatever you drop, know that you dropped it. The SDKs report back what they discarded and why, so client-side drops appear as a number rather than as silence, and spike protection records what it shed. An engineer who cannot tell "there were no events" from "we threw those events away" will misread the next incident badly. ## Alerting on state changes, not on counts Sentry's issue alerts can be triggered by an issue's **state changing** -- a new issue is created, or a previously resolved issue regresses -- or by volume, such as an issue being seen more than a certain number of times or affecting more than a certain number of users within a window. The first family is what an error monitor knows that a plain metric threshold does not. 1. **A new issue means a fingerprint the project has never seen.** It fires once, when something genuinely new appears, which is exactly what you want to know after a deploy. 2. **A regression is a state transition.** Somebody closed it; it came back. That is information, not volume. 3. **An event-count threshold mostly tracks traffic.** It fires on a busy afternoon for an issue you triaged last week, stays silent on a low-traffic service where three occurrences matter, and re-fires every time the count crosses the line again. Volume conditions still have a place -- "affecting more than a few hundred users in an hour" is a reasonable severity signal, and it is the right shape for an issue you already know about. But as the *default* trigger they produce repetition rather than news, and the repetition is what teaches people to ignore the channel. ## Be honest about what it costs Every mechanism here trades completeness for signal: - event counts stop being a census, so "how many users hit this" becomes an estimate whose scale factor you must remember; - an event filtered at ingest cannot be recovered by changing your mind later; - a class dropped in the SDK is invisible until someone reads the configuration, so the filter list needs the same review as any other production rule. That is the trade worth stating out loud in an interview: a smaller, deliberately chosen stream that people actually read, against a complete one that nobody does.
- Why is a low error sample rate riskier than a low trace sample rate?Errors are rare and unevenly distributed, so a uniform rate is most likely to discard the three occurrences of the serious bug while keeping the thousands of the trivial one, and every issue's count becomes a scaled estimate. Transactions are dense and repetitive, so a small fraction still represents the distribution. Where volume must come down on the error stream, drop known-worthless classes outright instead of sampling everything.
- What is the practical difference between an SDK ignore list and a server-side inbound filter?The ignore list runs inside the process, so the event is never serialized or sent -- it saves the round trip as well as the storage, but changing it needs a deploy in every service. An inbound filter is applied after the event arrives, so the client has already paid the network cost, but you can change it immediately and centrally for a whole project without shipping code. Most teams need both.
- Why is "seen more than N times in an hour" a poor default Sentry alert condition?Event counts track traffic more than they track breakage. The condition fires on a busy afternoon for an issue somebody already triaged, never fires at all on a low-traffic service where three occurrences are serious, and re-fires each time the count crosses back over the line. Conditions keyed to state changes -- the issue is new, or it regressed after being resolved -- fire once per meaningful change instead.
Filtering is choosing which smoke detectors to unplug: the one over the toaster, never the one in the server room.
saying these in an interview costs you the question
- Turns down the error sample rate as the first response to volume
- Alerts on every event rather than on new or regressed issues
- Assumes dropped or filtered events can be recovered afterwards
- Thinks an SDK ignore list and an inbound filter save the same thing
- Treats a client-key rate limit as a form of deduplication
- Samples errors and still quotes the event count as an exact number