skip to content

How does an AWS X-Ray sampling rule decide whether a given request is traced? Explain the reservoir and fixed-rate fields, how rule priority works, and how one reservoir is shared across many instances of the same service.

level: seniorimportance: nice to knowfreq 33%

answer

  1. a floor and a percentage above it
  2. per second, and per service
  3. lowest number is evaluated first
  4. the service hands out quotas
  5. new instances borrow until told otherwise

basics

~20 s

An X-Ray sampling rule matches requests by service, host, method and URL path, then traces a per-second reservoir of them and a fixed percentage of the rest. Rules are evaluated by priority, lowest number first, and the reservoir is handed out to instances as quotas by the X-Ray service.

solid answer

~50 s

Each rule has matchers — `ServiceName`, `ServiceType`, `Host`, `HTTPMethod`, `URLPath`, optionally `ResourceARN` and attributes — plus two knobs. `ReservoirSize` is a guaranteed number of traces **per second**; `FixedRate` is the fraction of requests sampled *after* the reservoir for that second is exhausted. So `ReservoirSize: 5, FixedRate: 0.05` means the first five matching requests each second are always traced, then five percent of the remainder. Rules are ordered by `Priority`, lowest number wins, and the first match applies; a built-in default rule sits at the bottom with a reservoir of one request per second and a five percent rate. The reservoir is a service-wide budget, not per instance: SDKs and the daemon poll X-Ray for the rule set and periodically report usage, and X-Ray hands each client a share of the reservoir as a quota. Until a client has a quota it may borrow one trace per second.

code

json · 15 lines
json
{
  "SamplingRule": {
    "RuleName": "checkout-post",
    "Priority": 100,
    "ReservoirSize": 5,
    "FixedRate": 0.10,
    "ServiceName": "checkout-api",
    "ServiceType": "*",
    "Host": "*",
    "HTTPMethod": "POST",
    "URLPath": "/orders*",
    "ResourceARN": "*",
    "Version": 1
  }
}

go deeper

for a junior

Know that X-Ray does not trace everything, and that a rule combines a guaranteed few traces per second with a percentage of the rest.

for a middle

Explain reservoir versus fixed rate precisely, that priority is evaluated lowest number first with first match winning, and that an undeletable default rule sits at the bottom.

for a senior

Describe the distributed part — clients poll for rules, report usage, receive a quota share of the reservoir, and borrow before their first quota — and use it to reason about lag during scaling and rule changes.

for a principal

Own the sampling policy as a cost and coverage tradeoff across the estate: guaranteed floors on customer-facing paths, zero on health checks, and a discipline for temporarily raising rates during incidents and lowering them again.

## Why the two knobs exist A plain percentage is a bad sampling policy at both ends of the traffic range. At 5% a low-traffic endpoint that gets ten requests a minute may produce no traces at all for hours — exactly the endpoint you most want examples of. At 5% a hot endpoint doing 20,000 requests per second produces 1,000 traces per second, which is expensive and adds nothing over 50. X-Ray's answer is a **reservoir plus a rate**: - `ReservoirSize` — a fixed number of matching requests **per second** that are always traced. This is the floor: quiet endpoints still yield examples. - `FixedRate` — a fraction between 0 and 1, applied to matching requests *after* that second's reservoir is used up. This is the ceiling-shaper: busy endpoints scale proportionally rather than linearly with traffic. `ReservoirSize: 1, FixedRate: 0.05` — the values of the built-in default rule — reads as "always give me one trace a second, plus 5% of everything else". ## Matching and priority A rule matches on the fields of the incoming request as the sampler sees them: ```json { "RuleName": "checkout-post", "Priority": 100, "ReservoirSize": 5, "FixedRate": 0.10, "ServiceName": "checkout-api", "ServiceType": "*", "Host": "*", "HTTPMethod": "POST", "URLPath": "/orders*", "ResourceARN": "*", "Version": 1 } ``` `*` is a wildcard and `?` matches a single character, so most rules are mostly wildcards with one or two meaningful matchers. **`Priority` orders evaluation, lowest number first, and the first matching rule wins** — later rules are not consulted, so a broad rule with a low priority number will shadow every specific rule beneath it. That is the single most common configuration mistake: someone adds a catch-all at priority 1 to "trace everything for a day" and every carefully-tuned rule below it goes dead. The **default rule** always exists, cannot be deleted, and sits below everything else. If you never configure anything, that is what you are running. ## How one reservoir works across a fleet The interesting part. `ReservoirSize: 5` means five traces per second *for the service*, not five per instance — otherwise a fleet that scaled from 4 to 400 tasks would silently multiply its tracing bill by a hundred. Making that work requires coordination, and X-Ray does it centrally: 1. Clients (the SDK's sampler, or the daemon acting on their behalf) call `GetSamplingRules` to fetch the rule set, and refresh it periodically so console changes take effect without a deploy. 2. Clients then call `GetSamplingTargets` on a short cycle, **reporting how many requests they saw and sampled**. 3. X-Ray divides the reservoir across the reporting clients in proportion to the traffic each reported, and returns each one a **quota** valid until an expiry time. 4. A client that has just started and has no quota yet may **borrow** — one trace per second — so a fresh instance is never blind while it waits for its first target. When a client cannot reach the sampling API at all, it falls back to local rules rather than dropping tracing entirely. The practical consequences: rule changes propagate on the order of minutes, not instantly; reservoirs are only honoured approximately during rapid scaling events, because quota allocation lags traffic; and a service whose instances see wildly uneven traffic will see the reservoir concentrate on the busy ones, which is usually what you want. ## Where sampling actually happens The decision is made **once, at the entry point of the trace**, and then propagated in the `Sampled` field of the trace header. Downstream services honour an incoming decision rather than re-deciding. That is what keeps a sampled trace complete end to end — sampling per service would produce traces with holes in the middle, which are worse than no trace at all. It also means the rule that matters for a given request is the rule evaluated by the *first* instrumented component in the chain. Tuning a rule on a downstream service will not change how often it appears in traces started upstream. ## Tuning it in production Good defaults: leave the default rule as a safety net; add rules with a healthy reservoir for low-traffic but high-value paths (checkout, auth, anything with a customer-facing SLO); add a rule with a low or zero rate for the endpoints that flood without informing — health checks especially, which can otherwise dominate your traces and your bill. And when you raise sampling to chase an incident, raise it on a *specific* rule with a *specific* priority, and put a calendar reminder on lowering it again.

  • Health-check requests are dominating your traces. How do you deal with that with sampling rules?
    Add a rule that matches the health-check path — a `URLPath` matcher such as `/healthz` — with `ReservoirSize` 0 and `FixedRate` 0, and give it a priority number lower than your general rules so it is evaluated first. The requests still serve, they just stop being traced. It is a cheap and large win, because health checks are usually the highest-frequency and least-informative traffic a service handles.
  • You changed a sampling rule in the console. Why did nothing happen for several minutes?
    Clients do not receive a push. The samplers poll for the rule set on an interval and separately report usage to obtain their reservoir quota, so a change takes effect only after the next refresh across the fleet. It is a good property — no deploy needed — but it means you should not stare at the console expecting an instant change, and you should not stack multiple edits inside one polling interval while debugging.
  • Does raising the sampling rate on a downstream service make it appear in more traces?
    No. The sampling decision is made once, by the first instrumented component in the chain, and propagated in the trace header's Sampled field; downstream services honour it rather than re-deciding. To see more traces that include a downstream service, raise sampling on the entry point whose requests reach it. Re-deciding downstream would produce traces with gaps in the middle, which is why the SDKs do not do it.

saying these in an interview costs you the question

  • Thinking the reservoir is per instance rather than per service
  • Assuming rules are evaluated highest priority number first
  • Believing every service re-runs sampling for the same request
  • Expecting rule changes to take effect instantly
  • Treating fixed rate as applying to all requests including the reservoir

context