In Jaeger, what does serving sampling configuration centrally to clients solve, and what can it not solve?
answer
- The rate lives outside the deployed binary
- Clients poll; nobody pushes to them
- Probabilistic and rate-limiting are different units
- It counts traces started, not spans
- A silent default when the fetch fails
basics
~20 sJaeger clients poll the backend for their sampling rate, so per-service and per-operation rates change without a redeploy. Central configuration cannot pick traces by outcome, cannot bound spans per trace, and does nothing for a client that never fetches it.
solid answer
~50 sRather than baking a rate into each service, Jaeger clients can be configured to poll the backend for one. The collector reads a JSON strategies file (`--sampling.strategies-file`) with a `default_strategy` plus `service_strategies` and `operation_strategies`, each a `type` of `probabilistic` (a probability) or `ratelimiting` (traces per second). Operationally this buys real things: change a rate during an incident without a build, pin a health-check path near zero while sampling a checkout path fully, and cap a single noisy service's contribution. What it cannot do matters as much. The decision is taken when the trace starts, so it cannot preferentially keep the request that errored. It is polled, so changes land a poll interval later with no acknowledgement. A client that never fetches it keeps its own built-in default and nothing surfaces that. And it bounds *traces started*, not spans — one wide fan-out still lands whole.
code
json · 19 lines{
"service_strategies": [
{
"service": "seed-catalogue-order-api",
"type": "probabilistic",
"param": 0.08,
"operation_strategies": [
{ "operation": "reserveStock", "type": "probabilistic", "param": 1.0 },
{ "operation": "healthz", "type": "probabilistic", "param": 0.0 }
]
},
{
"service": "bulk-import-worker",
"type": "ratelimiting",
"param": 2
}
],
"default_strategy": { "type": "probabilistic", "param": 0.003 }
}go deeper
Recall that a service's sampling rate does not have to be compiled into it. Jaeger clients can ask the backend what rate to use, and that configuration lives in one JSON file rather than in each service's code.
Explain the mechanics: the client polls, the file has a default plus per-service and per-operation overrides, and the two strategy types express a probability and a traces-per-second cap respectively. Be clear that both count traces, not spans.
Demonstrate the operational judgement. Say why you would raise a rate mid-incident, why a rate limit suits a service whose volume you do not control, and be able to diagnose the silent fallback when a client never fetches a strategy at all.
Own the policy. Decide who is allowed to change the file, how you stop one team's volume from consuming the ingest budget, and where hand-tuned per-operation rates should give way to adaptive sampling in an estate too large to curate by hand.
## Where the sampling rate lives In the simplest deployment every service is built with its rate baked in — sample one trace in a thousand, say — and changing it means a code change and a redeploy of every service affected. Jaeger's alternative is a **remote sampler**: the client is configured only with "ask the backend", polls an HTTP endpoint at a fixed interval, and applies whatever strategy comes back. In Jaeger v1 the local `jaeger-agent` served that endpoint on port `5778`; the collector is started with a strategies file (`--sampling.strategies-file`) and that file is what gets served. The file is JSON. It carries a `default_strategy` applied to any service it does not name, and a list of `service_strategies` that override it, each of which may carry `operation_strategies` that override it again for individual endpoints. A strategy has a `type` and a `param`: | `type` | What `param` means | Behaviour when traffic spikes | |---|---|---| | `probabilistic` | Probability a new trace is sampled, `0.0` to `1.0` | Sampled volume rises in proportion to traffic | | `ratelimiting` | Traces started per second for that service | Sampled volume is capped regardless of traffic | Note the units carefully, because this is where candidates slip: both knobs govern **traces started**, not spans emitted. ## What it genuinely solves - **Rate changes without a deploy.** During an incident on a seed-catalogue ordering service you can lift one service from a background rate to full capture, wait one poll interval, and get traces — with no build, no rollout, no restart. - **Per-service and per-operation granularity from one place.** A health-check endpoint handling most of the request count can be pinned near zero while a checkout path that must stay inside a 320 ms p99 budget is sampled at one hundred per cent, and the two live in the same reviewable file. - **A cap on a noisy neighbour.** When one team generates seventy per cent of the volume, a `ratelimiting` strategy on their services bounds what they can contribute to ingest no matter what their traffic does — something a single global probability cannot express. - **The rare-endpoint problem.** A fleet-wide probability of 0.003 samples an endpoint serving eleven requests an hour essentially never. A per-operation override, or adaptive sampling, gives that endpoint a floor. - **An auditable record.** The effective sampling policy for the whole estate is one artefact in version control rather than a property scattered across dozens of repositories. ## What it cannot solve 1. **It decides at trace start.** The strategy is applied when a trace begins, before anything is known about how the request turned out. It cannot preferentially keep the request that errored or the one that blew the latency budget, because at decision time neither has happened yet. 2. **It is polled, not pushed.** A change takes effect one poll interval later, per process, and there is no acknowledgement. During a rollout some replicas are on the old strategy and some on the new, and nothing in the UI tells you which produced a given trace. 3. **A service that never asks keeps its own default.** If the remote sampler is not configured, points at the wrong host, or is blocked by a network policy, the client silently falls back to its built-in rate. This failure is invisible: you see traces, just not the proportion you think you configured. 4. **It bounds traces, not span volume.** Two traces per second sounds cheap until one of them is a fan-out producing four thousand spans. Nothing in a strategies file limits the width or depth of a sampled trace, so ingest cost is only loosely coupled to the number you set. 5. **It does not make the decision consistent by itself.** A rate configured centrally still has to be honoured the same way at every hop for a trace to arrive whole; a service that makes its own fresh decision produces fragments regardless of what the file says. 6. **It cannot express a condition.** "Sample everything for customers on the new pricing plan" is not something `type` and `param` can say. ## Adaptive sampling Instead of reading a static file, the collector can compute strategies itself: it observes the throughput of each service and operation, and adjusts each probability so the resulting sampled traces converge on a target rate. That is what stops a newly launched endpoint from being invisible and stops a suddenly popular one from flooding ingest, without anyone editing a file. The cost is real — the collectors need somewhere to keep that state, so it is available only where the backend supports it, and the effective probability for any given operation becomes a moving number you have to look up rather than one you can read off a config. Treat it as the thing you reach for when the estate is too large to hand-tune, not as the default.
- When would you choose a rate-limiting strategy over a probabilistic one for a service?When you care about bounding cost rather than preserving statistical proportion. A probability keeps a fixed share of traffic, so a spike multiplies your ingest; a rate limit caps traces per second whatever the load does. Use it for a service whose volume you do not control — a batch worker, or a team generating most of the fleet's spans — and probabilistic where you need representative sampling for analysis.
- A team swears they set their service to full sampling, yet you see almost no traces from it. What do you check?Whether the client is actually fetching the strategy at all. If the remote sampler is unconfigured, pointed at the wrong host, or blocked by a network policy, it falls back to its built-in default and reports no error. Confirm the file names that exact service string, that the process polls the endpoint successfully, and that the replicas have been running longer than one poll interval.
- Why does a rate-limiting strategy not bound your storage bill?Because it limits traces started, not spans stored. Two traces per second is cheap if each has twelve spans and expensive if a fan-out produces four thousand. Storage tracks sampled traces multiplied by mean spans per trace multiplied by span size, and nothing in the strategies file constrains the second term — you have to control that through instrumentation depth instead.
saying these in an interview costs you the question
- Thinks the backend pushes strategies to clients instantly
- Reads the rate-limiting param as a percentage
- Claims central config keeps the error traces you want
- Assumes an unreachable sampling endpoint stops tracing entirely
- Believes capping traces per second caps span volume