How would you halve a Datadog bill dominated by custom metrics and indexed logs without losing incident coverage?
answer
- Two billed units dominate everything else
- Custom metrics counted per name and tag combination
- Logs billed for ingestion and again for indexing
- Decide after ingest which logs get indexed
- Cut dimensionality before cutting whole signals
basics
~20 sAttack the two billed units directly: custom metrics, counted per unique metric-name-and-tag-value combination, and indexed logs, billed separately from ingestion. Bound what tags may contain, aggregate before submitting, and decide after ingest which logs are worth indexing.
solid answer
~50 sDatadog's two expensive levers are **custom metrics** and **indexed logs**, and they are cut by different means. A custom metric is billed per unique combination of metric name and tag values, so the lever is dimensionality: drop tags carrying an unbounded identifier, aggregate in the application before submitting, and use the platform's ability to keep an ingested custom metric while dropping tags from what is queried. Logs are billed once for ingestion and again for indexing, the design that lets you send everything then decide: **exclusion filters** on a log index drop or sample matching records out of the index while they remain ingested, so debug-level noise from a healthy service costs ingestion only. Long-tail retention moves to archives with rehydration on demand. APM has the same shape, with its own ingestion sampling and retention filters. The judgement is which signals must be queryable instantly versus reconstructable later.
go deeper
Know that Datadog charges for custom metrics your code submits and for logs you keep searchable, and that adding a tag with many possible values is the usual way a bill grows unexpectedly.
Be able to explain the two billed units: a custom metric as a unique metric-name-and-tag-value combination, and logs charged for ingestion separately from indexing. Name the controls - bounded tags, aggregation before submission, exclusion filters on an index.
Show you would measure first and cut by contribution, protect anything a monitor evaluates, prefer dropping dimensions to dropping signals, and know where unindexed data still lives - the live stream, archives, rehydration.
Own it as policy: per-team visibility of consumption so cost sits with the teams creating it, defaults that make the cheap thing the easy thing, and a written record of the blind spots you accepted and how each one is recovered during an incident.
## Where the money actually goes Datadog charges for a lot of things, but on most bills two lines dominate, and both grow with engineering behaviour rather than with traffic: 1. **Custom metrics** - the metrics your own code submits, counted by **unique combinations of metric name and tag values**. One metric name with a bounded set of tags is cheap. The same name with a tag whose values are unbounded is not, and the growth is multiplicative across tags rather than additive. 2. **Indexed logs** - logs are billed once for **ingestion** (getting them into the platform) and again for **indexing** (making them searchable and retaining them). This split is deliberate and is the single most useful thing to know about Datadog cost, because it means the expensive decision happens *after* the data has arrived, not at the emitting service. Take a ferry-timetable booking platform with a 41-service estate whose telemetry budget was just halved. The instinct is to turn off collection service by service, which trades money for blindness in exactly the places nobody is watching. The better move is to keep collecting and change what is *retained and dimensioned*. ## The custom metric levers - **Bound what a tag may contain.** A tag holding a booking reference, a session identifier or a raw URL path is the classic budget destroyer. Replace the identifier with a bounded category: a route template rather than a path, a status class rather than a code, a ferry route code rather than a passenger id. - **Aggregate before submitting.** Submitting one point per event and letting the platform aggregate is a design choice you are paying for. Aggregating in the application, or letting the Agent's flush interval do it, reduces submission volume without changing what a dashboard shows. - **Watch metric types that expand.** A histogram submitted through DogStatsD is aggregated Agent-side into several sub-metrics - count, average, median, maximum and a percentile - so one submission name becomes several billed metrics for every tag combination. Knowing this before the invoice arrives is the difference between a designed system and an accident. - **Drop dimensions after ingest.** The platform can keep a custom metric while restricting which tags it is queryable by, so a metric that must exist for an alert does not have to be queryable by every dimension it was submitted with. - **Delete what nobody queries.** Metrics outlive the dashboards that motivated them. A periodic sweep for metrics with no queries and no monitors referencing them typically finds a surprising amount. ## The log levers | Control | What it changes | What you keep | |---|---|---| | Exclusion filter on an index | Matching records are not indexed | Still ingested; visible in the live stream | | Sampling inside an exclusion filter | A fraction of matching records is indexed | Statistical coverage of high-volume noise | | Multiple indexes with different retention | Short retention for noisy sources | Long retention only where it is needed | | Archives to object storage | Cheap durable storage outside the index | Rehydration on demand for audits | | Metrics generated from logs | A cheap counter replaces a search | Trend visibility without indexing volume | The pattern is always the same: **ingest broadly, index narrowly, archive the rest**. Debug-level output from a service that is behaving normally does not need to be searchable at three in the morning six weeks later; it needs to be available while the incident is live and recoverable afterwards if a regulator asks. ## Making the call A halving exercise is a prioritisation exercise, and it is a principal-level question because there is no correct answer, only a defensible one. The way to structure it: 1. **Measure before cutting.** Find the metric names and log sources contributing the most billed units. In almost every estate the distribution is brutally skewed and a handful of sources account for most of the bill. 2. **Separate alerting signals from investigation signals.** Anything a monitor evaluates must stay at full fidelity - cutting it converts a cost problem into an outage you do not detect. Investigation signals can be sampled, archived, or reconstructed. 3. **Cut dimensionality before cutting signals.** Dropping a tag usually saves more than dropping a metric, and costs less coverage. 4. **Give the saving back as a budget.** Publish per-team consumption so the teams that create the cost can see it. Central cuts are re-added within a quarter; visible ownership is not. 5. **Accept named blind spots.** Write down what you stopped indexing and what the recovery path is - rehydration, a live stream, a log source that can be turned back up during an incident - so the decision is a documented tradeoff rather than a surprise. The answer that fails is *turn off tracing on the least important services*. It saves little, because those services are not what is expensive, and it costs exactly the correlation the platform was bought for.
- What do you lose when you drop a tag from a custom metric after ingest?Any query, dashboard or monitor that groups or filters by that tag stops working from that point forward, because the dimension is no longer available to query. What you keep is the metric itself at lower dimensionality, which is usually enough for alerting on the aggregate. The migration order matters: find the monitors that reference the tag first, or the cost saving lands as a silent alerting gap.
- How do you keep an audit-grade record of logs you have stopped indexing?Route them to archives in your own object storage, where retention is cheap and can be set to whatever a compliance requirement demands, and rely on rehydration to bring a bounded time range back into a searchable index when someone actually needs it. During a live incident the unindexed stream is still visible as it arrives, so the loss is historical searchability rather than visibility.
- A team objects that sampling their logs makes debugging impossible. How do you answer?Distinguish the cases. For error and warning records, do not sample - they are low volume and high value. For high-volume routine records, a sample preserves the statistical picture and the rare interesting line is usually reachable through the trace that produced it. Then commit to a specific escape hatch: the ability to raise indexing for their source during an incident, and rehydration afterwards.
Ingestion is the postal delivery and indexing is the filing cabinet: it is cheap to receive every envelope and expensive to file all of them, so you decide at the desk which ones are worth a folder.
saying these in an interview costs you the question
- Counts a custom metric per metric name, ignoring tag combinations
- Says the fix is to stop collecting logs from less important services
- Believes a log that is not indexed has been lost entirely
- Thinks a DogStatsD histogram costs exactly one custom metric
- Proposes shortening every retention window as the only lever
- Cuts a signal an existing monitor evaluates without checking first