In New Relic, how does a NRQL query select, filter, facet and bucket results over time?
answer
- Aggregation happens when you ask
- The clause list reads almost like SQL
- One clause groups by an attribute's values
- Another clause cuts the window into buckets
- FACET groups, TIMESERIES buckets, SINCE bounds
basics
~20 sNRQL aggregates records at query time: SELECT chooses the aggregation, FROM the data type, WHERE filters, FACET splits results by an attribute's values, SINCE and UNTIL set the window, and TIMESERIES cuts that window into buckets.
solid answer
~50 sA NRQL query reads records out of NRDB and aggregates them when you ask. `SELECT` names the aggregation - `count(*)`, `average(duration)`, `percentile(duration, 95)`; `FROM` names the data type; `WHERE` filters on attributes; `SINCE` and `UNTIL` bound the window. The two clauses that shape the result are `FACET` and `TIMESERIES`. `FACET` groups: one aggregate per distinct value of an attribute, bounded to a group count you raise with `LIMIT`. `TIMESERIES` buckets: the same aggregation computed independently inside each slice of the window, which turns a number into a line. `COMPARE WITH` re-runs the query shifted back to overlay an earlier period. That is a different kind of language from one built on labelled time series, where a series exists because a label was chosen before the data was written. In NRQL any attribute on the record is a candidate grouping - flexibility bought on read and paid for in scan cost.
code
nrql · 6 linesSELECT percentile(duration, 95)
FROM Transaction
WHERE appName = 'tour-logistics-api'
FACET name
SINCE 6 hours ago
TIMESERIES 5 minutesgo deeper
Know the shape of a NRQL query well enough to read one aloud: what is being counted or averaged, out of which data type, filtered how, and over what time window.
Explain what FACET and TIMESERIES each do to the result set, and why an aggregation that runs at query time can group by an attribute nobody planned for in advance.
Show judgement about query cost: which filters actually narrow the scan, when a dashboard panel is too wide to re-run on every load, and how facet limits hide the outlier you wanted.
Own read-time aggregation as a platform property. Decide what every team must attach to a record, what you refuse to accept, and how query breadth and ingest volume land in the bill.
## The clause model A NRQL query is read left to right and each clause does one job. The order in a typical query is: 1. `SELECT` — the aggregation to compute, such as `count(*)`, `average(duration)`, `percentile(duration, 95)` or `uniqueCount(...)`. Without an aggregation you get raw records back instead. 2. `FROM` — which data type you are reading: `Transaction`, `Log`, `Span`, `Metric` and so on. 3. `WHERE` — a predicate over the attributes on those records. 4. `FACET` — split the result into one aggregate per distinct value of an attribute. 5. `SINCE` / `UNTIL` — the time window, written in relative English such as `SINCE 6 hours ago`. 6. `TIMESERIES` — cut that window into buckets and return one aggregate per bucket, optionally with an explicit width. 7. `LIMIT` — how many rows or facets come back. `COMPARE WITH` deserves a mention of its own: it re-runs the same query shifted back by an offset and returns the earlier result alongside the current one, which is how a chart gets a "same time last week" overlay without a second query. Two clauses carry most of the meaning. **`FACET` groups; it does not filter.** Adding `FACET` to a query does not reduce the records considered, it changes one aggregate into many — one per attribute value — and the platform returns a bounded number of groups unless you raise it with `LIMIT`, which is exactly why a facet chart can silently omit the group you were looking for. **`TIMESERIES` buckets; it does not change the maths.** The same aggregation is computed independently inside each bucket, which is what turns a single number into a line. ## Aggregation happens when you ask The important property is *when* the work is done. NRQL runs the aggregation at query time over the records that matched. The store keeps records with their attributes; it does not keep a pre-built answer. Consider a tour-logistics service that peaks at 5,400 requests per second and whose telemetry budget was just halved. Someone asks which destination city is responsible for the slow shipments. If the destination is an attribute on the request records, the answer is one `FACET` away and nobody had to decide in advance that it mattered. That is the whole benefit of read-time aggregation, and it is worth a lot under a halved budget, because you did not have to pre-commit to carrying every breakdown you might one day want. ## Why this differs in kind from a labelled-time-series language | | NRQL over a record store | a query language over labelled time series | |---|---|---| | Unit of data | one record with arbitrary attributes | one series identified by a name plus labels | | When aggregation happens | at query time, over matching records | largely at write time; the series already exists | | Grouping by a new dimension | works if the attribute was on the record | needs that label to have existed on the series | | Extra dimensions cost | query scan cost, paid per query | storage and memory, paid continuously | | Time bucketing | a clause on the query | a step or range given to the engine | | Natural question | "what happened, sliced how I like" | "how does this series behave over time" | Neither model is better. The record model buys unplanned breakdowns and pays for them in scan cost; the series model buys cheap, fast, repeatable charts and pays for them by making you choose your dimensions before the data is written. The single most useful sentence in an interview is that a dimension you did not plan for is a *query* problem in the first model and a *data collection* problem in the second. ## Where it bites in practice - **A missing time window is still a window.** Every NRQL query is bounded; if you did not write `SINCE`, one was chosen for you, and you should know that rather than be surprised by it. - **Facet limits mislead.** A dashboard showing "the top services" is showing the top *n* by the aggregation you chose, and the interesting outlier is often not in the top *n*. - **Wide queries on dashboards are recurring cost.** A panel that scans a month of records is not paid once; it is paid every time someone opens the dashboard and every refresh after that. - **Aggregating first is not always right.** `count(*)` over an hour hides the shape; the same query with `TIMESERIES` frequently shows that the "steady" rate is two bursts. ## What a strong answer sounds like Name the clauses and what each one does to the *result set*, not just what it is called. Then make the model claim: NRQL aggregates records on read, so the flexibility is in the query and the cost is in the scan; a labelled-time-series language selects pre-existing series, so the flexibility is decided at write time and the cost is in the store. Anyone who can state that contrast cleanly has understood both kinds of system, not just memorised a clause list.
- What changes when you add TIMESERIES to a query that already has FACET?You go from one aggregate per group to one aggregate per group per bucket, so the result is a set of lines rather than a set of rows. The grouping is still bounded, so the chart shows the top groups by the aggregation you chose, and an interesting outlier can be missing from the picture entirely unless you raise the limit or filter to it.
- You need last week's shape beside this week's on one chart. What does NRQL give you?COMPARE WITH. It re-runs the same query against a window shifted back by the offset you name and returns both results together, so the comparison uses identical filters, facets and bucketing by construction. That is materially safer than writing a second query by hand, where one differing predicate quietly invalidates the comparison.
- Why can NRQL group by an attribute nobody planned to chart when a labelled-time-series store cannot?Because the aggregation runs on read. NRQL scans records that already carry the attribute and groups them at query time, so any attribute present is a candidate dimension. A labelled-time-series store fixes a series' identity when the data is written, so a dimension that was not a label then does not exist now, and no query can recover it.
NRQL treats telemetry as a table you group on demand; a labelled-time-series language treats it as a shelf of pre-cut ribbons you can only pick up whole.
saying these in an interview costs you the question
- Assumes NRQL is ordinary SQL against a relational schema
- Treats FACET as a filter rather than a grouping
- Thinks TIMESERIES changes the aggregation rather than bucketing it
- Believes results are limited to what was pre-aggregated at write time
- Leaves the time window implicit and cannot say what applies