skip to content

New Relic

Full-stack SaaS observability where everything lands in one telemetry database queried with NRQL: APM, infrastructure, browser and distributed tracing. Interviewers usually probe NRQL and how you slice traces to find the slow service.

on this pageshow

explore

questions

4

In New Relic, why does it matter that events, metrics, logs and spans all land in one database?

level: juniorimportance: must knowfreq 66%

answer

  1. One store behind every New Relic screen
  2. The telemetry database is called NRDB
  3. One query language reaches every data type
  4. Records are a timestamp plus attributes
  5. Joins depend on shared attribute values

basics

~20 s

New Relic writes every signal - events, metrics, logs and spans - as timestamped, attributed records in one store, NRDB, and reads them all with one language, NRQL. Correlating a log with a span becomes a query, not an export.

solid answer

~50 s

New Relic's storage model is a single telemetry database, **NRDB**. APM transaction records, infrastructure samples, log lines and trace spans are all records there: a timestamp, a data type and a bag of attributes. One language, **NRQL**, reads all of them, so `FROM Transaction`, `FROM Log` and `FROM Span` differ only in which records you are reading. For an investigator that removes a seam. You can go from an error rate to the log lines of the failing requests without exporting one product's output into another, and chart a metric beside a log count because a single engine answered both. The price is that correlation is only as good as the attributes you send. Two records meet because they share a value - a trace identifier, a service name, a host. If one emitter spells the service differently, the records sit in the same database and still never join.

code

nrql · 3 lines
nrql
SELECT count(*) FROM Transaction SINCE 30 minutes ago
SELECT count(*) FROM Log SINCE 30 minutes ago
SELECT count(*) FROM Span SINCE 30 minutes ago

go deeper

for a junior

Be ready to say what NRDB and NRQL are: one telemetry database behind every New Relic screen, and one query language that reads events, metrics, logs and spans with the same clauses.

for a middle

Explain the record shape - timestamp, data type, attributes - and show how you move from a chart to the underlying records, naming the attribute that lets two different signals be joined.

for a senior

Demonstrate you have made correlation work in practice: an attribute vocabulary agreed across agents and forwarders, and the debugging you do when logs and traces refuse to line up.

for a principal

Own the tradeoff. Query flexibility and one bill against a query language, data shape and alert estate that are all one vendor's, plus ingest volume as the axis every team's carelessness lands on.

## The storage model in one sentence New Relic's back end is a single telemetry database, usually written **NRDB**. Everything the platform ingests is written there as a record: a timestamp, a data type, and an open-ended set of attributes — key/value pairs — carried on that record. A request record reported by an APM agent, a host sample reported by the infrastructure agent, a forwarded log line and a span from a distributed trace are all the same kind of thing at rest. They differ in which data type they belong to and which attributes they happen to carry, not in which engine stores them. The consequence interviewers are really probing is the query surface. Because there is one store, there is one query language — **NRQL** — and the clause grammar does not change as you move between signals: - `FROM Transaction` reads the completed-request records an APM agent reported. - `FROM Log` reads forwarded log lines. - `FROM Span` reads the spans that make up distributed traces. - `FROM Metric` reads metric data points. - `FROM SystemSample` reads host samples from the infrastructure agent. Filtering, grouping by an attribute value and bucketing over time are written the same way in every one of those. ## What it buys an investigator | During an investigation | With one store and one language | With separate systems per signal | |---|---|---| | Error rate to the failing requests' log lines | one more query, filtered on a shared attribute | export or copy an identifier into a second product | | A metric and a log count on one chart | two queries answered by one engine | two tools, two time pickers, two notions of "now" | | Ad hoc breakdown by an attribute | available if the attribute is on the record | available only where that product indexed it | | Alerting on something you just discovered | the same query becomes the condition | re-express the finding in the other tool's language | The practical effect is that the seam between products disappears. You do not "leave APM and go to logs"; you change which data type you are reading. That is why New Relic interviews so often start here — the model explains why almost every workflow on the platform ends in a query rather than in a product switch. ## What the model demands of the data you send A single store does not correlate anything by itself. Two records meet only because they share an attribute value, so the model pushes work onto the emitters: 1. **A shared identity vocabulary.** Everything that should join must agree on how a service, a host and a request are spelled. A log forwarder that labels the service one way and an agent that labels it another produce records that sit centimetres apart in the same database and never join. 2. **Request identity on every signal that has one.** A trace identifier present on spans and on log records is what turns "this request was slow" into "here are that request's log lines". If the identifier is only on the spans, the join is impossible no matter how good the query is. 3. **Attributes at write time.** Aggregation happens when you ask, but attributes do not appear retroactively. An attribute nobody attached is an attribute nobody can group by, ever, for data already stored. 4. **Restraint.** Because attaching an attribute is easy, teams attach everything. Every attribute is stored on every record; volume-priced ingest turns an unconsidered high-cardinality attribute into a bill rather than into an error message. ## Where the single store costs you - **Aggregation is paid per query.** Nothing is pre-computed by default, so a wide dashboard query over a long window scans a lot of records. Narrow filters and shorter windows are not style preferences, they are the cost control. - **Volume is the pricing axis.** One store means one ingest pipeline and one place where "just add a field" shows up financially. - **Retention is a per-data-type decision.** Records of different types are not necessarily kept for the same length of time, so a correlation that works today may not work against last quarter. - **Portability.** Dashboards, alert conditions and saved investigations end up written in one vendor's language against one vendor's data shape. That is a real strategic cost, and it is the honest answer to "why would you not want this". ## Talking about it well A weak answer describes the model as a feature list. A strong one states the mechanism and then immediately names the condition it depends on: everything is a timestamped, attributed record in NRDB; NRQL reads all of it with the same clauses; correlation works exactly to the extent that your emitters agree on attribute names and values. That last sentence is the one that separates someone who has used the platform from someone who has read about it, because the failure people actually hit is not "the query is hard" but "the two things I want to join do not share a value".

  • What has to be true of your telemetry for that single store to actually pay off?
    A shared vocabulary across emitters. The join between two records is an attribute value, so the service name an agent reports and the one a log forwarder attaches have to match, and a request identifier such as a trace id has to be present on every signal you want to line up. Attributes cannot be added retroactively, so this is a write-time decision.
  • If everything lands in one database, why is a metric still stored differently from an event?
    Because the shapes answer different questions. An event record is one row per occurrence and keeps every attribute of that occurrence, which is what makes ad hoc slicing possible but makes volume track traffic. A metric data point is already aggregated over an interval, so it is cheap and bounded but can only be sliced by the dimensions it was given.
  • You inherit an account where logs never line up with the services that emitted them. Where do you look first?
    At the attributes on the log records themselves, not at the query. Check whether the forwarder attaches a service-identifying attribute at all, whether it spells the service the way the agent does, and whether the trace identifier survives from the application into the log line. Almost every failure of this kind is missing or mismatched attributes at the source.

It is the difference between four filing cabinets with four sets of keys and one archive where every document is filed under the same tags - but only if whoever filed them agreed on the tags.

saying these in an interview costs you the question

  • Thinks each New Relic product keeps its own separate database
  • Believes NRQL only queries metrics, not logs or spans
  • Assumes signals correlate automatically with no shared attributes
  • Treats a metric data point and an event record as identical
  • Says one store means storage and ingest are effectively free
open as a page

In New Relic, how does a NRQL query select, filter, facet and bucket results over time?

level: middleimportance: must knowfreq 74%

basics

~20 s

NRQL aggregates records at query time: SELECT chooses the aggregation, FROM the data type, WHERE filters, FACET splits results by an attribute's values, SINCE and UNTIL set the window, and TIMESERIES cuts that window into buckets.

open as a page

In a New Relic distributed trace, how do you attribute latency when only some services are instrumented?

level: seniorimportance: should knowfreq 51%

basics

~20 s

New Relic attributes time only to the spans it received. Work in an uninstrumented hop shows up as unexplained time inside its caller's span, and the platform cannot say whether that was network, queueing, or the missing service's own work.

open as a page

What is an entity in New Relic, and how does the platform assign telemetry to one?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

An entity in New Relic is anything the platform monitors and can name - a service, host, container, database or browser app - identified by an entity GUID. Incoming telemetry is matched to one by the identifying attributes it carries.

open as a page