skip to content

Segment

A customer-data platform: instrument identify and track once, then fan the same events out to dozens of destinations, with identity resolution and a tracking plan holding the schema steady. Interviewers ask because it is the standard answer to "every tool has its own SDK and none of them agree".

on this pageshow

explore

questions

5

What is the difference between Segment's identify and track calls?

level: juniorimportance: must knowfreq 80%

answer

  1. one call is about the person
  2. the other is about the action
  3. traits persist, properties describe one occurrence
  4. userId plus traits versus event name plus properties

basics

~10 s

Segment's identify call records who a user is: a userId plus traits such as email or plan. track records what a user did: an event name plus properties describing that one action.

solid answer

~50 s

`identify(userId, traits)` tells Segment who the current user is and attaches durable attributes to that profile — email, name, plan, company. `track(event, properties)` records a single action, such as `Order Completed`, with properties describing that action: `order_id`, `revenue`, `currency`. The rule of thumb is that traits describe the person and persist, while properties describe one event and do not. The rest of the spec fills the gaps: `page` and `screen` for views, `group` for the account a user belongs to, and `alias` for destinations that need an explicit profile merge. Every message, whichever call it is, carries a `messageId`, a `userId` and/or an `anonymousId`, timestamps and a `context` block. In the warehouse they land separately: an `identifies` table plus a `users` table of latest traits, and a `tracks` table plus one table per event name.

code

javascript · 11 lines
javascript
analytics.identify('usr_9f3a', {
  email: '[email protected]',
  plan: 'pro',
  createdAt: '2026-01-14T09:12:00Z'
});

analytics.track('Order Completed', {
  order_id: 'ord_5c8d',
  revenue: 49.0,
  currency: 'USD'
});

go deeper

for a junior

Be ready to state the split in one sentence and give an example of each: identify carries a userId and traits, track carries an event name and properties.

for a middle

Explain why the distinction matters downstream — traits upsert onto a profile while properties stay on the event row — and describe how each call lands in the warehouse tables.

for a senior

Show judgment about instrumentation hygiene: a closed event vocabulary, no personal data in names or keys, identify only when traits actually change, and a naming convention the whole team follows.

for a principal

Own the schema as a shared asset: who approves new events, how naming conventions are enforced across teams and platforms, and how volume-priced calls influence what you agree to instrument at all.

## What the Segment spec is Segment is a customer-data platform: you instrument your application once against Segment's small API, and Segment fans each message out to whatever analytics, marketing and warehouse destinations you have configured. That only works because the API is deliberately tiny and semantic. The whole spec is a handful of call types — `identify`, `track`, `page`, `screen`, `group`, `alias` — and the first thing an interviewer checks is whether you know which of them describes a **person** and which describes an **action**. ## identify — who the user is `identify` takes a `userId` and a `traits` object. The `userId` should be your own stable internal identifier for the user (a database id or an opaque uuid), not an email address, because emails change and are personal data you may later need to redact. `traits` are attributes of the person that persist between events: email, first and last name, subscription plan, signup date, company. Semantically, `identify` says "from now on, this session belongs to this user, and here is the current state of their profile". Destinations treat traits as an upsert onto a profile record rather than as a new row of behaviour. That is why sending a trait you did not intend to change is not free — you may overwrite a good value with a stale one. ## track — what the user did `track` takes an event **name** and a `properties` object. The name is the action: `Order Completed`, `Signup Started`, `Report Exported`. Segment's own conventions favour an Object-Action, past-tense name in title case, and its e-commerce spec fixes names like `Product Added` and `Order Completed` so that destinations expecting an e-commerce shape can recognise them. `properties` describe **that occurrence**: which order, how much revenue, which currency, which plan was chosen. Properties do not persist onto a profile; they belong to the event row. The classic beginner error is to encode a variable in the event name — `Order Completed 49 USD` — which produces an unbounded set of event names, breaks every funnel and cannot be aggregated. The value goes in a property; the name stays a small, closed vocabulary. ## The other calls `page` (web) and `screen` (mobile) record a view, with properties such as path, referrer and title; the browser library typically fires `page` automatically. `group` associates the current user with an account or organisation and carries **account-level** traits — useful in B2B, where plan and seat count belong to the company, not to each individual. `alias` exists for destinations whose identity model requires an explicit instruction to merge an anonymous profile into a known one; whether you need it depends entirely on the destination, and several vendors have reworked their identity handling, so check the destination's current documentation rather than adding `alias` reflexively. ## What every message carries Regardless of call type, the payload includes a `messageId` (Segment uses it to collapse duplicate deliveries of the same message within a bounded window), a `userId` and/or `anonymousId`, a `type`, timestamp fields, and a `context` object holding the library, IP, user agent, page information and campaign/UTM parameters. Most of that is populated for you by the client library; on server-side calls through the HTTP tracking API you supply it yourself. ## How this lands in the warehouse Segment's warehouse destination materialises the same distinction. `identify` calls accumulate in an `identifies` table, and Segment maintains a `users` table holding the most recent traits per user. `track` calls accumulate in a `tracks` table containing the common columns for every event, alongside one table per event name holding that event's properties as columns. `page` and `screen` calls get their own tables. So a data engineer reading the warehouse can answer "what is this user's plan" from `users` and "what did they do" from the per-event tables — the shape of your queries follows directly from which call the instrumentation used. ## Where teams go wrong Firing `identify` on every page load with identical traits inflates volume and, on usage-priced plans, cost, without adding information. Using `track` to set attributes (`track('Plan Changed', {plan: 'pro'})` with no corresponding `identify`) leaves profiles stale in every destination that reads traits. And putting personal data into event names or property keys makes later deletion and consent work far harder than putting it in a trait you can redact in one place.

  • In Segment's warehouse tables, why can timestamp differ from originalTimestamp?
    `originalTimestamp` is the client's clock when the call was made, `sentAt` the client clock when the batch left the device, and `receivedAt` Segment's own clock on arrival. `timestamp` is the skew-corrected value derived from those, so it survives a badly set device clock. Use `timestamp` for behavioural analysis and `receivedAt` when you need arrival order for an incremental warehouse load.
  • When would you use group rather than putting the company on identify traits?
    In B2B, where plan, seat count and industry belong to the account rather than to each person. `group` attaches those traits once to the account and associates the user with it, so account attributes are not duplicated across every member's profile and can be updated in one call. Destinations that model accounts separately consume the group call directly.
  • Why should the userId be an internal id rather than an email address?
    Emails change, are reused, and are personal data. A stable opaque internal id keeps profiles from splitting when someone changes their address, keeps you from spraying an identifier across dozens of vendors, and makes deletion requests tractable — you redact the email trait in one place instead of rekeying every downstream profile.

identify updates the entry in a contacts app; track appends a line to that contact's call log.

saying these in an interview costs you the question

  • Says track sets user attributes on the profile
  • Encodes values such as amounts into the event name
  • Fires identify on every page load with unchanged traits
  • Uses email as the userId
  • Thinks properties persist onto the user profile like traits

context

open as a page

In Segment, how do device-mode and cloud-mode destinations differ?

level: middleimportance: must knowfreq 65%

basics

~20 s

A cloud-mode Segment destination receives events from Segment's servers, so Segment can filter, transform, retry and replay them. A device-mode destination gets the vendor's own SDK bundled into the client and receives events directly from the browser or app.

open as a page

In Segment, how does an anonymousId get linked to a userId, and what breaks it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Segment's client library assigns an anonymousId and stores it locally; the first identify call sends that anonymousId together with the userId, letting downstream tools stitch the two. Skipping analytics.reset on logout or omitting the anonymousId on server calls breaks the link.

open as a page

When would you route product events through Segment rather than instrument each vendor SDK directly?

level: principalimportance: should knowfreq 48%

basics

~20 s

Route through Segment when the set of downstream tools changes often, when several teams need one agreed schema, and when you want history replayed into tools added later. Instrument directly when the destination set is small, stable and client-dependent.

open as a page

In Segment Protocols, what does a tracking plan do to a non-conforming event?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

A Segment Protocols tracking plan declares the allowed events and their property types; by default a non-conforming event is still delivered and recorded as a violation. Only the source's schema controls turn that into dropping the event or stripping the offending property.

open as a page