skip to content

In Segment, how does an anonymousId get linked to a userId, and what breaks it?

level: seniorimportance: should knowfreq 52%

answer

  1. a random id exists before you know who they are
  2. one call carries both identifiers at once
  3. logout must clear the cached identifier
  4. server calls have no browser storage to read

basics

~20 s

Segment's client library assigns an anonymousId and stores it locally; the first identify call sends that anonymousId together with the userId, letting downstream tools stitch the two. Skipping analytics.reset on logout or omitting the anonymousId on server calls breaks the link.

solid answer

~50 s

Segment's browser library generates an `anonymousId` on first visit and persists it, attaching it to every call. When you finally call `identify(userId, traits)`, the message carries **both** ids, which is the signal destinations use to merge the anonymous history into the known profile; the library then caches the `userId` and keeps sending both. Three things commonly break it. First, not calling `analytics.reset()` on logout: the cached `userId` survives, so the next person on that browser has their events attributed to the previous user. Second, server-side calls that send only a `userId` and never the browser's `anonymousId` — pre-login activity stays on a separate anonymous profile. Third, identifying too late, so the whole acquisition funnel is anonymous. Note also that Segment forwards both ids; how they are merged is each destination's own identity model, and several vendors have changed theirs.

code

text · 5 lines
text
1  track     anonymousId=a1b2                         "Product Viewed"
2  identify  anonymousId=a1b2  userId=usr_77          {email: ada@...}
3  track     anonymousId=a1b2  userId=usr_77          "Order Completed"
4  -- user logs out, analytics.reset() is NOT called --
5  track     anonymousId=a1b2  userId=usr_77          "Product Viewed"   <- next visitor

go deeper

for a junior

Know that Segment gives an unknown visitor a generated anonymousId, and that identify is what attaches a real userId to it.

for a middle

Explain that the link is made by one message carrying both ids, that the library then caches the userId, and that reset clears it — plus why server libraries have no such storage.

for a senior

Show that you have debugged this: spot the missing reset from warehouse symptoms, plumb the anonymousId into server calls, and know that merge semantics belong to each destination rather than to Segment.

for a principal

Own identity as policy: which identifier is authoritative, how profiles are prevented from over-merging on shared values like a support email, and what the deletion story looks like across every downstream vendor.

## Where the anonymousId comes from The first time a visitor loads a page carrying Segment's browser library, the library mints a random `anonymousId` and persists it client-side (in the browser's cookie/local storage under an `ajs_anonymous_id` key). Every `page` and `track` call from then on carries it. That is what makes pre-login behaviour analysable at all: you cannot know who the visitor is, but you can tie their pageviews, searches and add-to-carts to one another. On mobile, the equivalent identifier is generated and stored by the mobile library; on server libraries there is no automatic value at all, which turns out to matter enormously (see below). ## The link is made by sending both ids together When the user logs in or signs up, you call `identify(userId, traits)`. The library attaches the existing `anonymousId` alongside the new `userId`, and that single message — *"these two identifiers are the same person"* — is the stitching signal. Afterwards the library caches the `userId` too, so subsequent calls continue to carry both. An important nuance for an interview: **Segment forwards the pair; it does not impose a merge.** Each destination applies its own identity model to that message. Some merge the anonymous profile automatically on seeing both ids. Some require an explicit `alias` call. Several vendors have reworked their identity handling over the years, so the correct answer for any given destination is "check that destination's current documentation", not a blanket rule. Segment's own profile layer (the product has been renamed over time — Personas, then Unify/Engage) does a further, richer resolution across identifiers such as email and phone, with rules about how many of each identifier a profile may accumulate, precisely so a shared value like a support email does not collapse thousands of people into one profile. ## Failure one: no reset on logout `analytics.reset()` clears the cached `userId` and traits and issues a fresh `anonymousId`. If you do not call it when a user logs out, the browser keeps the previous `userId` and keeps sending it. On a shared or kiosk machine, the next visitor's browsing is attributed to the person who logged out. This is not a cosmetic error: it corrupts per-user funnels, poisons any audience built from behaviour, and in a B2B account it can attribute one customer's activity to another. The symptom in the warehouse is a user whose events continue at implausible volume and from implausible contexts after their session should have ended, often with a device or IP that changes underneath a constant `userId`. ## Failure two: server-side calls with no anonymousId Server libraries do not have the browser's storage. A server-side `track` that supplies only a `userId` is fine for a logged-in user, but a server-side call made **before** login — an event from a checkout service, say — either has no id at all (the API requires a `userId` or an `anonymousId`) or invents a new one, creating a fresh anonymous profile that will never join the real one. The fix is to plumb the browser's `anonymousId` through to the server (read it from the cookie, pass it in your API call) and send it explicitly on the Segment call. This is the most common structural mistake in a hybrid client/server instrumentation, and it produces the maddening symptom that client-side and server-side events for the same user never appear in the same funnel. ## Failure three: identifying too late If the only `identify` fires after an email confirmation step at the end of onboarding, everything before it lives under the anonymous identifier. Destinations that stitch on `anonymousId` will recover it; destinations that do not will simply never attribute the funnel's early steps to the user. The guidance is to identify as soon as you have a stable id — at login, at signup submission, on session restore from a token — and to re-identify on every page load of an authenticated session so that a page loaded in a fresh tab is not anonymous. ## Cross-device is a different problem One browser's `anonymousId` says nothing about the same person's phone. Cross-device stitching is only deterministic once the same `userId` appears on both devices, which means it depends on the user logging in on each. Anything beyond that is probabilistic and belongs to a profile layer, not to the SDK. ## Operational checks A short list worth running on any Segment implementation: does logout call `reset()`; do server events carry the browser's `anonymousId`; does `identify` fire on session restore, not just at login; are there `userId` values with wildly implausible event counts; and does the ratio of anonymous to identified events for logged-in journeys match expectations. Each of those maps directly to one of the failure modes above.

  • How do you get server-side events onto the same profile as browser events for a not-yet-logged-in visitor?
    Read the anonymousId the browser library stored, pass it to your backend with the request, and set it explicitly on the server-side Segment call. A server call that mints its own identifier creates a second anonymous profile that will never merge. Once the user authenticates, send the userId as well so the pair is linked.
  • Why is cross-device stitching not solved by the anonymousId?
    The anonymousId is per browser and per storage context, so a phone and a laptop start with different values and nothing connects them. Deterministic cross-device linkage requires the same userId to appear on both, which means the user must log in on each. Anything more is probabilistic matching handled by a profile layer, not by the SDK.
  • What symptom in the warehouse suggests a missing reset on logout?
    A userId whose event stream continues at implausible volume, spanning contexts that change underneath a constant identifier — different IPs, user agents or locales on the same user within one day, often on shared or kiosk devices. Compare the distinct-context count per userId to spot it quickly.

The anonymousId is a numbered cloakroom ticket; identify is the moment you write your name on it, and reset is handing the ticket back so the next guest gets a clean one.

saying these in an interview costs you the question

  • Believes Segment merges profiles for the destination automatically
  • Never calls reset on logout
  • Sends only userId on server-side calls
  • Thinks anonymousId links a user across devices
  • Fires identify only once, deep in the signup funnel

context