skip to content

In Mixpanel, how does calling identify() link a user's anonymous events to their account?

level: middleimportance: must knowfreq 52%

answer

  1. two identifiers describe the same person
  2. the browser has one before login, you have another
  3. the login call joins the pair
  4. clusters, not renames
  5. logout needs its own call

basics

~10 s

Mixpanel tags pre-login activity with a device-generated distinct_id. Calling identify with your application's user id links that device id to the user id, so anonymous and logged-in events resolve to one user in reports.

solid answer

~50 s

Every Mixpanel client generates an anonymous `distinct_id` on first load and stamps it on events. When the person authenticates you call `mixpanel.identify('u_42')` with your own stable user id. Under Simplified ID Merge — the model newer projects use — that call links the device id and the user id into one identity cluster, so the pre-signup pageviews and the post-login events are counted as the same user and appear in a single funnel. Older projects run Original ID Merge, where `alias` performed that link and had to be called exactly once, in the anonymous-to-known direction, at signup. Two rules matter either way: identify with an id that never changes and is never shared (a database id, not an email), and call `mixpanel.reset()` on logout so the next person on a shared browser starts a fresh anonymous id instead of being merged into the previous account.

code

javascript · 9 lines
javascript
// Before login: events carry the library's anonymous device id
mixpanel.track('Pricing Page Viewed');

// After authentication: join the device id to the account id
mixpanel.identify(user.id);            // opaque, immutable, never shared
mixpanel.people.set({ $email: user.email });

// On logout: new anonymous id for whoever uses this browser next
mixpanel.reset();

go deeper

for a junior

Know that Mixpanel gives anonymous visitors a generated distinct_id and that identify is the call made after login with your own user id.

for a middle

Explain the merge mechanically: device id and user id joined into one identity cluster, what reset does on logout, and why the identifier must be stable and unshared.

for a senior

Diagnose a broken funnel from the raw event stream — missing identify on one path, server events with a different distinct_id, a shared id swallowing many users — and know that merges are effectively one-way.

for a principal

Own the identity contract across web, mobile and backend: one canonical id source, an agreed lifecycle for anonymous-to-known transitions, and a documented position on what happens to analytics when accounts merge or are deleted.

## What distinct_id is Every event Mixpanel ingests carries a `distinct_id` — the identifier of the actor the event belongs to. All user-scoped analysis, from unique counts to funnels to retention, is grouped by it. The whole identity problem in Mixpanel reduces to one question: does the platform know that the anonymous visitor of Tuesday and the logged-in customer of Thursday are the same `distinct_id`? When the browser library loads for the first time it generates an anonymous identifier and persists it in local storage or a cookie. Every event tracked before login carries it. Your application, meanwhile, has its own stable user id. Identity resolution is the act of joining the two. ## Simplified ID Merge In the ID-merge model Mixpanel uses for newer projects, the client tracks two properties: a device identifier (`$device_id`) and, once known, a user identifier (`$user_id`). Calling: ```javascript mixpanel.identify('u_42'); ``` sends the identification, and Mixpanel links the device id and the user id into a single identity cluster. Events from either identifier — including the ones already ingested under the anonymous device id — resolve to one canonical user. That is what makes a `Landing Page Viewed → Signed Up → Activated` funnel work when the first step happened before an account existed. A device that later sees a second user identify on it joins that second cluster for subsequent events, which is why logout handling matters. ## Original ID Merge and alias Projects created under the older model behave differently. There, `identify` simply switched the id used for subsequent events, and `mixpanel.alias('u_42')` was the call that told Mixpanel the current anonymous id and the new user id were the same person. `alias` carried real constraints: call it once per user, at signup, and only from the anonymous id to the known id — calling it repeatedly or in the wrong direction created identity graphs that were hard to unpick. If you inherit an old project, find out which merge model it runs before writing a line of tracking code; the correct call sequence is different and the documentation for one model is actively misleading for the other. ## Logout and shared devices `mixpanel.reset()` clears the stored identity and super properties and generates a fresh anonymous id. Without it, a second person using the same browser after a logout tracks events that Mixpanel attributes to — or merges into — the first person's cluster. On kiosk, support-desk and family-shared devices this produces users with impossible behaviour and inflated per-user counts. Call `reset` in the same code path that clears your own session. ## Server-side and multi-surface tracking Server-side events have no browser storage to draw on, so the calling code must set `distinct_id` explicitly on each event. If your backend stamps only the internal user id and the web client is still on its anonymous device id, the two streams stitch together only if the client has already identified. A common failure mode is backend events for a signup flow that never join the pre-signup web events because the web `identify` fires after the server event, or not at all on that path. Mobile SDKs have the same shape as the web one — a device-scoped anonymous id until `identify` is called. ## Choosing the identifier The id you pass to `identify` should be stable for the lifetime of the account, opaque, and never reused. Emails fail the first test — people change them — and usernames fail the third. Sequential integers are acceptable but leak account counts to anyone who reads a URL or an export. Above all, never identify two different humans with the same value: merges in Mixpanel are effectively one-way, and unpicking a cluster that swallowed a shared 'admin' id usually means deleting and re-importing the affected event range. ## What to check when identity looks wrong Symptoms are recognisable. A funnel where step one has vastly more users than the population that could have reached it suggests anonymous events are not merging. A small number of users with implausible event counts suggests a shared id or a missing `reset`. Users appearing twice — once anonymous, once identified — with adjacent timestamps suggests `identify` is firing on the wrong page or after the events you care about. In each case, look at the raw event stream for a single session and read the identifiers directly rather than reasoning from the aggregate report.

  • Why is an email address a poor argument to pass to identify?
    It changes. When a user updates their email you would start identifying them under a new id, splitting one person into two users and breaking every retention and funnel report that spans the change. Use an opaque, immutable primary key from your own database and treat the email as a profile property instead.
  • What goes wrong on a shared kiosk browser if you never call reset on logout?
    The stored identity survives the logout, so the next person's events are tracked under the previous account and can be merged into that identity cluster. You get users with impossible behaviour, inflated per-user counts, and cross-contaminated funnels. Calling `mixpanel.reset()` in the same path that clears your session prevents it.
  • How do you stitch server-side events to a web session that has not identified yet?
    You cannot rely on the SDK — server events must carry an explicit `distinct_id`. Either pass the client's current device identifier through to the backend so the server stamps the same value, or make sure the client identifies before the server-side event fires, so both streams end up in the same identity cluster.

The browser knows the visitor by the coat they arrived in; your application knows them by their membership number. Identify is the cloakroom ticket that ties the two together for everything they did that evening.

saying these in an interview costs you the question

  • Passes the user's email address to identify
  • Thinks identify rewrites already-ingested events in place
  • Never calls reset on logout
  • Uses alias on every login rather than once at signup
  • Assumes server-side events inherit the browser's distinct_id

context