skip to content

In Amplitude's HTTP API v2, what does the insert_id field on an event do?

level: middleimportance: should knowfreq 40%

answer

  1. Uploads are at-least-once, so retries happen
  2. A key that makes a resend harmless
  3. Must be identical across attempts, unique per event
  4. Not a UUID minted at each send
  5. Window is bounded, so backfills still duplicate

basics

~10 s

insert_id is a client-supplied unique key on an Amplitude event. Amplitude drops a later event carrying an insert_id it has already seen within its recent deduplication window, so a retried upload does not double-count.

solid answer

~50 s

Event upload is at-least-once: a mobile client times out, a server-side batch fails halfway, a queue redelivers. `insert_id` is the idempotency key that makes those retries safe — Amplitude discards an incoming event whose `insert_id` it has already ingested within a bounded recent window. Two things matter in practice. First, the value must be **deterministic for the same logical event**: derive it from something stable such as a hash of user id, event type, event time and a per-event sequence number, computed once when the event is created, not regenerated per upload attempt. A fresh UUID per retry defeats the whole mechanism. Second, the dedupe window is bounded, so replaying an old archive or backfilling last quarter's events will land duplicates regardless — treat `insert_id` as protection against retries, not as a general-purpose deduplication service.

code

json · 12 lines
json
{
  "api_key": "REDACTED",
  "events": [
    {
      "user_id": "user-90210",
      "event_type": "Order Completed",
      "time": 1755777600000,
      "insert_id": "user-90210:Order Completed:1755777600000:0007",
      "event_properties": { "cart_total": 42.5 }
    }
  ]
}

go deeper

for a junior

Be able to say that insert_id is a per-event unique key Amplitude uses to ignore a duplicate upload, and that the official SDKs populate it for you.

for a middle

Explain why the value must be computed once at event creation and reused on every retry, and what breaks when it is regenerated per attempt or derived from something too coarse.

for a senior

Show that you know the window is bounded and plan backfills and long-delayed replays accordingly. Be ready to describe how you would detect and repair a duplicated date range after the fact.

for a principal

Own the end-to-end idempotency story: which stage of the pipeline is authoritative, how keys are generated consistently across mobile, web and server producers, and what reconciliation you run to prove counts are right.

## Why the field exists Amplitude receives events over HTTP from clients it does not control: browsers on flaky networks, mobile SDKs that buffer offline and flush later, and server-side pipelines that batch to `/2/httpapi` or the batch endpoint. Every one of those transports is at-least-once. A client sends a batch, the response is lost in transit, the client cannot distinguish "never arrived" from "arrived and the ack was dropped", so it retries. Without a deduplication key you would over-count exactly the events that matter most — the ones sent during the network trouble a user was already experiencing. `insert_id` is the idempotency key that closes that hole. It is an optional string field on each event object in the request body, alongside `user_id`, `device_id`, `event_type`, `time`, `event_properties` and `user_properties`. When Amplitude ingests an event whose `insert_id` it has already seen recently, it drops the newcomer instead of storing a second copy. ## The property that makes it work The key has to be **stable across attempts and unique across events**. That single sentence is the whole interview answer, and it is where implementations fail: - **Wrong:** generate a UUID at send time. Every retry gets a new key, so every retry is a new event. This is the single most common bug, and it looks correct in code review because a UUID is obviously unique. - **Wrong:** reuse something too coarse, like the `device_id` or the session id. Now genuinely distinct events collide and get silently dropped — under-counting, which is far harder to notice than over-counting. - **Right:** compute the value once, at the moment the event object is constructed, and carry it through the buffer, the retry and the batch. A hash of user identifier plus event type plus event time plus a monotonically increasing per-event sequence number is the standard recipe; a UUID minted at event-creation time and stored with the buffered event works equally well. Amplitude's own client SDKs set `insert_id` for you. The field matters most when you are the client: a server-side ingestion job, a warehouse-to-Amplitude push, a custom mobile pipeline, or an event-replay tool. ## What the field does not do The dedupe lookback is bounded — Amplitude compares against recently ingested events, not against the entire history of the project. Three consequences follow: 1. **Replaying an old archive will duplicate.** If you resend last quarter's events after an outage, matching `insert_id` values may fall outside the window and land as new rows. Plan a backfill on the assumption that Amplitude will accept the duplicates, and decide up front whether you would rather delete and reload a date range or accept the skew. 2. **A retry that lingers for days is not protected.** A mobile client that buffers offline for a long trip and flushes on reconnection may fall outside the window on redelivery. 3. **It is Amplitude's protection, not your pipeline's.** If you also export Amplitude data into a warehouse, the deduplication that happened at ingest does not relieve the export loader of being idempotent on its own key. ## Related ingestion pitfalls on the same endpoint Two neighbours come up in the same conversation, because they produce the same symptom — "we sent events and the chart is wrong". **Identifier length.** Amplitude enforces a minimum length on `user_id` and `device_id` values (documented as five characters) and rejects events whose identifiers are shorter, unless the request supplies the `min_id_length` option to relax it. Teams pushing integer primary keys as user ids discover this the first time user `42` is silently rejected while user `10023` succeeds. **Time versus upload time.** The `time` field is the event's own timestamp in epoch milliseconds. It is the field charts bucket on, which is why a client with a badly skewed clock can drop events into next week or last year. Amplitude reconciles client and server upload times to correct for clock skew, but a client that stamps `time` from an untrusted device clock and buffers for hours is still a source of confusing history. ## How to talk about it in an interview Lead with the guarantee: transport is at-least-once, `insert_id` makes ingestion idempotent for retries. Then show you know the failure mode in both directions — regenerating the key per attempt causes over-counting, reusing too coarse a key causes silent under-counting — and close with the boundary: the window is finite, so backfills and long-delayed replays need their own plan. That progression (guarantee, both failure directions, boundary) is what separates a candidate who has read the docs from one who has operated the pipeline.

  • What goes wrong if you build the insert_id from only the user id and the event type?
    Genuinely distinct events collide. Every subsequent `Order Completed` by that user within the dedupe window looks like a duplicate of the first and is dropped, so you silently under-count repeat behaviour. Under-counting is worse than over-counting because nothing errors and no alert fires — the funnel simply looks flat. Include the event time and a per-event sequence number so distinct occurrences stay distinct.
  • Your team replays a month of archived events into Amplitude after an outage. Does insert_id protect you?
    Not reliably. Deduplication compares against recently ingested events within a bounded window, so month-old keys are likely outside it and the replay lands as new rows. Plan the backfill explicitly: scope it to a date range, confirm whether the range is already partially loaded, and prefer resending only the gap rather than the whole month.

saying these in an interview costs you the question

  • Generates a new UUID for insert_id on every upload attempt
  • Believes insert_id deduplicates against the project's entire history
  • Uses device_id or session id as insert_id, colliding real events
  • Assumes Amplitude dedupes automatically without the field
  • Thinks server-side dedupe removes the need for idempotent warehouse loads

context