skip to content

What are the 12 bytes of a MongoDB ObjectId, and what can you infer from one?

level: middleimportance: must knowfreq 75%

answer

  1. twelve bytes, twenty-four hex characters
  2. one part of it is a clock
  3. seconds, not milliseconds
  4. five random bytes per process, three counting
  5. getTimestamp() decodes the first four bytes

basics

~20 s

An ObjectId is 12 bytes: a 4-byte Unix timestamp in seconds, a 5-byte value random per process, and a 3-byte incrementing counter. From one you can read its creation time to the second and infer rough insertion order.

solid answer

~50 s

An ObjectId is a 12-byte BSON value, printed as 24 hexadecimal characters. The first 4 bytes are a big-endian Unix timestamp in **seconds**, the next 5 bytes are a random value generated once per process, and the last 3 bytes are a counter that starts at a random value and increments for each ObjectId that process creates. `getTimestamp()` decodes the first four bytes, so you get creation time at one-second resolution for free — which is why sorting by `_id` approximates insertion order. It is only an approximation: the timestamp is generated **client-side** by the driver when `_id` is absent, so clock skew between application servers, and the random block deciding ties within the same second, both break strict chronological order. ObjectIds are also predictable, so they are not a substitute for an unguessable token.

code

javascript · 3 lines
javascript
const id = new ObjectId()
id.toString()      // 24 hex characters = 12 bytes
id.getTimestamp()  // Date decoded from the first 4 bytes

go deeper

for a junior

Recall that _id defaults to an ObjectId, that it is 12 bytes shown as 24 hex characters, and that its first bytes hold a creation timestamp you can read back.

for a middle

Be ready to name all three fields — timestamp seconds, per-process random value, incrementing counter — and explain why that layout allows uncoordinated generation across many application processes.

for a senior

Show where the approximation breaks in production: one-second resolution, client-side generation under clock skew, and ordering ties across processes. Say what you would store instead when exact ordering matters.

for a principal

Own the identifier strategy: when ObjectId's time-prefixed monotonic shape helps (time-range scans, cheap created-at) versus when you need externally meaningful or unguessable keys, and what that choice costs in index size and cross-system identity.

## What an ObjectId is `ObjectId` is a distinct BSON type — not a string and not a UUID. It occupies **12 bytes** on disk and in indexes, and is displayed as 24 hexadecimal characters, e.g. `ObjectId("6512a3f19c4b2d1a7e0f8b3c")`. It is the value MongoDB uses by default when you insert a document without an `_id` field, which is why nearly every MongoDB collection you meet has ObjectId primary keys. Its design goal is to be generatable **without coordination**: many application processes must be able to mint identifiers simultaneously, with no round trip to the database and no shared sequence, and still practically never collide. ## The three fields In current drivers the 12 bytes are laid out as: - **Bytes 0–3 — timestamp.** A 4-byte big-endian count of **seconds** since the Unix epoch. Resolution is one second, not one millisecond. - **Bytes 4–8 — a 5-byte random value.** Generated once per process and reused for every ObjectId that process creates. It is what keeps two processes minting ids in the same second apart. Older drivers filled these five bytes with a machine identifier plus a process id instead; current drivers use a single random value, so an ObjectId no longer leaks anything about the host. - **Bytes 9–11 — a 3-byte counter.** Initialised to a random value at process start and incremented for each ObjectId that process generates, so ids created within the same second by the same process are still distinct and ordered. ## Where the value is generated When you insert a document with no `_id`, the **driver** normally creates the ObjectId on the client before sending the write; if the field is still missing when the write reaches the server, the server supplies one. The practical consequence is that the embedded timestamp reflects the **application server's clock**, not the database's. On a fleet with drifting clocks the timestamps drift too. ## What you can infer - **Creation time to the second**, via `getTimestamp()` in mongosh (drivers expose an equivalent). This is genuinely useful: a collection with ObjectId `_id` values carries a free created-at approximation. - **Rough insertion order**, because the `_id` index is sorted and the timestamp is the high-order prefix. - Nothing about the host machine with modern drivers — beyond the fact that two ids sharing the same 5-byte block came from the same process run. ## What you must not rely on - **Exact ordering.** Within one second, order is decided by the per-process counter only if both ids came from the same process; across processes the random block breaks the tie arbitrarily. Two events a few hundred milliseconds apart on different app servers can sort in the wrong order. - **Secrecy.** An ObjectId is largely predictable: the timestamp advances with the clock and the counter increments by one. Using one as a password-reset token, an unguessable share link, or any capability is a security defect; use a cryptographically random value for that. - **A guarantee of uniqueness.** Uniqueness is probabilistic (random per-process block plus counter), and in practice sufficient, but it is not a mathematical guarantee the way a coordinated sequence is. ## Practical consequences Storing the 24-character hex form as a **string** instead of an `objectId` doubles the storage and index size for the same value and breaks `getTimestamp()` and type-correct comparisons — always store it as the real type. Conversely, if your identifiers come from an external system, you can put them in `_id` directly rather than carrying both an ObjectId and a natural key. Because the timestamp is the high-order prefix, ObjectIds are **monotonically increasing over time**, which makes `_id` range scans over a time window nearly free but also means all new inserts land at one end of the index.

  • Can you rely on sorting by _id to reproduce exact insertion order?
    No. The timestamp prefix has one-second resolution, the counter only orders ids from the same process, and the value is generated on the client, so clock skew between application servers can invert order. It is a good approximation for reporting and pagination, but not an audit-grade sequence — store an explicit event time or sequence if you need one.
  • Is an ObjectId safe to expose as a public identifier in a URL?
    Exposing it is fine for identification, but never treat it as a secret. The timestamp advances predictably and the counter increments by one, so nearby ids are guessable and the value leaks the record's creation time. Anything that must be unguessable — reset tokens, share links — needs a cryptographically random value instead.
  • Is an ObjectId generated by the server or the client?
    The driver generates it client-side when the document has no _id, and the server fills one in only if the field is still absent on arrival. That is why the embedded timestamp reflects the application server's clock, and why the value is known to the application before the write is acknowledged.

saying these in an interview costs you the question

  • Says an ObjectId embeds the machine's MAC address
  • Claims ObjectIds sort in exact millisecond-accurate global order
  • Treats an ObjectId as an unguessable secret or security token
  • Calls an ObjectId a UUID or a random 12-byte value
  • Assumes the server always generates the _id value

context