skip to content

How do Python's uuid.uuid4, uuid.uuid1 and uuid.uuid5 differ in where their bits come from?

level: juniorimportance: must knowfreq 62%

answer

  1. Three different sources of bits
  2. Random, clock, or digest
  3. One of them repeats on purpose
  4. uuid1 carries a host node value
  5. uuid5 = namespace UUID plus name

basics

~20 s

uuid.uuid4 fills 122 bits from the operating system's random source. uuid.uuid1 encodes a timestamp plus the host's 48-bit node address. uuid.uuid5 hashes a namespace UUID together with a name, so the same name always produces the same UUID.

solid answer

~50 s

All three return a 128-bit `uuid.UUID`; the version nibble records how the bits were made. `uuid.uuid4()` takes sixteen bytes from the OS random source and stamps version 4, leaving 122 random bits — no state, no coordination, nothing derivable. `uuid.uuid1()` builds a value from a 60-bit clock reading, a 14-bit clock sequence and a 48-bit node id from `uuid.getnode()`, so the value carries when and roughly where it was made and can be read back through `UUID.time` and `UUID.node`. `uuid.uuid5(namespace, name)` is deterministic: it digests the namespace UUID's bytes plus the name with SHA-1 and keeps the first sixteen bytes, so the same pair always yields the same UUID on any machine. Reach for `uuid4` for an opaque surrogate id, `uuid5` when you want to recompute an id from stable inputs, and `uuid1` only when embedded time and host actually help you.

code

python · 6 lines
python
import uuid

random_id = uuid.uuid4()
name_id = uuid.uuid5(uuid.NAMESPACE_URL, "https://example.com/a")
print(random_id.version, name_id.version, uuid.uuid1().version)
print(name_id == uuid.uuid5(uuid.NAMESPACE_URL, "https://example.com/a"))

go deeper

for a junior

Be ready to say in one breath that uuid4 is random, uuid1 is time-plus-host, and uuid5 is a digest of a namespace and a name, and that uuid5 is the only one that repeats. Know that these functions return UUID objects, not strings.

for a middle

Explain the mechanics: 128 bits with a stamped version nibble, uuid4 pulling from the OS random source, uuid1 assembling a clock reading with uuid.getnode(), uuid5 hashing namespace bytes plus the UTF-8 name and truncating the digest.

for a senior

Show the judgement about which one belongs in a given system: what a uuid1 leaks to an outside caller, why a derived uuid5 is not a secret, and when recomputability is worth more than opacity.

for a principal

Own the convention for a whole codebase — one default generator, one pinned namespace constant for derived ids, and a written rule about what identifiers may reveal to parties outside your trust boundary.

A UUID is 128 bits — sixteen bytes. Six of them are spoken for: four bits hold the **version** nibble and two or three hold the **variant** that says which specification the layout follows. The version is the interesting part for an interview, because it names the strategy that produced the remaining bits, and Python's `uuid` module exposes one function per strategy. ### uuid.uuid4 — random bits `uuid.uuid4()` reads sixteen bytes from the operating system's random source, overwrites the version and variant fields, and hands back the result. That leaves 122 free bits. There is no shared state, no clock, no coordination between machines, and nothing in the value that can be traced back to where it was made. Two processes on two continents can each mint millions of them and the chance of a collision stays negligible. The price of that is that the value is *only* an identifier: you cannot recompute it, you cannot derive it from your data, and it tells you nothing about ordering. ### uuid.uuid1 — clock plus node `uuid.uuid1()` composes three things: a 60-bit timestamp counting 100-nanosecond intervals since 1582-10-15, a 14-bit clock sequence that is re-randomised when the clock appears to go backwards, and a 48-bit node identifier obtained from `uuid.getnode()`. `getnode()` returns the machine's hardware address when it can read one, and otherwise a random 48-bit value with the multicast bit set, stable for the life of the process. Uniqueness here comes from coordination — distinct hosts, a clock that mostly moves forward — rather than from entropy. The inputs are readable back out of the value through `UUID.time`, `UUID.node`, `UUID.clock_seq` and `UUID.fields`, which is simultaneously the feature and the drawback: a version 1 UUID handed to an outside party tells them when it was created and, if the node came from real hardware, gives them a stable per-machine fingerprint. ### uuid.uuid5 — namespace and name `uuid.uuid5(namespace, name)` is the deterministic one. It concatenates the namespace UUID's sixteen bytes with the name (a `str` is encoded as UTF-8; `bytes` is accepted directly on 3.14), takes a SHA-1 digest, keeps the leading sixteen bytes and stamps version 5 and the variant. The same namespace and name give the same UUID on any machine, in any process, forever. `uuid.uuid3` is the identical construction over an MD5 digest with version 3; prefer `uuid5` in new code. The module ships four namespace constants — `uuid.NAMESPACE_DNS`, `uuid.NAMESPACE_URL`, `uuid.NAMESPACE_OID` and `uuid.NAMESPACE_X500` — and nothing stops you from pinning a constant of your own instead; the namespace simply partitions names so that the same string under two namespaces yields two different ids. Determinism is what you are buying: an identifier you can recompute from business data instead of storing a lookup table. It is also the limitation. The value is *derived*, not secret — anyone who can guess the inputs can reproduce it — so it is not a substitute for a token that has to be unguessable. ### Choosing between them An opaque surrogate identifier with no story attached: `uuid4`. An identifier you want to regenerate from stable inputs so a repeated operation lands on the same row: `uuid5`. An identifier whose creation time you genuinely want embedded: `uuid1` will do it, though Python 3.14 also added `uuid.uuid6`, `uuid.uuid7` and `uuid.uuid8` from the newer specification, plus the `uuid.NIL` and `uuid.MAX` constants. ### Traps worth naming out loud "Guaranteed unique" is wrong for all of them. `uuid4` is *collision-improbable*; `uuid1` is unique only while its assumptions about distinct nodes and a sane clock hold. `UUID.version` reads the stamped nibble rather than inferring anything, so a `UUID` you build yourself from arbitrary bits with `uuid.UUID(int=...)` can report a version nobody generated. And the generator functions return `UUID` objects, not strings — the 36-character hyphenated form only appears when something calls `str()` on one. ```python import uuid print(uuid.uuid4().version, uuid.uuid1().version) print(uuid.uuid5(uuid.NAMESPACE_DNS, "example.com")) print(uuid.uuid5(uuid.NAMESPACE_DNS, "example.com")) ``` The two `uuid5` lines print the same value every time you run the file; two `uuid4` calls never will.

  • What does uuid.uuid3 do that uuid.uuid5 does not?
    Nothing structurally different — it is the same namespace-plus-name construction, but the digest is MD5 instead of SHA-1 and the stamped version is 3. Both truncate the digest to sixteen bytes and overwrite the version and variant bits, so neither is reversible. Use `uuid.uuid5` for new code and keep `uuid.uuid3` for interoperating with ids somebody already generated that way.
  • Where does uuid.uuid1 get its node value, and what happens when no hardware address is readable?
    From `uuid.getnode()`, which tries the platform's interfaces for a 48-bit hardware address. When it cannot read one it returns a random 48-bit value with the multicast bit set, and caches it, so every `uuid1` from that process shares a node id that is stable for the process but meaningless across restarts. That fallback is why you cannot treat the node field as a reliable machine identity.
  • Does picking uuid.NAMESPACE_DNS instead of uuid.NAMESPACE_URL change how strong a uuid5 value is?
    No. The namespace is just sixteen bytes prefixed to the name before hashing, so it partitions the name space: the same string under two namespaces produces two unrelated UUIDs. It adds no entropy and no secrecy. Choosing a namespace is a modelling decision about which family of names you are in, and pinning your own constant UUID as a private namespace is perfectly normal.

uuid4 is a lottery ticket, uuid1 is a timestamped postmark with the sender's address on it, and uuid5 is a filing rule: give it the same drawer and the same label and it always points at the same slot.

saying these in an interview costs you the question

  • Claims uuid4 values are guaranteed unique rather than collision-improbable
  • Thinks uuid5 is random and cannot be reproduced
  • Believes a uuid1 value carries no host or time information
  • Calls uuid5 a reversible encoding or encryption of the name
  • Says uuid4 seeds its randomness from the current time
  • Assumes the generator functions return strings rather than UUID objects

context