How do you store a uuid.UUID as sixteen bytes and rebuild the object later?
answer
- One number, several spellings
- Sixteen against thirty-six
- The constructor takes one keyword only
- bytes, hex, int, fields, bytes_le
- Bad text raises ValueError
basics
~20 sWrite u.bytes, which is the raw sixteen-byte value, and rebuild it with uuid.UUID(bytes=...). The hex form is 32 characters, int is a 128-bit integer, and str(u) is the 36-character hyphenated text — all four round-trip to the same object.
solid answer
~40 sA `uuid.UUID` is 128 bits with several read-only views onto the same value: `u.bytes` is those sixteen bytes big-endian, `u.hex` is 32 hex characters without hyphens, `u.int` is the value as one Python integer, and `str(u)` is the familiar 36-character hyphenated form. The constructor takes exactly one of the matching keywords — `uuid.UUID(bytes=...)`, `uuid.UUID(hex=...)`, `uuid.UUID(int=...)`, `uuid.UUID(fields=...)`, `uuid.UUID(bytes_le=...)` — and the first positional argument is the hex string, so `uuid.UUID("f81d4fae-7dec-11d0-a765-00a0c91e6bf6")` works and tolerates hyphens, braces and a `urn:uuid:` prefix. Bad input raises `ValueError`. Storing the sixteen bytes rather than the 36-character text more than halves the stored size and makes equality a plain byte comparison; the cost is that the column is unreadable without decoding it back.
code
python · 6 linesimport uuid
u = uuid.UUID("f81d4fae-7dec-11d0-a765-00a0c91e6bf6")
print(len(u.bytes), len(str(u)), len(u.hex))
print(uuid.UUID(bytes=u.bytes) == u, uuid.UUID(int=u.int) == u)
print(uuid.UUID("f81d4fae7dec11d0a76500a0c91e6bf6") == u)go deeper
Know that a UUID object is not a string: str(u) gives the 36-character text and uuid.UUID(text) parses it back. Remember that malformed text raises ValueError rather than returning nothing.
Explain the four spellings — bytes, hex, int and the hyphenated text — and that they are views of one 128-bit value with exact inverses in the constructor. Be precise about which constructor keyword matches which view.
Argue the storage decision: sixteen bytes at rest with text at the edges, why mixed text presentations break equality at the storage layer, and how a bytes/bytes_le mix-up corrupts identifiers without raising anything.
Set the convention across services — one canonical wire form, one at-rest form, parsing at the boundary — so that identifiers written by one component are readable by every other without a per-team decode ritual.
`uuid.UUID` is an immutable value object wrapping a single 128-bit integer. Everything you can get out of it is a *view* of that one number, and everything you can put into the constructor is a way of *spelling* it. Knowing the four spellings, and which of them you actually persist, is most of what this question is after. ### The read-only views Given `u = uuid.UUID("f81d4fae-7dec-11d0-a765-00a0c91e6bf6")`: * `u.bytes` — the sixteen bytes, most significant first. This is the canonical binary form. * `u.bytes_le` — the same sixteen bytes with the first three fields byte-swapped to little-endian. It exists for the Microsoft GUID layout, and mixing it up with `u.bytes` silently produces a different, wrong-looking identifier that still parses. * `u.hex` — 32 lowercase hex characters, no hyphens. * `u.int` — one Python `int` holding all 128 bits. * `str(u)` — the 36-character hyphenated text; `u.urn` wraps that as `urn:uuid:...`. * `u.fields`, `u.version`, `u.variant` — the structural decomposition. The object is genuinely immutable: assigning to an attribute raises `TypeError`. Because it is immutable it is hashable and orders by its integer value, so `UUID` objects work directly as dict keys, in sets, and in `sorted()` without converting to strings first. ### The constructor `uuid.UUID` takes exactly one of `hex`, `bytes`, `bytes_le`, `fields` or `int`; passing none or more than one is a `TypeError`. `hex` is also the first positional parameter, which is why `uuid.UUID(some_string)` reads naturally. The text parser is forgiving about presentation and strict about content: it strips a `urn:uuid:` prefix, surrounding braces and hyphens wherever they appear, then requires exactly 32 hex digits. Anything else — a truncated string, a stray letter, sixteen-plus-one bytes — raises `ValueError`. That strictness is worth leaning on: parsing user input through `uuid.UUID` is a genuine validation step, not a formality, and it belongs in a `try`/`except ValueError`. There is also `uuid.UUID(int=..., version=N)`, which stamps the version and variant bits over whatever you supplied. Use it only when you are deliberately constructing a value in a known layout. ### Sixteen bytes against thirty-six characters The hyphenated text form is 36 characters, and stored as ASCII that is 36 bytes plus whatever length overhead the container adds; the binary form is 16. So the text costs more than twice as much everywhere it is repeated — in a stored value, in a message payload, in a memory-resident collection of millions of ids. It also makes equality a string comparison rather than a sixteen-byte one, and if the text was ever written in two different presentations — upper case in one place, braces in another — two spellings of the same identifier stop comparing equal at the storage layer even though `uuid.UUID` would parse both to the same object. The counter-argument is human: a binary blob does not read back in a console session, does not paste into a support ticket, and needs a decode step in every ad-hoc query. The usual compromise is binary at rest with the text form at the edges — `str(u)` on the way out to a log line or an API response, `uuid.UUID(...)` on the way in, and `u.bytes` in between. ### Round-tripping Every view has an exact inverse, and this is worth demonstrating rather than asserting: ```python import uuid u = uuid.uuid4() assert uuid.UUID(bytes=u.bytes) == u assert uuid.UUID(int=u.int) == u assert uuid.UUID(u.hex) == u assert uuid.UUID(str(u)) == u ``` The one that is *not* symmetric is `bytes_le`: `uuid.UUID(bytes=u.bytes_le)` parses happily and gives you a different UUID. If you take bytes from a system that hands out GUIDs in the little-endian layout, feed them to the `bytes_le` keyword, and remember that `u.bytes` is what everything else means by "the bytes of this UUID".
- What is uuid.UUID's bytes_le attribute for, and why is confusing it with bytes dangerous?`bytes_le` is the same value with the first three fields byte-swapped, matching the little-endian GUID layout some platforms use. It is dangerous precisely because nothing complains: sixteen bytes in the wrong order still parse into a perfectly valid `UUID`, just a different one. Data written through one attribute and read through the other silently becomes a distinct identifier, and only a mismatch downstream reveals it.
- Can you rely on uuid.UUID to validate an identifier arriving from an external caller?Yes, for shape. `uuid.UUID(text)` accepts hyphens, braces and a `urn:uuid:` prefix but insists on exactly 32 hex digits and raises `ValueError` otherwise, so wrapping the parse in `try`/`except ValueError` is a real check. What it does not tell you is whether the id refers to anything, or that the caller is allowed to see it — authorisation is a separate step.
- Why can uuid.UUID objects be used directly as dictionary keys?Because they are immutable and hashable — the class refuses attribute assignment with `TypeError`, and equality and hashing are defined over the 128-bit integer. That also means they sort by that integer, so you can put them in sets, use them as keys, and call `sorted()` on them without converting to `str` first, which avoids the mismatch you get when some strings carry hyphens and others do not.
saying these in an interview costs you the question
- Says a UUID takes 36 bytes to store because it prints that way
- Mixes up bytes and bytes_le and assumes both round-trip
- Converts UUIDs to strings before using them as dict keys
- Thinks the constructor accepts hex and bytes together
- Expects malformed text to return None instead of raising ValueError
- Believes UUID objects are mutable value holders