When would you use binascii.hexlify instead of bytes.hex() in Python?
answer
- Same digits, different container
- One returns text, one returns bytes
- The method gained formatting options
- The inverses disagree about whitespace
- Two characters per byte, always
basics
~10 sOnly when you want bytes out: they compute the same hex digits, but binascii.hexlify returns bytes and bytes.hex() returns a str. Prefer bytes.hex(), which also takes a separator argument for readable dumps.
solid answer
~40 sThey produce identical digits; the difference is the wrapper type. `binascii.hexlify(b'\xde\xad')` returns `b'dead'`, while `b'\xde\xad'.hex()` returns the `str` `'dead'` — so reach for `hexlify` only when the consumer wants bytes and you would otherwise re-encode. In new code prefer the method, which since Python 3.8 accepts a separator and a group size: `digest.hex(' ', 2)` gives `'dead beef'`, and a negative group size counts from the left instead of the right. The inverses differ in strictness: `bytes.fromhex` takes a `str` and skips ASCII whitespace, so `bytes.fromhex('de ad be ef')` works, while `binascii.unhexlify` rejects any non-hex character with `binascii.Error`. Both require an even number of digits. Remember the cost: hex doubles the payload where Base64 grows it by about a third.
code
python · 15 linesimport binascii
digest = bytes.fromhex("de ad be ef") # whitespace is skipped
print(digest) # b'\xde\xad\xbe\xef'
print(digest.hex()) # deadbeef
print(digest.hex(" ", 2)) # dead beef
print(digest.hex(" ", -3)) # deadbe ef
print(binascii.hexlify(digest)) # b'deadbeef'
print(binascii.unhexlify("deadbeef") == digest) # True
for bad in (b"dead beef", b"deadbee"):
try:
binascii.unhexlify(bad)
except binascii.Error as exc:
print(bad, "->", exc)go deeper
Know data.hex() and bytes.fromhex(text) as the everyday pair, and that hex costs two characters per byte. Never use str() on a bytes object to get a hex string — that gives you the repr.
Explain that binascii.hexlify returns bytes while the method returns a str, that they compute the same digits, and that the method's sep and bytes_per_sep arguments cover formatting. Mention that a negative group size counts from the left.
Bring in the strictness split: fromhex skips whitespace and unhexlify does not, so one of them is safe to point at pasted input. Justify hex only for short human-facing values and Base64 for payloads.
Decide the representation at the interface level — hex where a value will be eyeballed, compared in a log or pasted into a ticket; Base64 where volume dominates — and record the choice so identifiers do not appear in three encodings across a fleet.
### The same computation, two wrappers Hexadecimal encoding is the simplest binary-to-text scheme there is: each byte becomes exactly two characters from `0-9a-f`. There is no grouping, no padding, no alphabet variant and no state — which is why it can be split anywhere and concatenated freely, unlike Base64. Python gives you two doors to the same computation. * `bytes.hex()` — a method on `bytes`, `bytearray` and `memoryview`, returning a `str`. * `binascii.hexlify(data)` — a function taking any buffer, returning `bytes`. The digits are identical. The honest answer to "when would you use `hexlify`" is: when the destination wants bytes — a socket write, a buffer you are assembling, an API that refuses `str` — and you would otherwise write `.hex().encode('ascii')`, paying for a second conversion. `hexlify` also predates the method, which is why so much existing code calls it; that is history, not a reason to write it in new code. ### What the method has that the function does not Since Python 3.8, `bytes.hex` takes two optional arguments, `sep` and `bytes_per_sep`: ```python b = bytes.fromhex("0011223344") b.hex() # '0011223344' b.hex("-") # '00-11-22-33-44' b.hex("-", 2) # '00-1122-3344' b.hex("-", -2) # '0011-2233-44' ``` A positive `bytes_per_sep` groups from the **right**, which is what you want when the value reads as a number and the low bytes should line up. A negative one groups from the **left**, which is what you want for a packet or record dump where the header fields start at offset zero. `binascii.hexlify` gained the same two arguments in the same release and returns the grouped digits as bytes, so the formatting power is equal and the return type stays the only real difference. ### The two inverses, and their different tempers Going back to bytes, the pair is not symmetric in strictness, and interviewers like this detail because it decides which one you can point at user input. `bytes.fromhex(s)` is a `bytes` classmethod taking a `str`, and it **skips ASCII whitespace** — spaces, tabs and newlines — anywhere between digit pairs. So `bytes.fromhex('de ad be ef')` and `bytes.fromhex('de\tad\nbe ef')` both give `b'\xde\xad\xbe\xef'`. That makes it the right tool for text a human typed or pasted, or for a hex dump you are re-reading. (Whitespace skipping widened from spaces to all ASCII whitespace in Python 3.7.) `binascii.unhexlify(s)` accepts bytes or an ASCII `str` and is strict: any non-hex character, whitespace included, raises `binascii.Error: Non-hexadecimal digit found`. Both functions reject an odd number of digits — `binascii.Error: Odd-length string` from `unhexlify`, `ValueError` from `fromhex` — because half a byte is not a byte. Note that `binascii.Error` subclasses `ValueError`, so a single `except ValueError` covers the whole family. ### Choosing hex or Base64 Both solve the same problem — getting arbitrary bytes through a text-only channel — and they trade size against convenience. Hex doubles the payload: one byte in, two characters out, a flat 100% overhead. Base64 costs about a third more than the input, since three bytes become four characters. On a 32-byte digest the difference is 64 characters against 44 and nobody cares; on a multi-megabyte blob the difference is the whole conversation. What hex buys for that overhead is total simplicity. It is case-insensitive on input, has no padding, no alphabet variants, no `validate` flag, and no alignment requirement — you can cut a hex string at any even offset and the pieces still decode independently. Base64 has a URL-safe alphabet, `=` padding rules and a three-byte alignment that make every one of those statements conditional. So the working rule is: **hex for short, human-facing, eyeball-comparable values** — digests, keys in a log line, magic numbers, a packet dump you will read in a terminal — and **Base64 for payloads**, where the size actually costs you something. ### The mistakes to avoid The recurring one is `str(b'\xde\xad')`, which yields `"b'\\xde\\xad'"` — the repr, complete with prefix, quotes and escapes. It looks like hex at a glance and is not. The second is hand-rolling `''.join(f'{c:02x}' for c in data)`, which is correct but slower and longer than `data.hex()`. The third is assuming the two families interoperate: `hexlify` output is bytes, so `binascii.unhexlify(digest.hex())` works only because `unhexlify` also accepts an ASCII `str`.
- What does bytes.fromhex tolerate that binascii.unhexlify rejects?ASCII whitespace between digit pairs. `bytes.fromhex('de ad be ef')` succeeds, while `binascii.unhexlify(b'de ad be ef')` raises `binascii.Error` — whitespace counts as a character, so you get either `Non-hexadecimal digit found` or `Odd-length string` depending on how many were there. Both reject an odd number of real hex digits. Use `fromhex` for anything a human typed or pasted.
- How do you print a readable hex dump of a record header?`data.hex(sep, bytes_per_sep)`, added in Python 3.8. `header.hex(' ', 4)` groups four bytes per chunk counting from the right, which suits values read as numbers; a negative count, `header.hex(' ', -4)`, groups from the left, which is what you want when field offsets start at zero. No loop and no third-party helper is needed.
- Why would you choose Base64 over hex for a large payload?Size. Hex is a flat 100% overhead — two characters per byte — while Base64 turns three bytes into four characters, about 33%. On a 32-byte digest that is 64 characters against 44 and it does not matter; on a multi-megabyte blob it is the difference between doubling the transfer and adding a third. Hex earns its overhead only when a human reads the value.
saying these in an interview costs you the question
- Thinks bytes.hex() returns bytes like hexlify
- Uses str() on bytes and calls the repr hex
- Believes the two produce different digit values
- Assumes bytes.fromhex accepts an odd digit count
- Hand-rolls a per-byte f-string loop instead
- Claims hex and Base64 cost the same size