skip to content

What do the 0b, 0o and 0x prefixes and underscores mean in Python numeric literals?

level: juniorimportance: nice to knowfreq 30%

answer

  1. Four ways to write the same object
  2. The value forgets how it was spelled
  3. Separators are for eyes, not for the machine
  4. They must sit between two digits
  5. Python 3 refused C's leading-zero octal

basics

~20 s

0b1010, 0o17 and 0x1f are binary, octal and hexadecimal ways of writing an ordinary int. Underscores such as 1_000_000 are visual digit separators with no runtime effect. A plain decimal literal may not start with a zero: 010 is a SyntaxError.

solid answer

~40 s

The prefixes are case-insensitive alternate spellings, not types: `0x1f` produces the plain `int` 31, and `repr()` prints it as `31` because the object keeps no memory of how it was written. To see it in hex again use `hex()`, `bin()`, `oct()` or a format specification such as `f"{31:#x}"`. Underscores arrived in **3.6** (PEP 515) purely as reader-facing grouping - `1_000_000` compiles to the same constant as `1000000` - and must sit between digits, with the one exception that they may follow a base prefix (`0x_ff`). Leading, trailing or doubled underscores are a `SyntaxError`. The same separators are accepted by the string parsers, so `int('1_000')` is 1000 and `float('1_000.5')` is 1000.5. `010` is a deliberate `SyntaxError` in Python 3 rather than octal 8.

code

python · 7 lines
python
print(0b1010, 0o17, 0x1f, 0xdead_beef, 1_000_000)
print(0x_ff, int("1_000"), float("1_000.5"))
print(repr(0x1f), hex(31), f"{31:#x}", f"{1000000:_d}")
try:
    eval("010")
except SyntaxError as e:
    print("SyntaxError:", e)

go deeper

for a junior

Be able to read the four literal forms and say what each is worth, and know that underscores are purely for readability. Remember that a leading zero on a plain decimal number is an error, not octal.

for a middle

Explain that the prefixes vanish at compile time into ordinary ints, name the formatting route back out (hex/bin/oct or a format spec), and state the underscore placement rules including the base-prefix exception.

for a senior

Show taste about when an alternate base earns its place - masks, protocol constants, colour values where the bit pattern is the meaning - and keep literal formatting consistent so readers are not decoding radixes on every line.

for a principal

This is style surface, and style surface is worth settling once: which radixes appear in domain constants, whether long numerals carry separators, and letting a formatter or linter enforce it rather than reviewers arguing it per patch.

### The three alternate bases Python source can write an integer in four radixes. Decimal needs no prefix; `0b` (or `0B`) introduces binary, `0o`/`0O` octal, and `0x`/`0X` hexadecimal. Hex digits may be upper or lower case. Every one of them produces an ordinary `int` object — there is no "hex integer" type, the prefix is purely a way of *writing* the value: ```python 0b1010 == 10 0o17 == 15 0x1f == 31 type(0x1f) is int # True repr(0x1f) == '31' # the object has no memory of how it was written ``` That last line is the point people miss. Once the source is compiled, the literal is a plain `int` constant; printing it gives decimal. If you want to see the value in hex again you have to ask, with the `hex()`, `bin()` and `oct()` builtins or a format specification: `f"{31:#x}"` gives `'0x1f'`, `f"{31:b}"` gives `'11111'`, and `f"{255:#06x}"` gives `'0x00ff'` with padding. Alternate bases earn their place where the *bit pattern* is the meaning: permission masks, colour values, protocol constants, byte-level flags. Written in decimal those values are unreadable; `0o644` and `0xff00ff` say what they mean. ### A plain decimal literal may not start with a zero `010` is a `SyntaxError` in Python 3, and the interpreter says why: *leading zeros in decimal integer literals are not permitted; use an 0o prefix for octal integers*. Python 2 read `010` as octal 8, following C, and that convention had caused enough quiet bugs — a zero-padded number pasted from a form or an ID becoming a different value — that Python 3 made it an error rather than choose either meaning. A single `0` is fine, as is a run of zeros with no other digits (`000`), and zero-padding is fine in a *string* (`int("010")` is 10). Only the bare decimal literal is refused. ### Underscores are visual separators and nothing else PEP 515, in **Python 3.6**, allowed single underscores inside numeric literals purely as grouping for the reader. They have no runtime existence at all: `1_000_000` and `1000000` compile to the identical constant, with no cost and no difference in `repr`. The rule is that an underscore must sit between two digits, with one convenience exception — it may also follow a base prefix: ```python 1_000_000 # 1000000 0xdead_beef # 3735928559 - grouping by 16-bit halves 0x_ff # 255 - allowed right after the prefix 1_000.000_1 # floats too 1_000j # and imaginary literals ``` What is rejected, all as `SyntaxError`: a leading underscore (`_1000` is a perfectly good *identifier*, so it can never be a number), a trailing one (`1000_`), doubled ones (`1__000`), and one adjacent to the `.` or the exponent marker (`1._5`, `1e_5`). The same separators are accepted by the *runtime* parsers, which is the part that catches people out in a good way: `int("1_000")` is 1000 and `float("1_000.5")` is 1000.5, in every base. So text that a human formatted for readability round-trips. The converse also exists in formatting — `f"{1000000:_d}"` produces `'1_000_000'`, and the `,` format type gives comma grouping instead. ### Where this shows up in interviews It is a small question and rarely a gate, but it separates people who have read the language reference from people who have only ever typed decimals. The three things worth having ready: the prefixes produce plain `int`s and the value has no memory of its literal form; underscores are compile-time cosmetics that also work in `int()` and `float()` string parsing; and `010` is a deliberate `SyntaxError`, not octal, because Python 3 refused to inherit C's ambiguity. If you can add *why* — that alternate bases exist for bit patterns, and separators exist because a nine-digit constant is unreadable — you have said everything the question is worth.

  • Is 1_000 any slower at runtime than 1000?
    No. Underscores exist only in the source text; the compiler produces the identical integer constant, so there is no parsing, no stripping and no cost at run time. The only place they survive is a string you parse yourself, where `int('1_000')` accepts them by the same rule.
  • Where exactly is an underscore rejected in a numeric literal?
    Leading (`_1000` is an identifier, not a number), trailing (`1000_`), doubled (`1__000`), and adjacent to the decimal point or the exponent marker (`1._5`, `1e_5`) - all `SyntaxError`. The only non-between-digits position allowed is immediately after a base prefix, as in `0x_ff`.
  • Why did Python 3 make 010 an error instead of choosing a meaning?
    Python 2 read it as octal 8, following C, and zero-padded numbers pasted from forms or identifiers silently became different values. Rather than pick either reading, Python 3 rejects the ambiguous form and the error message points at the explicit `0o` prefix. Zero-padding in a string is unaffected: `int('010')` is 10.

saying these in an interview costs you the question

  • Thinks 0x1f is a string rather than an int
  • Says 010 is octal 8 in Python 3
  • Believes underscores change the value or cost time
  • Expects int('1_000') to raise ValueError
  • Assumes repr of 0x1f prints 0x1f
  • Claims hex literals need a special type for arithmetic

context