How does int(s, base) parse text, and why is base 0 special?
answer
- The second argument only makes sense for text
- Two to thirty-six, or one special value
- Letters carry the digits above nine
- Zero means read it as source code
- That is why '010' is refused
basics
~20 sint(s, base) reads the text as a numeral in that radix, which must be 0 or 2 through 36, with letters serving as digits above 9. Base 0 means infer the radix from a 0b/0o/0x prefix, and it rejects a decimal numeral with leading zeros.
solid answer
~40 sThe second argument selects the radix and applies only to text - `str`, `bytes` or `bytearray`. Passing a base with a number raises `TypeError: int() can't convert non-string with explicit base`, which is why `int(3.7, 10)` fails. Legal values are `0` or `2`-`36`; letters `a`-`z` (case-insensitively) are the digits above 9, so `int('zz', 36)` is 1295. A `0b`/`0o`/`0x` prefix is allowed when it matches the base, so `int('0xff', 16)` is 255 while `int('0x1f', 10)` raises `ValueError`. **Base 0** means read the text the way Python source would: the prefix picks the radix, no prefix means decimal, and Python's own rule against redundant leading zeros applies - `int('010', 0)` raises, though `int('010', 10)` is 10.
code
python · 6 linesprint(int("ff", 16), int("0xff", 16))
print(int("0b1010", 0), int("0o17", 0), int("101", 2))
try:
int("010", 0)
except ValueError as e:
print("base 0 rejects leading zeros:", e)go deeper
Know that int() takes an optional second argument naming the radix, that int('ff', 16) is 255, and that hex text will not parse without it. The default is base 10.
Explain the mechanics: the 0-or-2-to-36 range, letters as digits above nine, prefixes allowed when they match the base, the TypeError when a base accompanies a non-string, and what base 0 infers.
Demonstrate judgement about which base a parsing boundary should demand - explicit for protocol-fixed formats so a wrong radix fails loudly, base 0 for human-authored config - and reject try-every-base parsing as a source of silently wrong values.
Own the convention across a codebase: where numeric text enters, which radixes are accepted, and how the choice is documented. Accepting several spellings of a number is an interface commitment, not a local convenience.
### The signature `int()` has two call shapes: `int(x)` converts a number or parses a string in base 10, and `int(x, base)` parses text in an explicit base. The second argument is **only** meaningful for text: `str`, `bytes` and `bytearray` are accepted, and anything else with an explicit base raises `TypeError: int() can't convert non-string with explicit base`. That is why `int(3.7, 10)` is an error rather than 3 — a float has no digits to interpret in a base. ### Legal bases and their digits `base` must be `0`, or an integer from `2` to `36` inclusive; anything else raises `ValueError: int() base must be >= 2 and <= 36, or 0`. Above base 10 the extra digits are letters, case-insensitively: `a`/`A` is 10 through `z`/`Z` is 35, which is why 36 is the ceiling. `int("zz", 36)` and `int("ZZ", 36)` both give 1295. Underscore separators and surrounding whitespace are tolerated in every base, exactly as in base 10, so `int(" ff_ff ", 16)` parses. ### Prefixes are allowed when they agree with the base A `0b`, `0o` or `0x` prefix may appear in the text when it matches the base you asked for. `int("0xff", 16)` is 255 and `int("0b1010", 2)` is 10. If the prefix disagrees with the base you get the ordinary parse failure: `int("0x1f", 10)` raises `ValueError`, because in base 10 the character `x` is not a digit. This matters when you accept identifiers copied from logs or configuration where the prefix may or may not be present — passing base 16 handles both `"ff"` and `"0xff"` with one call. ### Base 0 means "read it the way Python source would" `base=0` is the interesting case and the one interviews turn on. It does not mean base 10 and it does not mean "guess". It means: interpret the text using Python's own integer-literal rules. The prefix decides the radix — `int("0b1010", 0)` is 10, `int("0o17", 0)` is 15, `int("0x1f", 0)` is 31 — and with no prefix the text is read as decimal. Crucially, base 0 also inherits Python 3's rule that a decimal literal may not carry redundant leading zeros, so: ```python int("010", 0) # ValueError: invalid literal for int() with base 0: '010' int("010", 10) # 10 - explicit base 10 is happy to ignore the zero int("000", 0) # 0 - a run of zeros with no other digits is fine ``` Note what base 0 does *not* do: it does not resurrect the C convention where a leading zero means octal. It refuses the ambiguity rather than resolving it, and this is exactly why `base=0` is the right choice for a configuration file that wants to accept `0x1f`, `0b1010` and `31` from a human — and the wrong choice for a fixed-width numeric field where `"007"` is legitimate data. ### Going the other way There is no general "int to base-N string" function in the standard library. The three power-of-two radixes have builtins — `bin()`, `oct()` and `hex()` — and the same conversions are available as format specifications, which is usually the nicer form because you control width and prefix: `format(255, "x")` is `'ff'`, `f"{255:#06x}"` is `'0x00ff'`, and `f"{255:b}"` is `'11111111'`. For an arbitrary base you write the divmod loop yourself. Round-tripping is then `int(hex_text, 16)`. ### Choosing a base in real code Three rules of thumb are worth stating out loud in an interview. Pass an **explicit base** whenever the format is fixed by a protocol or a file format — a hex colour, a permission mask, a checksum — so a stray prefix or a wrong-radix value fails loudly. Pass **base 0** when the text is authored by a person who may reasonably write any Python literal form. Pass **no base at all** for ordinary decimal user input. And never build a parser that tries several bases in a loop until one succeeds: it turns a typo into a plausible wrong number, which is the failure mode explicit bases exist to prevent.
- How do you turn an integer back into a base-16 or base-2 string?Use the `hex()`, `oct()` and `bin()` builtins for a prefixed form, or a format specification when you want control: `format(255, 'x')` gives `'ff'` and `f"{255:#06x}"` gives `'0x00ff'`. The standard library has no general base-N formatter, so an arbitrary radix means writing a `divmod` loop yourself.
- When would you pass base 0 rather than base 10 for a configuration value?When a human authors the value and any Python literal form should be acceptable - `31`, `0x1f` and `0b11111` all meaning the same thing. Base 0 is wrong for fixed-width numeric fields where `'007'` is legitimate data, because it refuses redundant leading zeros rather than resolving them.
- Is it reasonable to try several bases in a loop until one parses?No. It turns a typo into a plausible wrong number: `'11'` parses in every base and means something different in each. Either the format is fixed by a protocol, in which case pass that base explicitly and let a mismatch fail, or the text carries its own prefix, in which case pass base 0.
saying these in an interview costs you the question
- Thinks any integer is a valid base, such as 64
- Says int('0xff') works without an explicit base
- Believes base 0 is just another spelling of base 10
- Expects int('010', 0) to give 8 the way C would
- Passes a base along with a float or an int
- Assumes digits above nine must be uppercase letters