skip to content

Why does int() refuse a 100,000-digit decimal string, and how is that limit changed?

level: seniorimportance: should knowfreq 25%

answer

  1. The integer is fine; its text is not
  2. Only bases that are not powers of two
  3. Conversion cost grows faster than the input
  4. A four-digit default, and a floor of 640
  5. One process-wide setting, not per call

basics

~20 s

Since Python 3.11 CPython caps int-to-text and text-to-int conversion in decimal at 4300 digits and raises ValueError beyond it, because that conversion is superlinear and makes a cheap denial of service. Change it with sys.set_int_max_str_digits, -X int_max_str_digits or PYTHONINTMAXSTRDIGITS.

solid answer

~40 s

The integer itself is not too big - `int` is arbitrary precision - it is the **decimal rendering** that is capped. Since 3.11, `int(s)` in base 10 and `str()`/`repr()`/`format()` of an int raise `ValueError` past `sys.get_int_max_str_digits()`, default 4300. Power-of-two bases are exempt because they are linear: `hex()`, `int(s, 16)`, `int.to_bytes()` and `int.from_bytes()` are unrestricted, as is all arithmetic. The cap exists because decimal conversion costs more than linear time, so a few hundred kilobytes of digits can burn minutes of CPU on one core. The right response to hitting it on untrusted input is to length-check the field before converting, not to call `sys.set_int_max_str_digits(0)` - that setting is interpreter-wide and disables the guard everywhere.

code

python · 11 lines
python
import sys

print(sys.get_int_max_str_digits())
big = 10**5000
try:
    str(big)
except ValueError as e:
    print("blocked:", e)
print("hex is fine:", len(hex(big)))
sys.set_int_max_str_digits(0)
print("digits:", len(str(big)))

go deeper

for a junior

Recognise the message 'Exceeds the limit for integer string conversion' and know it comes from the interpreter, not your code. The integer is not too large; only its decimal text form is capped.

for a middle

Explain which operations are limited and which are exempt, name the 4300-digit default, and describe the three ways to change it - the sys function, the -X option and the environment variable.

for a senior

Show the production judgement: diagnose it as a parse-time guard rather than a memory problem, bound untrusted numeric fields before conversion, map the ValueError to a client error at the boundary, and log lengths rather than values.

for a principal

Own the tradeoff between a hard interpreter guardrail and a workload that legitimately needs huge decimal numerals - which processes get a raised limit, what compensating input bounds they carry, and why binary encodings are the better long-term answer for large integer transport.

### The limit itself Since **CPython 3.11** the interpreter refuses to convert between an `int` and its *decimal* text form when the numeral is longer than a configured number of digits. The default is **4300 digits**, and crossing it raises `ValueError: Exceeds the limit (4300 digits) for integer string conversion`. The current value is readable with `sys.get_int_max_str_digits()`, and `sys.int_info` reports both the default and the minimum value the limit may be set to (640). The limit is a property of the *text* conversion, not of the integer. Python's `int` is still arbitrary precision, and arithmetic on huge values is untouched. What is blocked is the pair of directions that cost more than linear time: * **Blocked:** `int(s)` and `int(s, 10)` (or any non-power-of-two base) on a long digit string; `str(n)`, `repr(n)`, `format(n)`, an f-string interpolation of `n`, and anything that serialises an int by way of its decimal text. * **Not blocked:** `hex(n)`, `bin()`, `oct()`, `int(s, 16)` and other power-of-two bases, `int.from_bytes()` and `int.to_bytes()`, comparisons, and all arithmetic. Binary-to-binary conversions are linear, so there is nothing to defend. ### Why the interpreter cares Converting between base 2 (how the integer is stored) and base 10 (how it is written) is not a linear operation — historically it was quadratic in CPython, and even with the subquadratic algorithms added later a big enough numeral burns real CPU inside a single opaque C call. That gives an attacker a very cheap amplification: a few hundred kilobytes of digits in a JSON body, a query parameter or a log line turns into seconds or minutes of interpreter time on one core, with no I/O to time out against. The cap was introduced in 3.11 and backported into security patch releases of 3.7 through 3.10 for precisely that reason; it converts an unbounded compute cost into an immediate, cheap exception. ### The symptom, and why it is easy to misdiagnose Consider a flight-schedule differ that ingests two daily schedule dumps and compares fields position by position. A field that should hold a small flight number arrives, in one feed, as a wall of digits — a corrupt export, or someone probing. On an older interpreter the symptom was an **intermittent timeout**: roughly one job in a thousand parked a worker for minutes on a single `int()` call, while the process sat at a **2.4 GB working set** from the two loaded dumps. Every dashboard pointed at memory pressure and GC, because the actual culprit was one C-level conversion that never yielded and never allocated much. On 3.14 the same input raises `ValueError` at the exact parse site, with the offending length in the message — a far better failure, and the reason the cap is worth having even in a service nobody attacks. ### What to do about it The wrong reflex is to treat the exception as a CPython bug and put `sys.set_int_max_str_digits(0)` at the top of the program. The setting is **interpreter-wide**, not per-call and not per-thread, so disabling it globally to make one legitimate code path work re-opens the hole on every untrusted path in the process. The ordered options: 1. **Bound the input first.** Untrusted numeric text should be length-checked and format-checked before conversion — a flight number is at most a handful of characters. Then the limit never fires and the rejection carries a useful message. 2. **Catch it at the boundary.** `ValueError` from `int()` on untrusted input is a client error: map it to a 400-style rejection, log the *length* and the field name, never the value. 3. **Raise the limit narrowly and deliberately** if a domain genuinely needs huge decimal numerals — cryptographic parameters, arbitrary-precision maths. Set it once at startup to a documented value with `sys.set_int_max_str_digits(n)` (n must be 0 or at least 640), or from outside the code with `-X int_max_str_digits=N` or the `PYTHONINTMAXSTRDIGITS` environment variable, and record why. 4. **Avoid decimal text altogether** for big values in hot paths. `int.to_bytes()` / `int.from_bytes()` and hexadecimal are linear and unrestricted, and are the right wire format for large integers anyway. ### The interview signal A senior answer distinguishes "the integer is too big" (never true — `int` has no size limit) from "its decimal *rendering* is too long"; names the operations on both sides of the fence; and treats the limit as a guardrail whose correct response is to bound the input, with raising the cap as a considered, scoped exception rather than a reflex.

  • Does the limit affect int.from_bytes() or hexadecimal parsing?
    No. Only conversions between an int and text in a base that is not a power of two are restricted, because those are the superlinear ones. `int(s, 16)`, `hex()`, `bin()`, `oct()`, `int.to_bytes()` and `int.from_bytes()` are linear and unrestricted, which makes them the right wire format for genuinely large integers.
  • Where else does this ValueError surface besides a direct int() call?
    Anywhere an int is rendered as decimal text: `str()`, `repr()`, `format()`, an f-string interpolation, percent formatting, JSON serialisation, and logging a large value. That is what makes it confusing in production - the exception can appear in a log statement rather than at the arithmetic that produced the number.
  • A domain legitimately needs 20,000-digit decimal values. What do you do?
    Raise the limit deliberately and narrowly: call `sys.set_int_max_str_digits(n)` once at startup with a documented value, or set it from outside with `-X int_max_str_digits=N` or `PYTHONINTMAXSTRDIGITS`. It must be 0 or at least 640. Then make sure untrusted input into that process is length-bounded separately, because the guard is now weaker for everything.

It is a speed limiter, not a fuel-tank limit: the vehicle can still carry any load, but one particular manoeuvre - spelling the number out in decimal - is capped because it is where an attacker can make you burn an unbounded amount of time.

saying these in an interview costs you the question

  • Calls sys.set_int_max_str_digits(0) globally to silence the error
  • Thinks the limit caps how large an int may be
  • Believes it applies to hex() or int.from_bytes()
  • Assumes the setting is per-call or per-thread
  • Says text-to-int conversion is linear in the digit count
  • Blames memory pressure and reaches for a bigger machine

context