skip to content

Why does int() on a 5,000-digit string raise ValueError, and how do you lift the cap?

level: juniorimportance: should knowfreq 28%

answer

  1. Decimal conversion is not linear
  2. There is a digit ceiling
  3. Base 16 never pays this cost
  4. Around four thousand digits
  5. A sys setter and PYTHONINTMAXSTRDIGITS

basics

~20 s

CPython caps conversion between a decimal string and an int at 4,300 digits, because that conversion is quadratic in the digit count and makes a cheap request expensive. Lift it with sys.set_int_max_str_digits(), PYTHONINTMAXSTRDIGITS, or -X int_max_str_digits.

solid answer

~40 s

Converting a decimal string to an int is not linear: CPython stores integers in base 2**30, and because 10 is not a power of two each decimal digit has to be folded in with a multiply-and-add, giving roughly quadratic cost. A few hundred kilobytes of digits in an untrusted field therefore buys minutes of CPU, so CPython 3.11 (backported to 3.10.7 and other security branches) added a default limit of 4,300 digits and raises `ValueError` past it. The limit covers `int(s)` and `str(n)`, plus `repr` and f-string formatting, in base 10 and other non-power-of-two bases. It does **not** cover hex, octal or binary, `int.from_bytes`/`int.to_bytes`, or arithmetic between ints. Raise it with `sys.set_int_max_str_digits(n)` at runtime, `-X int_max_str_digits=N` at launch, or the `PYTHONINTMAXSTRDIGITS` variable; 0 disables it, and any positive value must be at least 640.

code

python · 10 lines
python
import sys

print(sys.get_int_max_str_digits())   # 4300
try:
    int("9" * 5000)
except ValueError as exc:
    print("rejected:", exc)

sys.set_int_max_str_digits(10_000)
print(len(str(int("9" * 5000))))      # 5000

go deeper

for a junior

Recall the number and the message: 4,300 digits by default, a ValueError that names the limit, and sys.set_int_max_str_digits() as the way to raise it. Knowing it applies to base 10 but not to hex is enough at this level.

for a middle

Explain the mechanics — CPython's base-2**30 internal digits, the quadratic multiply-and-add for non-power-of-two bases — and name all three ways to change the limit plus the 640 floor and the 0-means-unlimited case.

for a senior

Show the judgment about who controls the string: default stays for untrusted input, a justified number for owned computations, an explicit length check in front either way. Be ready to note that the setting is process-wide and not thread safe.

for a principal

Own the encoding decision. Argue for carrying large integers as hex or bytes across service boundaries so the quadratic never appears, and treat a request to disable the cap in a shared image as a policy question about which processes touch untrusted input.

## The cost that the limit is protecting Python's `int` is arbitrary precision, and CPython stores one internally as an array of digits in base 2**30. Moving between that representation and a *string* of digits is only cheap when the string's base is a power of two: converting `0xdeadbeef` is a regrouping of bits, so hex, octal and binary conversions are linear in the length of the string. Base 10 is not a power of two. Every decimal digit has to be folded into the accumulating binary value with a multiply-and-add over the whole number so far, and the schoolbook algorithm CPython uses for this is O(d**2) in the number of decimal digits. That constant is small enough that nobody notices at ten or a thousand digits, and catastrophic at a million: a 400 KB run of digits is on the order of 10**12 digit operations. The security shape follows directly. Any place a service calls `int()` on something a client sent — a JSON number, a form field, a header, a path segment, a CSV column — is a place where a small request body converts into minutes of CPU. In CPython that work happens in C with the GIL held, so it is not merely one slow request; it stalls the whole interpreter. The same applies in reverse: `str(n)` on an integer the attacker made large (say, by supplying a big exponent to a `**`) is just as quadratic. ## The knob CPython 3.11 introduced a default ceiling of 4,300 decimal digits, and the fix was backported to 3.10.7, 3.9.14, 3.8.14 and 3.7.14. Exceeding it raises `ValueError` with a message that names the limit and points at the setter: ``` ValueError: Exceeds the limit (4300 digits) for integer string conversion: value has 5000 digits; use sys.set_int_max_str_digits() to increase the limit ``` Three ways to change it, all process-wide: ```python import sys sys.set_int_max_str_digits(10_000) # at runtime print(sys.get_int_max_str_digits()) # read it back ``` ``` python -X int_max_str_digits=10000 report.py PYTHONINTMAXSTRDIGITS=10000 python report.py ``` Passing `0` disables the check entirely. Any other value must be at least 640 — below that, `set_int_max_str_digits` itself raises `ValueError: maxdigits must be >= 640 or 0 for unlimited`. `sys.int_info` reports both the 4,300 default and the 640 floor for the running build. ## What is and is not covered Covered: `int(s)` where the base is 10 or another non-power-of-two base; `str(n)`; `repr(n)`; `f"{n}"` and `"%d" % n`; anything that ultimately calls those. Not covered: `int(s, 16)`, `int(s, 8)`, `int(s, 2)` and the matching `hex`/`oct`/`bin` output; `int.from_bytes` and `int.to_bytes`; every arithmetic operation between integers, however large the operands; and conversion from a float. Those are all linear or already bounded, so there is nothing for the cap to protect. That asymmetry is also the practical workaround. If you genuinely need to move very large integers between systems — a key, a checksum, an accumulated counter — carry them as hex or as bytes rather than as decimal text. Both directions stay linear and neither trips the limit. ## Living with it The first encounter is usually a false positive rather than an attack: a script prints a large factorial, a notebook renders a big power, a test fixture holds a long decimal literal read from a file, and code that worked on an older interpreter now raises. The correct response depends on who controls the string. - **You control it** (a computation, a fixture, a report): raise the limit, and prefer raising it at launch with `-X int_max_str_digits` or the environment variable so the setting is visible in the process's configuration rather than buried in a module import. - **A client controls it**: leave the default alone. If the domain honestly needs longer values, raise the limit to a specific number that the domain justifies, and add an explicit length check on the field before conversion — the limit is a backstop, not input validation. - **Never reach for `0`** to make a failing test pass in code that also serves untrusted input. Disabling the check removes the only bound on that quadratic. Two operational notes. The setting is per *process*, not per thread, so a "set it, convert, set it back" wrapper is not thread safe — another thread's conversion can run inside your window. And it is read at interpreter start when supplied by the environment variable or `-X`, so setting `os.environ` from inside a running program has no effect on the current process; it only affects children you spawn afterwards.

  • Which conversions does the digit limit deliberately leave alone, and why those?
    Hex, octal and binary in both directions, `int.from_bytes` and `int.to_bytes`, and arithmetic between integers. A power-of-two base maps onto CPython's internal base-2**30 digits by regrouping bits, so those conversions are linear and there is no quadratic blow-up to bound. The cap exists only for bases that are not powers of two, where each digit costs a multiply-and-add over the whole accumulated value.
  • A service must accept 20,000-digit decimal values from clients. How do you configure it?
    Set the limit to a specific number the domain justifies — `-X int_max_str_digits=25000` or `PYTHONINTMAXSTRDIGITS` at launch, so it is visible in the process configuration — and add an explicit length check on the field before converting, rejecting anything longer than the contract allows. Never pass 0. Better still, negotiate a hex or byte encoding for the field, which stays linear in both directions.
  • Why is a set-convert-restore wrapper around sys.set_int_max_str_digits risky?
    The limit is process-wide, not thread-local. While your wrapper has the ceiling raised, any other thread doing a conversion runs under the raised value, so a request handler on another thread loses its protection for the duration. If you need a wide limit for one computation, prefer setting it once at process start for a process that only does that work.

Reading a number in hex is like re-cutting a ribbon into different-sized pieces; reading it in decimal is like re-weighing the whole pile each time you add a coin.

saying these in an interview costs you the question

  • Claims the limit applies to all integer arithmetic
  • Thinks int() on a decimal string is linear in length
  • Sets the limit to 0 globally to make a test pass
  • Believes hex and binary strings hit the same cap
  • Says the cap prevents integer overflow in Python
  • Assumes os.environ can change it mid-process

context