Why is passing data-derived text as re.sub's replacement string unsafe, and what is the fix?
answer
- The replacement has a grammar too
- Backslashes in data are not literal
- One escaping helper guards the wrong side
- A function's return value is taken as-is
- Silent newlines are worse than the exception
basics
~20 sThe replacement string is a template that re.sub scans for group references and escapes, so a backslash in the data is interpreted rather than inserted. It corrupts output or raises re.error. Pass a callable returning the literal text instead.
solid answer
~40 s`re.sub` treats a string `repl` as a template with its own grammar: `\1` and `\g<name>` are group references, `\\` is a literal backslash, `\n` is a newline, and an unknown letter escape is an error. Text you did not write can contain any of those. `"C:\genome\1new"` inserts group 1 where you meant a directory; a stray `\g` raises `re.error` ("missing <"); a `\1` when the pattern has one group raises "invalid group reference"; a `\n` silently becomes a real newline. The fix is to bypass the template: pass `lambda m: value`, because a callable's return value is inserted verbatim with no escape processing. `re.escape` is **not** the fix — it escapes regex metacharacters for the *pattern*, a different grammar. If you must keep the string form, double the backslashes in the value first.
code
pycon · 8 lines>>> import re
>>> value = "C:\\genome\\1new"
>>> re.sub(r"x", value, "x")
Traceback (most recent call last):
...
re.PatternError: missing < at position 4
>>> re.sub(r"x", lambda m: value, "x")
'C:\\genome\\1new'go deeper
Remember that the replacement string is not plain text: backslashes in it mean something. If a replacement comes from a variable rather than a literal you typed, ask someone before shipping it.
Explain the template grammar — group references, \g<...>, letter escapes — and demonstrate the callable form as the way to insert literal text. Know that re.escape belongs to the pattern side.
Show how you would find and prevent this in a long batch job: test substitutions with hostile values, default to a callable for any data-derived replacement, and recognise the silent newline case that never raises.
Own the policy. Decide where user-configurable rewrite rules are allowed at all, how they are validated at load time rather than mid-run, and when a text-rewriting stage should be replaced by structured parsing with typed fields.
### Two grammars, one call `re.sub(pattern, repl, string)` parses two different mini-languages. The `pattern` is a regular expression, where `.`, `*`, `(`, `[` and friends are metacharacters. The string `repl` is a **replacement template**, where almost nothing is special *except* backslash sequences: * `\1` … `\99` — insert the text captured by that group. * `\g<1>`, `\g<name>` — the explicit forms; `\g<0>` is the whole match. * `\\` — one literal backslash. * `\n`, `\t` and the other standard string escapes. * `\` followed by any other ASCII letter — an error, not a literal. Everything else is copied through. The failure mode follows directly: the moment part of `repl` comes from data rather than from your source, you have handed that data a grammar. ### What actually goes wrong ```pycon >>> import re >>> re.sub(r"x", "C:\\genome\\1new", "x") Traceback (most recent call last): ... re.error: missing < at position 4 ``` The value looks like a plain Windows path. The template reader sees `\g` and demands `<`, so a six-hour nightly annotation run dies at hour five on the one record whose field held a path. Change the letter and you get a different failure for the same reason: `\1` when the pattern defines one group inserts that group's text instead of a literal `1`; `\1` with no groups raises "invalid group reference"; `\d` raises "bad escape". Worst of all is the silent case — a `\n` in the data becomes an actual newline in the output, so the file has one extra line per bad record and nothing raises at all. Corruption that does not raise is the expensive kind: you find it when a downstream consumer misaligns columns, weeks later. Note that the exception class here is `re.error` — also exposed under the name `re.PatternError` since Python 3.13, which is what a 3.14 traceback prints — and its message points at the *replacement*, not the pattern — which is why the first debugging instinct, staring at the regex, wastes the most time. ### The fix Pass a callable. Its return value is inserted verbatim, with no template scan: ```python import re value = r"C:\genome\1new" re.sub(r"x", lambda m: value, "x") # 'C:\\genome\\1new' — literal ``` This is not a trick; it is the documented difference between the two forms of `repl`, and it is the reason "when do you need a callable?" has an answer beyond "when you need arithmetic". If you are forced to keep the string form — a config file that ships templates, say — then escape the value for *that* grammar by doubling backslashes (`value.replace("\\", "\\\\")`), and be aware you have just made the value non-round-trippable if it also contained legitimate group references. ### `re.escape` is the other grammar The reflex answer in interviews is "use `re.escape`", and it is wrong here. `re.escape` prepares text for use in a **pattern**: it backslash-escapes regex metacharacters so a literal string matches itself. It knows nothing about `\g<name>`, and running it over a replacement adds backslashes that the template will then try to interpret. The correct pairing is: data going into the pattern gets `re.escape`; data going into a string replacement gets escaped for the template, or, better, does not go into a string replacement at all. ```python re.sub(re.escape(user_pattern), lambda m: user_value, text) ``` ### The related ambiguity Even with hand-written templates, the numeric form has a parsing hazard: `\10` means group 10, not group 1 followed by `0`. If the pattern has one group, that raises; if it has ten, you silently get the wrong group. `\g<1>0` is unambiguous, and is the reason to prefer the `\g<...>` spelling in any template that abuts digits. Named groups plus `\g<name>` remove the class of problem entirely. ### Operating advice Three habits make this a non-issue. First, treat the replacement side as untrusted whenever any of it is data — default to a callable. Second, unit-test the substitution with a hostile value (a backslash, a `\g`, a `\1`, a newline) rather than with clean sample rows; the clean rows will never catch it. Third, ask whether a regex is needed at all: if you are swapping a fixed substring for another fixed substring, `str.replace` has no grammar on either side and cannot be attacked by a backslash. `Match.expand` exists for the case where you genuinely want template semantics applied on purpose — for user-configurable rewrite rules, for example — and it makes that intent explicit rather than accidental. ### In an interview State that `repl` is a template, name two concrete failures (a raised `re.error` on `\g`, a silent group insertion on `\1`), give the callable as the fix, and explicitly reject `re.escape` as belonging to the pattern side. That last correction is what makes the answer read as experience rather than recall.
- What is re.escape for, and why does it not solve this?`re.escape` backslash-escapes regex metacharacters so a literal string can be embedded in a **pattern** and match itself. The replacement template is a different grammar — it cares about `\1`, `\g<...>` and letter escapes, not about `.` or `[`. Running re.escape over a replacement adds backslashes the template then tries to interpret, so it makes the problem worse rather than better.
- How do you insert group 1 immediately followed by a literal zero in a replacement template?Write `\g<1>0`. The bare `\10` is parsed as a reference to group 10, so it either raises "invalid group reference" or silently pulls the wrong group in a pattern that has ten. The `\g<...>` form ends the reference explicitly, which is why it is the safer default in any template where a digit or word character follows.
- When would you deliberately want template semantics on text you did not write?When the text is a user-supplied *rewrite rule* rather than user data — a configurable find-and-replace where the author is meant to write `\g<name>`. Then use `Match.expand` or the string repl on purpose, validate the template at load time by compiling and test-expanding it, and report a bad rule as a configuration error instead of letting it fail mid-run.
saying these in an interview costs you the question
- Reaches for re.escape to sanitize the replacement string
- Believes only the pattern needs escaping
- Thinks a callable's return value is scanned for group references
- Says \g<1> means something different from \1
- Blames the regex when the error names the replacement
- Assumes a bad replacement always raises rather than corrupting