Why does re.split return the separators when the split pattern contains a capture group?
answer
- Splitting can be made lossless
- The parentheses decide what comes back
- Grouping and capturing are separate jobs
- A colon after the question mark
- A branch not taken contributes None
basics
~20 sre.split interleaves the text of every capturing group into the result list, by design, so you can reassemble the string. Make the delimiter group non-capturing with (?:...), or drop the parentheses, to get only the fields.
solid answer
~40 sIt is documented behaviour, not a bug: if the split pattern contains capturing groups, the text captured by each group is included in the returned list between the surrounding fields. With one capturing group you get field, separator, field, separator, field; with N groups you get N entries between each pair of fields, and a group that did not participate contributes `None`. That is genuinely useful when you must rebuild the original text after editing the fields — `"".join(parts)` round-trips. When you only wanted the fields, the parentheses were there for grouping, not capturing, so write `(?:...)` instead: it groups an alternation or scopes a quantifier without opening a capture slot and without shifting the numbers of the groups that follow it.
code
pycon · 5 lines>>> import re
>>> re.split(r"\s*([,;])\s*", "gene1 , gene2; gene3")
['gene1', ',', 'gene2', ';', 'gene3']
>>> re.split(r"\s*(?:[,;])\s*", "gene1 , gene2; gene3")
['gene1', 'gene2', 'gene3']go deeper
Recall that re.split cuts a string on a pattern, and that parentheses in that pattern change what comes back. Know the non-capturing form (?:...) and when you would type it.
Explain the interleaving precisely — captured text between fields, one entry per group, None for a group that did not participate — and why the design makes the split round-trippable with "".join.
Demonstrate the maintenance angle: capture only what you read, keep the rest non-capturing so numbering cannot shift, and handle the empty-string edges from leading or trailing delimiters before they become a field-count bug.
Own the choice of tool. Decide when varying delimiters justify a regex split at all, versus a real format reader with quoting and escaping rules, and set the team convention for capture discipline in shared patterns.
### The behaviour `re.split(pattern, string, maxsplit=0, flags=0)` cuts `string` at every non-overlapping match of `pattern` and returns the pieces. The wrinkle that surprises people is the interaction with capturing groups: **if the pattern contains capturing groups, the captured text is included in the result list**, interleaved between the fields. ```pycon >>> import re >>> re.split(r"\s*([,;])\s*", "gene1 , gene2; gene3") ['gene1', ',', 'gene2', ';', 'gene3'] >>> re.split(r"\s*(?:[,;])\s*", "gene1 , gene2; gene3") ['gene1', 'gene2', 'gene3'] ``` Both patterns split at the same places. The only difference is whether the delimiter's parentheses capture. Note also that only the *captured* part is inserted — the surrounding whitespace, matched outside the group, is consumed and discarded either way. The list therefore alternates field, capture(s), field, and with one capturing group its length is `2n + 1` for `n` splits. ### Why the design is right The interleaved form makes the split **lossless**. Fields and separators together are exactly the original string, so `"".join(parts)` reconstructs it. That is how you edit fields while preserving whatever separators the input actually used — a tab here, a comma-space there — instead of normalising them on the way out. Tokenisers built on `re.split` rely on it: you get the tokens and the punctuation that joined them in one pass. If the pattern has more than one capturing group, all of them appear, in numeric order, between each pair of fields. A group inside an alternation branch that was not taken did not capture anything, so it contributes `None`: ```pycon >>> re.split(r"(a)|b", "1a2b3") ['1', 'a', '2', None, '3'] ``` Code that assumes "every odd index is a separator string" then hits `None` and fails on the second delimiter kind. Either give every branch a capture, or filter, or restructure the pattern. ### `(?:...)` — grouping without capturing Parentheses in a regex do two jobs at once: they group sub-patterns so a quantifier or an alternation applies to the whole thing, and they capture. `(?:...)` performs only the first. Use it whenever you need the grouping but not the text — which, in a split pattern, is almost always. Non-capturing groups matter beyond `re.split` for a reason that shows up in maintenance. Group numbers follow opening parentheses, so inserting a capturing group in the middle of a pattern renumbers every group after it: `m.group(3)` now points somewhere else, and a `\2` backreference means a different thing. Every group you *do not* capture is one fewer thing that can shift. The corollary is the practical rule: capture what you intend to read, group everything else with `(?:...)`, and name the captures you keep so an edit cannot silently repoint them. ### The other knobs `maxsplit` bounds the number of splits and returns the remainder as the final element: `re.split(r"[,;]", "a,b;c", maxsplit=1)` gives `['a', 'b;c']`. A leading or trailing delimiter produces an empty string at that end of the list — `re.split(r",", ",a,")` is `['', 'a', '']` — which differs from nothing at all and is a common off-by-one in field parsing. Since Python 3.7, a pattern that can match the empty string is allowed in `re.split` (it raised `ValueError` before), which makes splitting on zero-width assertions such as a word boundary possible. ### When not to reach for it `re.split` earns its keep on multi-character, alternating or irregular delimiters. If the delimiter is a fixed substring, `str.split` is faster and clearer; if the format is CSV with quoting rules, a regex is the wrong tool entirely and will fail on a quoted comma. Pick `re.split` when the delimiter genuinely varies, and keep the pattern non-capturing unless you actually want the separators back. ### In an interview Answer in one sentence — capturing groups are included in the result on purpose, so the split round-trips — then show the `(?:...)` fix, then mention the `None` for a non-participating group. That last detail is the one that shows you have debugged this, not just read the note in the docs.
- Beyond keeping re.split output clean, why prefer (?:...) for grouping?Because it does not consume a group number. Adding a capturing group in the middle of a pattern renumbers every group after it, silently repointing `m.group(3)` and any `\2` backreference. Non-capturing groups also store nothing, so `groups()` stays meaningful and the engine has one less buffer to fill. Capture what you read; group everything else with `(?:...)`.
- How do you reconstruct the original text after re.split with a capturing delimiter?Join the list back with an empty string: `"".join(parts)`. Because the captured separators are interleaved with the fields, the concatenation is exactly the input — provided the pattern captured the whole delimiter. Anything matched outside the group, such as surrounding whitespace, was consumed and is gone, so capture everything you need to round-trip.
saying these in an interview costs you the question
- Says re.split always discards the delimiters
- Calls the interleaved separators a bug in re
- Thinks (?:...) is purely a performance tweak
- Expects no empty strings from a leading delimiter
- Assumes every separator slot holds a string, never None
- Adds a group mid-pattern and keeps the old numeric indexes