skip to content

When is str.translate with str.maketrans better than chained str.replace calls?

level: seniorimportance: nice to knowfreq 22%

answer

  1. Chained calls can see each other's output
  2. One pass over the text, not N
  3. The table is built once and reused
  4. Maps code points, not words
  5. A third argument deletes characters

basics

~20 s

str.translate applies a whole per-character mapping in one pass, so replacements never cascade into each other and cost stays linear in the text. Chained str.replace calls each scan the string again and can re-transform an earlier call's output.

solid answer

~50 s

`str.translate(table)` walks the string once and looks up every character in a mapping from code point to replacement — another string, an integer code point, or `None` to delete it. `str.maketrans` builds that table from a dict, from two equal-length strings, or from two strings plus a third listing characters to delete. Two properties matter. First, **no cascading**: `'a-b'.replace('a','b').replace('b','c')` gives `'c-c'` because the second call sees the first call's output, while the equivalent table gives `'b-c'`. Second, **one pass**: a chain of N replaces scans and rebuilds the string N times, whereas translate does it once, and the table can be built once at module level and reused. The limit is that translate maps single characters only — for multi-character sequences you still need `str.replace`. `bytes` and `bytearray` have their own `translate` and `maketrans` taking a 256-entry byte table.

code

python · 3 lines
python
s = 'a-b'
print(s.replace('a', 'b').replace('b', 'c'))   # 'c-c' - the new b was replaced again
print(s.translate(str.maketrans('ab', 'bc')))   # 'b-c' - one pass, no cascade

go deeper

for a junior

You are unlikely to be asked this. Know that a per-character translation table exists and that string methods return new strings; reaching for replace is a fine answer at this level.

for a middle

Be able to build a table with the static helper in both the two-string and the dict form, and explain that a missing key leaves the character alone while a None value deletes it.

for a senior

Lead with correctness, not speed: chained replaces can re-transform their own output when the search and replacement alphabets overlap, and the resulting bug is order-dependent and easy to miss in a fixture set. Then add the single-pass cost argument and the deletion form.

for a principal

Judge when hand-rolled normalisation is the right layer at all. A per-character table is cheap and auditable for a fixed rule set, but once the rules grow context-sensitive the answer is a parser or a normalisation library, not a longer chain of string calls.

### What `translate` actually does `str.translate(table)` iterates the string once, and for each character looks up its **ordinal** (its Unicode code point, as `ord` would give it) in `table`. The lookup can produce: - a **string** — substituted in place of the character (it may be longer than one character, or empty), - an **integer** — treated as the code point of the replacement character, - `None` — the character is deleted, - nothing at all (a `KeyError` from the mapping) — the character is passed through unchanged. The table is any object supporting `__getitem__` on an integer key, which is why a plain `dict` works. `str.maketrans` is the static helper that builds one correctly, in three forms: - `str.maketrans(mapping)` — a dict whose keys are single-character strings or ordinals, and whose values are strings, ordinals or `None`. - `str.maketrans(from_chars, to_chars)` — two equal-length strings, mapped position by position. - `str.maketrans(from_chars, to_chars, delete_chars)` — as above, plus a third string of characters to drop entirely. ### Property one: replacements never cascade This is the correctness argument, and it is the one worth having ready. ```python s = 'a-b' s.replace('a', 'b').replace('b', 'c') # 'c-c' s.translate(str.maketrans('ab', 'bc')) # 'b-c' ``` The chained version is wrong for the intended mapping because the second `replace` cannot tell an original `b` from the `b` the first call just produced. Every chain of `replace` calls whose replacement alphabet overlaps its search alphabet carries this bug. `translate` is immune by construction: each *input* character is examined exactly once, and whatever it becomes is never re-examined. The failure is order-dependent, which makes it nasty. Reorder the two calls and the output changes; the version that happens to be correct for today's inputs breaks when a new character is added to the mapping. It is the kind of defect that reaches a large regression pack — a 340-case sweep over records from an inventory sync between two systems, say — and passes every case, because none of the fixtures happens to contain a character that appears on both sides of the mapping. It only fails when a feed arrives whose text carries a character an earlier rule introduces. ### Property two: one pass, and a table you build once A chain of N `replace` calls performs N full scans and builds N intermediate strings, each of which is allocated and thrown away — `str` is immutable, so there is no in-place option. `translate` performs one scan and builds one result, and because the table is an ordinary object you can build it once at import time and reuse it for every record. In a normalisation step that runs over every row of a feed, that difference is the whole cost of the step. Do not oversell it as the headline, though: the correctness argument is the one an interviewer is really testing, and the performance claim only becomes material at volume. ### Deletion is the underrated form The three-argument form is the cleanest way to strip a set of characters from anywhere in a string — not just from the ends, which is what `strip` does: ```python drop = str.maketrans('', '', ' -_') 'SKU 1120-A_B'.translate(drop) # 'SKU1120AB' ``` That is the canonical way to canonicalise identifiers before comparing them, and it beats a chain of three `replace` calls on both counts above. Equivalently, a dict mapping each ordinal to `None` does the same job when you want to build the set programmatically. ### The limits `translate` is strictly **per character**. It cannot map a two-character sequence to something else, cannot do context-sensitive substitution, and cannot reorder. The moment your rule is about a *word* or a *sequence*, `str.replace` is the correct tool and the chain question becomes "in what order, and does the alphabet overlap?" — a question you answer by making the replacements disjoint, or by doing them in one pass some other way. `bytes` and `bytearray` have their own `translate` and `maketrans`, but the contract differs: the table is a 256-byte object mapping every possible byte value, and deletion is a separate `delete` argument rather than a `None` value. Do not carry the `str` mental model across unexamined — and never build a `str` table and hand it to a bytes method. ### Interview framing This is a differentiator rather than a gate: plenty of strong engineers have never needed `translate`. What earns credit is naming the cascading-replace hazard first and the single pass second, showing the deletion form, and knowing the boundary — per character only, so multi-character rules stay with `replace`.

  • What does str.translate do with a character that is not in the table?
    It passes it through unchanged. The lookup is by ordinal, and a missing key — a KeyError from the mapping — means 'leave this character alone', which is why a small dict works as a table for a large alphabet. Mapping a key to None is different: that deletes the character. So absent means keep, None means drop.
  • Can str.translate replace a two-character sequence such as a carriage-return newline pair?
    No. The lookup is per character, so a table can never see a sequence. It can map each character independently — including mapping one to an empty string — but a rule that depends on adjacency belongs to str.replace or a real parser. That per-character boundary is the main reason translate does not replace replace outright.
  • How does bytes.translate differ from str.translate?
    The table is a 256-byte object mapping every possible byte value rather than a sparse mapping of code points, deletion is a separate delete argument instead of a None value, and the arguments must be bytes-like. bytes.maketrans builds the table from two equal-length bytes objects. Handing a str table to a bytes method is a TypeError, which is the right outcome — the two are different alphabets.

Chained replaces are like editing a document in successive passes, where each pass can revisit text an earlier pass wrote; translate is a single read-through with a lookup card, where every original character is decided once and never revisited.

saying these in an interview costs you the question

  • Assumes chained replace calls cannot affect each other
  • Thinks translate can map multi-character sequences
  • Confuses an absent table key with a None value
  • Rebuilds the translation table inside a hot loop
  • Uses strip to remove characters from the middle
  • Passes a str translation table to a bytes method

context