skip to content

Dict API, Views, and Merging

Everyday dict work: safe lookups with get and setdefault, the live views keys/values/items return, and merging with update, {**a, **b} or |. Interviewers probe views because people expect lists.

part ofPythonoverview, primer and where to startread it →
on this pageshow

questions

4

How do `dict.get` and `dict.setdefault` differ from indexing a Python dict with `d[key]`?

level: juniorimportance: must knowfreq 78%

answer

  1. Three ways to reach one key
  2. Only one of them writes
  3. The default is always built
  4. get returns, setdefault inserts

basics

~20 s

Indexing a missing key raises KeyError. dict.get returns a default instead - None unless you pass one - and never changes the dict. dict.setdefault returns the existing value or inserts your default and returns that, so it mutates.

solid answer

~40 s

`d[key]` goes through `__getitem__` and raises `KeyError` when the key is absent, which is what you want when a missing key means a bug. `dict.get(key)` returns `None` for a missing key and `dict.get(key, default)` returns the default; either way the dict is left untouched, so `get` is the read-only safe lookup. `dict.setdefault(key, default)` is a read *and* a write: it returns the current value if the key exists, otherwise it inserts `default` under that key and returns the object it just stored. That makes it the one-liner for accumulating groups - `groups.setdefault(dept, []).append(name)` - because the list handed back is the one now living in the dict. The trap is that the default argument is evaluated on every call, even on a hit, so an expensive default gets built and thrown away.

code

python · 8 lines
python
counts = {}
print(counts.get("apples"), counts.get("apples", 0))
print(counts)

groups = {}
for name, dept in [("ada", "eng"), ("bo", "ops"), ("cy", "eng")]:
    groups.setdefault(dept, []).append(name)
print(groups)

go deeper

for a junior

Be ready to say without hesitating which of d[key], dict.get and dict.setdefault raises, which returns a default, and which one writes to the dict. This is a screening question and slowness here is itself the signal.

for a middle

Explain the mechanics: __getitem__ raising KeyError, the default argument being evaluated eagerly on every setdefault call, and why the object setdefault returns is the very one now stored in the dict.

for a senior

Show judgement about when a missing key should be loud. In an import path a silent get(field, 0) can bury an upstream schema change for weeks, so argue for KeyError or an explicit sentinel wherever absence means a defect rather than an option.

for a principal

Own the convention rather than the call. Pick one house style for missing-key handling - loud at boundaries, defaults only where absence is genuinely expected - so error behaviour is predictable across services instead of decided per author.

A Python dict offers three ways to reach a key, and interviewers ask about them because each one has a different failure mode. **Indexing — `d[key]`.** This calls the mapping's `__getitem__`. If the key is absent, `dict` raises `KeyError` with the key as its argument. That is the right behaviour whenever a missing key means a bug: a config value your service cannot start without, a column your import promised would be there. Loud is good. The idiomatic Python way to handle the miss is EAFP — wrap the access in `try` / `except KeyError` — rather than testing `key in d` first and then indexing, which hashes the key twice and leaves a race window whenever anything else can touch the dict between the two statements. **`dict.get(key)` and `dict.get(key, default)`.** A read-only lookup that never raises and never writes. With one argument it returns `None` for a missing key; with two it returns the second argument. The dict is untouched either way — a very common misunderstanding is that `get` with a default somehow stores that default. It does not. Two subtleties matter. First, the default is an ordinary expression, evaluated before the call, so `d.get(k, expensive())` pays for `expensive()` on every call regardless of whether the key exists. Second, `get` cannot distinguish "key absent" from "key present with the value `None`". When that distinction matters, use a private sentinel object and compare with `is`, or use `try` / `except KeyError`. **`dict.setdefault(key, default)`.** The one that both reads and writes. If the key is present, it returns the stored value and changes nothing. If the key is absent, it inserts `default` under that key and returns the object it just stored. The default argument itself defaults to `None`, so `d.setdefault(k)` inserts `None`. Because the returned object is the one now living in the dict, `setdefault` is the classic one-liner for accumulating groups: ```python groups = {} for name, dept in rows: groups.setdefault(dept, []).append(name) ``` Each miss creates a fresh list, stores it and hands it back; each hit hands back the list already there, and the `append` lands in the dict either way. The trap is the same eager evaluation as with `get`, and here it costs more, because people reach for `setdefault` precisely when the default is a freshly constructed container: `d.setdefault(k, [])` builds a new empty list on every iteration and throws it away on every hit. That is cheap for a list and expensive for anything that opens a file, hits a network or does real work — build such a default only after checking for the key. **Removal: `pop` and `popitem`.** `d.pop(key)` removes the key and returns its value, raising `KeyError` if it is absent; `d.pop(key, default)` returns the default instead of raising. `d.popitem()` takes no key at all: it removes and returns the most recently inserted `(key, value)` pair, and raises `KeyError('popitem(): dictionary is empty')` on an empty dict. That LIFO behaviour has been guaranteed since Python 3.7, when insertion order became part of the language; before then `popitem` removed an arbitrary pair. It makes `while d: k, v = d.popitem()` a safe way to drain a dict, since the dict is never iterated. **Bulk writes: `update` and `fromkeys`.** `d.update(other)` copies another mapping's entries into `d`, and also accepts an iterable of key/value pairs or keyword arguments; it mutates in place and returns `None`. `dict.fromkeys(keys)` builds a new dict mapping every key to `None`; `dict.fromkeys(keys, value)` maps every key to *the same object*, which is a genuine trap when that object is mutable — `dict.fromkeys(names, [])` gives every name one shared list, so appending through one key appears under all of them. Build per-key mutable values with a loop or a dict comprehension instead. **Choosing between them.** Use indexing when absence is an error you want to see. Use `get` when absence is normal and you only want to read. Use `setdefault` when absence should be repaired in place by inserting a container you are about to fill, and when the default is cheap to construct. A subclass of `dict` can also define the `__missing__` hook, which plain `dict` does not implement: `__getitem__` calls it on a miss, letting the subclass supply or insert a value instead of raising. The one habit to unlearn is defaulting everything. A pipeline that reads every field with `get(field, 0)` will happily process a file whose schema changed under it and produce plausible, wrong numbers for weeks. Reserve defaults for the fields that are genuinely optional.

  • When does `dict.pop` raise KeyError, and how is `dict.popitem` different?
    `d.pop(key)` removes the key and returns its value, raising `KeyError` if it is absent; passing a second argument returns that default instead. `d.popitem()` takes no key at all - it removes and returns the most recently inserted `(key, value)` pair, LIFO since Python 3.7, and raises `KeyError('popitem(): dictionary is empty')` on an empty dict. That makes `while d: k, v = d.popitem()` a safe way to drain a dict.
  • What goes wrong with `dict.fromkeys(names, [])`?
    The value expression is evaluated once, so every key maps to the *same* list object; appending through one key shows up under all of them. `dict.fromkeys(names)` with no value maps each key to `None`, which is safe because `None` is immutable. When the per-key value is mutable, build the dict with a loop or a comprehension so each key gets its own.
  • How do you tell a missing key apart from a key whose stored value is None?
    `dict.get` cannot: both return `None`. Either use `try: d[key] / except KeyError`, or pass a private sentinel - `MISSING = object()` and `if d.get(key, MISSING) is MISSING` - comparing with `is`. Testing `key in d` also works but costs a second hash lookup when you then read the value.

get is asking the front desk whether a room is booked under your name; setdefault is asking, and having the desk book it for you on the spot if it was free.

saying these in an interview costs you the question

  • Says `dict.get` raises KeyError on a missing key
  • Thinks `dict.setdefault` only reads and never writes
  • Believes setdefault's default is built only when the key is missing
  • Assumes `dict.get(key)` returning None proves the key is absent
  • Calls `key in d` followed by `d[key]` a single lookup
  • Thinks `dict.fromkeys(keys, [])` gives each key its own list

context

open as a page

What does `dict.keys()` return in Python 3, and how does that object behave as the dict changes?

level: middleimportance: must knowfreq 60%

basics

~20 s

A dict_keys view object, not a list. The view is a live window onto the dict, so keys added or removed later show through it. It supports len, iteration, fast membership and set operations, but not indexing.

open as a page

How does the `|` merge operator combine two Python dicts, and how does it differ from `dict.update`?

level: middleimportance: should knowfreq 45%

basics

~20 s

a | b returns a new dict and leaves both operands alone; it exists since Python 3.9. a.update(b) mutates a and returns None. Both let the right-hand side win on duplicate keys, and both copy one level deep.

open as a page

Why does a Python dict raise RuntimeError when a payroll import deletes stale keys mid-iteration, and how do you fix it?

level: seniorimportance: should knowfreq 42%

basics

~20 s

A dict iterator records the dict's size and re-checks it on every step, so deleting or inserting a key makes the next step raise RuntimeError. Iterate a snapshot such as list(d), or rebuild the dict with a comprehension, and mutate from there.

open as a page