Why can `if rec_id not in seen: seen.add(rec_id)` let two threads emit the same catalogue record twice?
answer
- The container never got corrupted
- Two threads, one gap, both pass
- Idempotent insert hides the evidence
- Claim under the lock, work outside it
- Shard ownership or upsert downstream
basics
~20 sBecause the membership test and the insert are two separate set operations. Both threads can test before either adds, both find the id absent, and both go on to emit the record. The set is fine; the claim was never exclusive.
solid answer
~50 sThis is check-then-act, the compound-operation bug wearing a dedupe costume. `not in` and `add` are each internally consistent, but between them another thread can run the same test on the same id, so two workers both decide the record is new and the downstream side effect happens twice. Notice what did *not* break: the set is not corrupted and holds each id exactly once — `set.add` is idempotent — which is why the corruption you would look for in a debugger is not there and the symptom shows up only downstream. The fix is to make the claim atomic: hold one `threading.Lock` across the test and the insert, release it before the slow work, and let only the thread that won the claim emit. Sharding ids across workers so one owner exists per id, or making the emit idempotent, are the alternatives.
code
python · 19 linesimport threading
from concurrent.futures import ThreadPoolExecutor
seen = set()
claim = threading.Lock()
emitted = []
def import_record(rec_id):
with claim: # test and insert are one critical section
if rec_id in seen:
return
seen.add(rec_id)
emitted.append(rec_id) # slow work happens outside the lock
ids = [i % 500 for i in range(8000)]
with ThreadPoolExecutor(max_workers=8) as pool:
list(pool.map(import_record, ids))
print(len(emitted)) # 500, every rungo deeper
Recognise the shape: a membership test and an insert are two steps, so two threads can both pass the test. Naming threading.Lock as the fix is the expectation here.
Explain why the evidence is missing — the set stays perfectly consistent because the insert is idempotent, so the only symptom is a duplicate side effect downstream. Put the lock around both operations.
Show the production reasoning: reproduce with contention rather than volume, keep the lock to the claim and off the slow write, and separate in-process exclusivity from retry safety after a long job fails partway.
Own the design choice — partitioning ids so no sharing exists at all versus locking a shared claim, and the standing requirement that anything a batch job writes is idempotent, since duplicate delivery will happen for reasons unrelated to threads.
### The scenario A museum-catalogue importer runs nightly for about six hours, fanning records out across a thread pool. Some catalogue entries reach the downstream store twice. The counts are small, they differ every night, and the duplicated ids look unremarkable. Somebody points at the dedupe: ```python if rec_id not in seen: seen.add(rec_id) emit(record) ``` and the room splits between "sets are thread-safe" and "then how is it duplicating?". Both halves are right about their own half. ### Why the set being safe is not the point `rec_id not in seen` is one set operation. `seen.add(rec_id)` is another. Each is internally consistent — under the GIL because a C-implemented set method ran to completion between bytecode boundaries, and on the free-threaded build (supported from 3.14, PEP 779) because the interpreter takes an internal per-object lock for the duration of the operation. Neither build lets you corrupt the set. But the invariant you actually care about is "exactly one thread ever passes this branch for a given id", and that invariant spans both operations. Thread A tests: absent. Thread B tests: absent. A adds, B adds — and because `set.add` is idempotent, the set now looks perfect, containing the id exactly once. Both threads then call `emit`. The data structure is pristine and the side effect fired twice. That asymmetry is why this bug survives code review: there is no corrupted state to find afterwards, only a duplicate in a system you do not own. ### Diagnosing it The fingerprint is characteristic. The duplication rate scales with worker count and disappears at one worker. It is nondeterministic run to run, and it got noticeably worse when the job moved to a free-threaded interpreter — not because free-threading broke anything, but because threads that used to take turns now genuinely overlap, so the microseconds between the test and the add became a window two workers routinely occupy at once. Reproduce it by hammering the same id from many threads in a tight loop rather than by replaying the six-hour input; you want contention, not volume. Then read the code by counting shared-object operations per invariant, which finds every sibling of this bug in the same pass: read-then-write counters, `get`-then-`setdefault`, "is the file already open?" caches, lazily-built singletons. ### Fixing it **Make the claim atomic.** One `threading.Lock` around the test and the insert, and nothing else. The thread that inserts has won the record; every other thread returns. Crucially, release the lock before the emit: the claim already guarantees exclusivity, and holding a lock across the slow downstream write would serialize the entire pool on the slowest step and hand back the parallelism you paid for. This is the answer to give first. **Or make the claim a single operation.** `dict.setdefault(rec_id, marker)` performs the lookup and the conditional insert as one container operation and returns whichever value is now stored, so a thread can tell whether it was the one who inserted. It is a real option, but say the caveat: CPython publishes no per-method atomicity table, so lock-free claims should be reserved for cases you can point at in the documentation, not inferred from how a method "probably" works. **Or remove the sharing.** Shard ids across workers by hash so each id has exactly one owner and the `seen` set is per-worker and unshared. No lock, no contention, and the invariant is enforced by the partition rather than by discipline. This is usually the strongest answer for a batch importer, where the input can be partitioned up front. **And regardless: make the downstream idempotent.** A six-hour job that crashes at hour five will be re-run, and a retry duplicates records for reasons that have nothing to do with threads. An upsert keyed on the catalogue id makes duplicate emission harmless, which turns a correctness bug into a wasted write. Locks protect one process's in-memory invariant; idempotence protects the outcome. A senior answer names both and does not confuse one for the other. ### The general shape Check-then-act, read-modify-write and test-then-insert are one bug with three names: an invariant that spans more than one operation on shared state, with nothing holding the two together. The per-object guarantee in the free-threaded build is exactly one operation wide, and it always was.
- Why should the downstream write happen outside the lock rather than inside it?Once a thread has inserted the id it owns that record exclusively, so no other thread can duplicate the work — the lock has already done its job. Holding it across the write would serialize every worker on the slowest step in the pipeline and give back the parallelism the thread pool exists for.
- Could `dict.setdefault` replace the lock here?It can: it performs the lookup and the conditional insert as one container operation and returns the value now stored, so the caller can tell whether it inserted. Use it deliberately, not by inference — CPython publishes no per-method atomicity table, so a lock is the answer whenever you would be guessing.
- The job is re-run after a crash at hour five. Does the lock help then?Not at all. The lock protects one process's in-memory set, which is empty on the next run, so a retry re-emits everything already written. Duplicate protection across runs has to live downstream — an upsert keyed on the catalogue id, or a durable record of what was committed.
Two volunteers each check the sign-up sheet, see nobody has taken the Tuesday shift, and both write their name. The sheet is legible and correct; two people still show up.
saying these in an interview costs you the question
- Insists sets are thread-safe so the code is correct
- Blames set.add for dropping or corrupting entries
- Holds the lock across the slow downstream write
- Says free-threading introduced the bug
- Adds a second lock instead of widening the first
- Treats an in-process lock as protection against job retries