skip to content

How does uuid.uuid5 with a namespace make a retried 6,800-row billing batch idempotent?

level: seniorimportance: should knowfreq 38%

answer

  1. The rerun must reach the same id
  2. Name the work, not the attempt
  3. A pure function of stable inputs
  4. Namespace plus a canonical name string
  5. uuid5 recipe frozen once shipped

basics

~20 s

Derive each charge's identifier with uuid.uuid5(namespace, name) from stable inputs such as the run key and the subscription id. The retry recomputes the same UUID, so a unique constraint or an upsert rejects the second write instead of charging twice.

solid answer

~50 s

`uuid.uuid4()` mints a fresh id on every attempt, so a batch that dies halfway and is rerun inserts a second row for work it already did. `uuid.uuid5(namespace, name)` is a pure function of its inputs: pin a namespace UUID as a module constant and build the name from the fields that identify the unit of work — the billing-run key plus the subscription id — and the retry computes byte-for-byte the same UUID. Write that id as the row's key, and the duplicate write fails a uniqueness check or collapses into an upsert rather than producing a second side effect. The care goes into the name: normalise the inputs, use a separator that cannot appear inside them, and never change the recipe or the namespace afterwards, because a changed recipe makes yesterday's ids unreachable. Remember the id is derived, not secret — anyone knowing the inputs can compute it.

code

python · 12 lines
python
import uuid

NAMESPACE = uuid.uuid5(uuid.NAMESPACE_DNS, "billing.example")


def charge_id(run_id: str, subscription_id: str) -> uuid.UUID:
    return uuid.uuid5(NAMESPACE, f"{run_id}/{subscription_id}")


first = charge_id("2026-09-01", "sub-4171")
retry = charge_id("2026-09-01", "sub-4171")
print(first == retry, first.version, first == charge_id("2026-09-02", "sub-4171"))

go deeper

for a junior

Take away the core idea: uuid.uuid5 returns the same UUID for the same namespace and name, while uuid.uuid4 returns a new one every call. That difference is what lets a rerun recognise work it already did.

for a middle

Explain the mechanics of the fix — a pinned namespace constant, a canonical name built from stable business fields, and a uniqueness constraint on the resulting id so the second write is rejected rather than duplicated.

for a senior

Show you have operated this: normalising inputs, choosing a separator that cannot be forged, keeping the recipe frozen, passing the same id to an external system as its idempotency key, and knowing the derived value is not a secret.

for a principal

Own the wider call — derived ids against random ids plus declarative constraints, who owns the name recipe, how it is versioned across services, and what happens to years of existing identifiers if it ever changes.

A nightly billing run walks 6,800 subscriptions and, for each one, records a charge. It dies at row 4,102 and somebody reruns it. If the identifier on each charge came from `uuid.uuid4()`, every row the rerun touches gets a brand-new id, the storage layer sees 6,800 rows it has never seen before, and 4,101 customers are charged twice. Nothing in the code is wrong in the ordinary sense; the identifier simply carries no memory of the first attempt. ### The shape of the fix `uuid.uuid5(namespace, name)` is a deterministic function: sixteen bytes of namespace, the UTF-8 name, a SHA-1 digest, truncated and stamped with version 5. Feed it the identity of the *work item* rather than the identity of the *attempt* and the id stops moving: ```python import uuid CHARGE_NS = uuid.UUID("3f2a6f5c-6a1b-4d7e-9c2b-8f1c9a0e4d55") def charge_id(run_key: str, subscription_id: str) -> uuid.UUID: return uuid.uuid5(CHARGE_NS, f"{run_key}|{subscription_id}") ``` Now the first attempt and the rerun both compute the same UUID for subscription `sub-4171` in run `2026-09-01`. Make that id the row's primary key, or put a uniqueness constraint on it, and the second write is refused or folded into the existing row. The duplicated side effect becomes a no-op instead of a second charge, and it becomes one *at the storage layer*, which is the only place that can see both attempts. The pattern is worth naming precisely: this is not deduplication after the fact, it is giving the operation a name that both attempts can independently derive. Any external system that also accepts an idempotency key can be handed the same value, so the guarantee extends past your own database. ### Where it goes wrong **The name recipe is a contract.** The instant you ship it, every id you have ever written depends on it. Adding a field, reordering the parts, switching the separator or changing the namespace constant produces a disjoint id space, and yesterday's rows become unreachable by tomorrow's code. Treat the recipe and the namespace as frozen, and version them explicitly (a `v2|` prefix inside the name) if they truly must change. **Separators can be forged.** `f"{a}{b}"` maps `("ab", "c")` and `("a", "bc")` to the same string and therefore the same UUID. Pick a separator that cannot appear in any component, or normalise and length-prefix the parts. This is the same collision hazard as any concatenated key. **Normalise the inputs.** `"SUB-4171"` and `"sub-4171"` are different names and produce unrelated ids. Case, whitespace and any date formatting have to be canonicalised before hashing, and the run key must be something the rerun genuinely reproduces — a business date or a run identifier, never `datetime.datetime.now()` or a fresh `uuid.uuid4()` allocated at start-up. **Derived is not secret.** A version 5 UUID is a digest of inputs, and truncation to 128 bits is not a security property. Anyone who can guess the run key and the subscription id computes the id, so it must not double as a bearer token or an unguessable capability. **Partial failures still need care.** A derived id makes the *write* idempotent; it does not make an external side effect idempotent unless that system honours the key too. If the run calls a payment provider, the same UUID should travel as the provider's idempotency key, and the local row should be written in the same transaction as anything else that must not repeat. ### The alternative, and when to prefer it The other way to get the same guarantee is a random `uuid4` id plus a separate natural-key uniqueness constraint on `(run_key, subscription_id)`. That is often better: the constraint states the invariant declaratively, and the id stays opaque. Derived ids earn their place when the identifier must be computed by more than one component without a round trip — two services, or a client and a server, that need to agree on a row's id before either has written it. ```python import uuid NAMESPACE = uuid.uuid5(uuid.NAMESPACE_DNS, "billing.example") def charge_id(run_id: str, subscription_id: str) -> uuid.UUID: return uuid.uuid5(NAMESPACE, f"{run_id}/{subscription_id}") print(charge_id("2026-09-01", "sub-4171") == charge_id("2026-09-01", "sub-4171")) print(charge_id("2026-09-01", "sub-4171") == charge_id("2026-09-02", "sub-4171")) ``` That prints `True` then `False`: the same work item, and a different one.

  • What breaks if someone later adds a field to the string you pass to uuid.uuid5?
    Every identifier changes. The function is a digest, so a new name produces an unrelated UUID and the rerun no longer collides with the rows the old code wrote — the duplicate protection silently disappears, and old rows become unaddressable by new code. Treat the name recipe and the namespace constant as a frozen contract, and if it must evolve, encode a version marker inside the name so both generations remain derivable.
  • Would a random uuid4 key with a uniqueness constraint on the business fields solve the same problem?
    Usually yes, and often better: the constraint states the invariant declaratively and the identifier stays opaque and unguessable. The derived id wins when two components must agree on the identifier without talking to each other first — a client that needs the id before the write lands, or a second service that computes the same id independently. Otherwise prefer the explicit constraint.
  • Can a uuid5 identifier double as the token in a URL that must not be guessable?
    No. It is a digest of inputs you have already told somebody, so a caller who knows the run key and the subscription id derives the same value without any access to your system. Truncating a digest to 128 bits is not a secrecy mechanism. If a value must be unguessable, generate it from the random source instead and treat it as a credential.

saying these in an interview costs you the question

  • Uses uuid4 per attempt and calls the run idempotent
  • Builds the uuid5 name from the current timestamp
  • Concatenates fields with no separator, allowing collisions
  • Treats a derived uuid5 value as unguessable or secret
  • Changes the namespace constant without migrating existing ids
  • Relies on a check-then-insert race instead of a uniqueness constraint

context