You are minting long-lived ingest keys for four hundred partner stations — how many random bits, what leading segment, and which one-way function at rest?
answer
- size it from a guessing budget
- 128 bits, cryptographic generator
- a marker a scanner can match
- non-secret id makes lookup indexed
- verifier cost matches credential entropy
basics
~20 sAt least 128 bits from a cryptographically secure generator; a fixed, recognisable leading segment plus a non-secret key identifier; and a fast cryptographic digest such as SHA-256 as the verifier. A deliberately slow password-hashing function is the wrong tool for a high-entropy machine key.
solid answer
~40 sThree decisions. **Size**: at least 128 bits of randomness from a cryptographically secure generator, encoded in a URL-safe alphabet — roughly 22 characters. Argue it from a guessing budget, not from taste: 128 bits is beyond any online or offline search. **Shape**: `<fixed marker>_<environment>_<key id>_<secret>`. The fixed marker is what lets automated leak scanners recognise your key in a public repository; the environment segment stops a test key reaching the live feed; the key identifier is non-secret and makes verification one indexed read. **Verifier**: a fast digest, optionally keyed with a server-held pepper. The slow, memory-hard function you would use for a human password exists to make a *small* search space expensive; a 128-bit random key has no small search space, and the slowness would land on every single ingest request instead.
code
pseudocode · 19 lines# mint
secret = base62(csprng_bytes(16)) # 128 bits
key_id = base62(csprng_bytes(4)) # non-secret handle
plain = "seis_live_" + key_id + "_" + secret
store(key_id = key_id,
verifier = hmac_sha256(pepper, secret),
display = "seis_live_" + key_id + "..." + last4(secret),
station = station_id,
scope = requested_scope)
return plain # shown once, never stored
# verify
marker, env, key_id, secret = split(presented)
if env != THIS_ENVIRONMENT: return 401
row = lookup_by_key_id(key_id) # one indexed read
if row is null: return 401
if not constant_time_equals(hmac_sha256(pepper, secret), row.verifier): return 401
if row.revoked_at is not null: return 401
return rowgo deeper
Remember three answers: the randomness comes from the cryptographic generator, at least 128 bits, and the stored value is a one-way digest rather than the key.
Pick the numbers and say what they trade against — 128 bits argued from a guessing budget, the four jobs the leading segment does, and the split between the non-secret identifier and the secret.
Make the verifier argument out loud: the slow function's work factor buys nothing against uniform 128-bit entropy and costs latency on every ingest request, while remaining exactly right for human passwords.
Treat the key format as a long-lived interface. The marker you choose gets embedded in partner scripts, scanner rules and support tooling, and changing it later is a fleet-wide migration.
## Sizing the secret The only defence a bearer secret has against guessing is that there is nothing to guess. So size it from a budget rather than a feeling. Suppose an attacker can push ten thousand candidate keys per second at your ingest endpoint and you never notice. Against a 32-bit secret they are through the space in a few days. Against 64 bits they are not, but 64 bits is a *floor* argued for values an attacker cannot collect offline, and a key that sits in a partner's configuration file for three years is exactly the kind of value that leaks in bulk. **128 bits is the working answer**: sixteen bytes from a cryptographically secure generator, encoded in a URL-safe alphabet, about 22 characters of secret. Two failure modes to name explicitly: - **The wrong generator.** A general-purpose pseudo-random function seeded from the clock produces strings that look random and are predictable to anyone who can guess the seed. The generator must be the cryptographic one. - **Confusing length with entropy.** A forty-character key built from a station name, a timestamp and eight random characters has eight characters of entropy, not forty. What an attacker searches is the random part. ## The leading segment, and what it actually buys A key that is a bare blob of base62 is indistinguishable from every other blob of base62 in the world. Giving it a fixed, recognisable leading marker buys four concrete things: 1. **Automated leak scanning works.** A recognisable pattern is what lets scanning tooling — yours, the code host's, a researcher's — spot your credential in a public repository, a build log or a container image layer and tell you. A pattern nobody can write a matcher for is a credential nobody will ever report to you. 2. **The console can show a stub.** `seis_live_7f3a91c4…4KQ2` is recognisable to an operator with three keys and reveals nothing. 3. **Environment mistakes fail loudly.** A `test` marker rejected by the live ingest endpoint turns 'the technician pasted the wrong key' from a silent data-quality incident into an immediate `401`. 4. **Verification is one indexed read.** A non-secret key identifier inside the string means you look the row up directly. Without it, verifying means digesting the candidate against every stored row. So the shape is something like `seis_live_7f3a91c4_<22 random characters>`: marker, environment, identifier, secret. ## The verifier at rest, and the rule that inverts here Everybody has learned 'never store a credential in plain text; use a slow, memory-hard password-hashing function'. Half of that carries over and half of it does not, and an interviewer is watching for whether you know which half. | | Human password | Long-lived machine key | |---|---|---| | Entropy | Perhaps 20-30 bits in practice | 128 bits, uniform, machine-generated | | Offline search after a table dump | Feasible; the space is small | Infeasible; there is no space to search | | Presented how often | Once per login | On every ingest request, thousands per second | | Latency budget per verification | 100ms is fine | 100ms is a catastrophe | | Right function | Deliberately slow and memory-hard | A fast cryptographic digest, e.g. SHA-256 | The slow function exists to buy *work factor* against a small search space. A 128-bit uniformly random key has no small search space, so the work factor buys nothing — while the cost is real and lands on the hot path. At 3,000 ingest requests per second, a 100ms verification is 300 cores of pure overhead for zero security gain, and it hands anyone who can send requests a cheap way to exhaust your capacity. Two refinements worth stating: - **A pepper.** Digesting with a server-held secret — an `HMAC` construction — means a stolen table alone is not even theoretically attackable, because the attacker lacks the pepper. It costs you a secret to manage and to rotate. - **Constant-time comparison.** Compare the computed digest against the stored one with a constant-time comparison. (*Why* that matters is general cryptographic hygiene and is taught on its own.) And do not let the lesson invert: for a **human-chosen** password the slow memory-hard function remains exactly right. The rule is not 'fast digests are fine for credentials'; it is 'the verifier's cost must be matched to the credential's entropy'. ## What you must not do - Store the key reversibly encrypted so it can be displayed again. - Derive the key deterministically from the station identifier and a service secret — one leak of that secret regenerates every key in the fleet. - Let the full credential reach an access log, an error report or a crash dump. - Reuse one key across the live and test endpoints because 'it is the same partner'.
- Why not just use a random identifier format that is already standard, instead of inventing a key shape?You can, for the random part — a standard random identifier is 122 bits of randomness, which is close enough. What it will not give you is the fixed marker a leak scanner can match, the environment segment, or a separate non-secret lookup handle. Those three are the reason to define your own shape around the randomness.
- What does the pepper cost you, and when would you skip it?It is one more secret to store outside the database, distribute to every verifying node and rotate — and rotating it means re-digesting every key row. Skip it when your keys are genuinely 128-bit random, because a plain digest of a 128-bit secret is already not attackable offline. Take it when you want a stolen table to be worthless even if key entropy later turns out to be lower than you believed.
- A teammate proposes using the slow password-hashing function anyway, 'to be safe'. What is your argument?Ask what attack it prevents. Against a 128-bit random key there is no offline search to slow down, so the work factor protects nothing — while the cost is paid on every request by every station. It also converts your ingest endpoint into an amplifier: anyone who can send unauthenticated requests makes your servers do expensive work. Safety here is entropy, not latency.
saying these in an interview costs you the question
- Builds the key from a station name and a timestamp with a little randomness appended
- Says a forty-character key is strong without asking how much of it is random
- Insists a slow memory-hard function is required for a 128-bit machine key
- Concludes from this that human passwords can be stored with a fast digest
- Generates the secret with a general-purpose pseudo-random function seeded from the clock
- Omits any recognisable marker, leaving leaked keys unmatched by any scanner