skip to content

A scanner runs both a known-format catalogue and an entropy heuristic — why does a human-chosen database password clear both?

level: middleimportance: must knowfreq 58%

answer

  1. two engines, one shared blind spot
  2. no issuer stamped a shape on it
  3. short values fail the length guard
  4. a per-character score needs length
  5. the setting's name is the third signal

basics

~20 s

Neither engine has anything to work with: no issuer stamped a recognisable prefix, length or checksum on a password a person typed, and at under twenty characters it sits below the minimum length an entropy score needs before it means anything.

solid answer

~50 s

A known-format engine finds credentials that an issuer deliberately made findable — a fixed prefix, a known length, sometimes a checksum inside the value. Nobody stamps anything on a password a person chose, so there is no shape to match. An entropy engine scores improbability, but it only scores strings above a minimum length, because a per-character score computed over a short run of text is unstable and, without that guard, the run fills with ordinary words and paths. A chosen password is short by construction and falls under the guard. The consequence is a specific blind class: account passwords, key passphrases and credentials embedded in connection strings — which happen to be the long-lived, widely shared ones. Catching them takes a third signal: the name of the setting the value is assigned to, read together with the value.

code

pseudocode · 18 lines
pseudocode
knownFormats   = [ issuerPrefix + fixedLength + checksum, ... ]   # shapes issuers stamp on
minLength      = 20                                               # this team's guard
entropyDial    = 4.0                                              # this team's bits-per-character dial

function scanValue(text, settingName):
    if matchesAny(knownFormats, text):
        report(text, reason = "known format")
        return

    if length(text) >= minLength and entropyBitsPerChar(text) >= entropyDial:
        report(text, reason = "high entropy")
        return

    # "Wint3r-Search-2026!" is 19 characters: it fails the length guard,
    # so the entropy test never runs on it, and no issuer stamped a shape on it either.

    if settingNameSuggestsSecret(settingName) and isLiteralValue(text):
        report(text, reason = "secret-shaped setting with a literal value")

go deeper

for a junior

Recall that a scanner looks for shapes it knows and for values that look generated, and that a password a person chose looks like neither of those.

for a middle

Explain both guards in the same answer: no issuer stamp for the catalogue to match, and a minimum length that the entropy engine applies before it scores anything.

for a senior

Point at the class this leaves uncovered — shared account passwords and key passphrases, the longest-lived credentials in the estate — and describe the context signal you would add and the review path it needs.

for a principal

The lever is upstream: where your organisation issues its own credentials, giving them a recognisable stamped shape makes every future scan cheaper, and that is a decision about issuance rather than about scanning.

## Two engines and one shared blind spot A credential scanner is usually two detectors under one name, and each keys on something different. | Engine | Keys on | Catches | Misses | |---|---|---|---| | Known-format catalogue | issuer-stamped shape: prefix, length, alphabet, checksum | credentials minted by systems that chose to make them recognisable | anything nobody stamped a shape onto | | Entropy heuristic | improbability of the character run, above a minimum length | long generated values with no word structure | short values, and values with word structure | | Context signals | the name the value is assigned to, and the file it sits in | chosen passwords and pasted values in recognisable settings | values in fields whose names say nothing | The first two engines are what most teams enable, and a password a person chose falls between them in a way that is not an accident of tuning. ## Why the format engine has nothing to match Findable credential formats exist because issuers decided to make them findable: a distinctive prefix, a fixed length, an alphabet, often a checksum so a scanner can reject a lookalike without calling anything. That is a property of the **issuing system**, not of credentials in general. A database account password has no issuer in that sense. Someone chose it, typed it into a form, and pasted it into a configuration file. There is no prefix, no fixed length, no alphabet and no checksum — nothing a catalogue can key on. The same is true of a key passphrase, of a shared service account password, and of a username and password pair embedded in a connection string. ## Why the entropy engine skips it The heuristic scores a run of characters for improbability, but almost every implementation applies a **minimum length** first, for two reasons: - A per-character score computed over a short run is unstable. A short string with mostly distinct characters scores high simply because it is short, not because it was generated. - Without the guard, the run fills with identifiers, short paths and ordinary words, and the output stops being readable. A password a person chose is short by construction — that is what makes it typeable — so it usually fails the length guard before its score is even considered. Where it does clear the guard, word structure and a narrow character set pull the score down towards prose. Either way the engine is silent, and it is silent for exactly the values people are most likely to have pasted somewhere. ## The class this makes invisible The blind spot is not random; it names a category: - database and message-broker account passwords; - passphrases protecting exported key material; - shared service account passwords used by several systems; - a username and password pair embedded in a connection string; - anything a person chose, typed and reused. Those are, as a group, the **long-lived** credentials: rarely replaced, shared across consumers, and often the ones with the widest reach. The detection method is weakest precisely where the exposure lasts longest. ## The third signal: read the name, not just the value What separates a chosen password from a random word in a file is rarely the value. It is the **context**: a setting whose name announces what it holds, assigned a literal value. - A name containing *password*, *secret*, *credential*, *token* or *key*, assigned a literal rather than a reference to somewhere else. - A connection string shape, where a credential sits in a known position inside the text. - The kind of file: a configuration or deployment description, rather than documentation or a test fixture. - Proximity: a host name and an account name in the lines above. Context signals are noisier than a format catalogue — placeholders, examples and documentation trip them constantly — which is why they need their own review path rather than being thrown into the same queue. But they are the only one of the three that can see a value that was chosen rather than generated. ## What a clean report is worth A scan that reports nothing has established that nothing matched, with the engines that ran, over the surface they read. It has not established that the file contains no credential, and the gap is a predictable shape rather than bad luck. An engineer who can name that shape — *no issuer stamp, under the length guard, no context rule enabled* — is describing a scanner accurately; one who reads a clean report as proof of absence is not.

  • Why do entropy engines apply a minimum length before scoring at all?
    Because a per-character score over a short run is unstable — a short string of mostly distinct characters scores high just for being short — and because without the guard the output fills with identifiers, short paths and words. The guard buys a readable run, and costs every short credential.
  • What does the context signal cost that the format catalogue does not?
    Precision. A name containing `password` next to a literal matches placeholders, documentation examples and test fixtures constantly. It earns its place, but it needs its own review path and its own accepted false-positive rate rather than sharing a queue with high-confidence format hits.
  • Does the same blind spot apply when scanning history rather than current files?
    Yes — history changes the surface, not the detection method. The same chosen password is equally invisible in a version from three years ago. Widening where you look does not widen what the engines can recognise.

saying these in an interview costs you the question

  • Reads a clean scan as proof the file holds no credential
  • Believes every credential is long and randomly generated
  • Says lowering the entropy threshold would catch it for free
  • Treats the two engines as one detector with one setting
  • Assumes a password a person typed is not worth finding