skip to content

A credential scanner's entropy heuristic flags a 40-character random build identifier on every run — why does that value score like a secret?

level: juniorimportance: should knowfreq 50%

answer

  1. randomness is a property, not a meaning
  2. the string, never its origin
  3. length and character spread decide it
  4. generated identifiers score like keys
  5. suppress the exact string and path

basics

~20 s

An entropy heuristic measures randomness and length, not meaning: a 40-character hexadecimal build identifier has the same character spread as a freshly minted key, so it clears the threshold. Nothing in the string says what produced it.

solid answer

~40 s

A scanner is usually two detectors. One matches known formats — the prefixes, lengths and checksums issuers stamp on the credentials they mint. The other scores a run of characters for improbability: high spread, no word structure, above a chosen length. That second engine reads the string and nothing else. It cannot see what wrote the value, and it has no way to ask whether anything would accept it, so a content-derived build identifier, a random test fixture and a real key all look identical to it. The right fix is narrow: confirm what produces the value, then record a reviewed decision against that exact string and path. Raising the threshold or excluding the file to make the noise stop also removes the long random values that genuinely are credentials.

go deeper

for a junior

Recall that the entropy engine scores the characters themselves — length and spread — and knows nothing about where the value came from or whether anything accepts it.

for a middle

Explain the split between a known-format catalogue and a randomness score, and say why a content-derived identifier and a minted key are indistinguishable to the second one.

for a senior

Show the operating judgment: dismiss by recording a reviewed decision against the exact string and path, never by moving a global threshold, and say what each shortcut would have cost in future coverage.

for a principal

Frame it as an attention budget. The number a team sets for a threshold decides which error it will live with, and the standard you publish for dismissing a finding is what keeps the scan believed a year later.

## Two engines wearing one name A credential scanner is usually two detectors under one label, and they answer different questions. A **known-format** engine matches a candidate string against a catalogue of shapes that issuers deliberately stamp onto the credentials they mint: a fixed prefix, a fixed length, a restricted character set, sometimes a checksum built into the value itself. An **entropy heuristic** has no catalogue at all. It takes a run of characters and asks a statistical question — how improbable is this run, given its length and the spread of characters inside it? Prose scores low. File paths and identifiers a person typed score low. Anything drawn from a generator scores high. That is the whole mechanism, and every limitation follows from it: - it never sees what produced the value; - it never checks whether any system would accept the value; - it has no notion of *secret*, only of *improbable*; - to a character-frequency measure, key material, a content digest and a random identifier are the same object. So a flag on a build identifier is not a scanner defect. It is the heuristic doing exactly what it is specified to do, to a value that genuinely has the property it looks for. ## Why the identifier scores like key material A release or build identifier is normally derived from a content digest or drawn at random, and it is then written into a file that ships with every release. Forty hexadecimal characters are spread evenly across sixteen symbols with no word structure — which is precisely the shape of a value that was generated rather than chosen. A minted credential is generated too, usually over a wider alphabet. Measured on character distribution alone, the two are not separable. | Value found in the tree | Why it scores high | Is it a secret? | |---|---|---| | Build or release identifier | derived from a digest; characters spread evenly | no | | Test fixture generated at random | drawn from a generator | no | | An encoded binary blob or asset digest | encoding flattens the character distribution | no | | An issued access credential | drawn at random, by design | yes | | Exported key material in a configuration file | key material is random by construction | yes | The useful form of the answer in an interview is the row structure itself: the heuristic separates *random* from *not random*, and secrecy is not a sub-case of randomness in either direction. ## The cost of a flag that never goes away The identifier is not flagged once. It is flagged in the current file, and again in every earlier version of that file, on every branch that carries it — so one harmless value can account for a large share of a first scan's output. Reviewer attention is the scarce resource here, not machine time. A reviewer who has dismissed the same string a dozen times reads the thirteenth one less carefully, and the finding that actually matters arrives in that queue. ## Silencing it without going blind 1. Establish what writes the value, so the dismissal is a fact rather than a guess. 2. Record a reviewed decision keyed to **that exact string and that exact path**, with who decided and when. 3. Leave the engine and its threshold alone. 4. Re-check when the producer changes, because the same path may later hold a different value. The alternatives all cost more than they save: - **Raising the threshold globally** removes exactly the long, high-spread values that real credentials are. - **Excluding the file** hides anything added to that file afterwards, which is a file someone already ships. - **Excluding the repository** is the same move with a bigger blast radius. - **Rewriting the identifier** changes a build artefact to please a tool, and the next generated value looks just as random. ## What a low score does not mean The signal runs one way only. A high score says *this looks generated*; it does not say *this is a credential*. A low score says *this looks chosen or written*; it does not say *this is safe*. A database password a person typed scores low and is still a live credential, which is why an entropy engine alone is never the whole scan. Thresholds and minimum lengths are dials that teams and platforms set, and they trade noise against misses; there is no universally correct number, and any team that claims one has simply picked the error it prefers.

  • Why does the same harmless identifier account for hundreds of findings in a first scan?
    Because the scan reads history, not just the current files. The value appears in every version of the file that carried it, on every branch that carried that file, and each occurrence is a separate hit unless the tool groups by value.
  • What does raising the entropy threshold buy, and what does it cost?
    It buys quiet: fewer generated-looking non-secrets are reported. It costs the finds nearest the boundary, and real credentials sit there — a shorter issued value or one over a narrower alphabet scores lower than a long digest. The threshold moves the error, it does not remove it.
  • If the value really is meaningless, why not delete it from the file?
    Because it is a build artefact something depends on, and the next generated value looks just as random. Deleting it treats the scanner's output as the requirement. The reviewed decision is the record that the value was examined; the file stays as the build needs it.

A metal detector at a door beeps at belt buckles and coins. It is measuring metal, not intent, and turning its sensitivity down to stop the buckles is how you walk a blade through.

saying these in an interview costs you the question

  • Says the scanner is broken because it flagged a non-secret
  • Treats every high-entropy string as a credential by definition
  • Raises the global entropy threshold to silence one noisy file
  • Assumes a value the heuristic ignores is therefore safe to commit
  • Deletes the build identifier so the finding stops appearing