skip to content

When do you need a custom secret scanning pattern in GitHub, and how do you roll one out?

level: seniorimportance: should knowfreq 35%

answer

  1. Built-in coverage is other people's formats
  2. Anchoring beats a clever expression
  3. Try it before you publish it
  4. Blocking pushes comes last, not first
  5. Prefixed tokens are easy to find

basics

~20 s

You need one when the credential is issued by you rather than a partner provider, so no built-in pattern matches it. Define it at repository, organisation or enterprise level, dry-run it against real repositories to measure false positives, then publish it and enable push protection for it.

solid answer

~50 s

GitHub's built-in coverage is **partner patterns** — formats published by the services that issue those credentials. Anything you issue yourself is invisible to it: internal service tokens, legacy database connection strings, signing keys, licence keys. A custom pattern closes that gap. You define it with a **secret format** regular expression plus optional **before secret** and **after secret** expressions that anchor the match to its surroundings, and additional match requirements to raise precision. The regex dialect is restricted — complex constructs are not supported — so keep patterns simple and anchored. Before publishing, run a **dry run** against representative repositories and read the hits: a noisy pattern is worse than none, because people learn to bypass. Once precision looks good, publish it, and only then enable **push protection for that pattern**, which is what actually stops the next leak. Scope matters: define at organisation level so every repository inherits it rather than per repository.

code

text · 10 lines
text
Pattern name:   Acme internal service token

Secret format:  acme_(live|test)_[A-Za-z0-9]{32}
Before secret:  \A|[^0-9A-Za-z]
After secret:   \z|[^0-9A-Za-z]

# Second pattern: an unprefixed password needs context to be precise
Secret format:  [A-Za-z0-9!@#$%^&*_-]{12,64}
Before secret:  (?i)(db_password|jdbc_password)\s*[:=]\s*["']?
After secret:   ["']?\s*(\n|\z)

go deeper

for a junior

Know that GitHub's detection covers credential formats published by partner providers, and that anything your own company issues needs a pattern someone defines.

for a middle

Be able to describe the pattern fields — secret format plus before and after context — and why anchoring on surrounding text is what makes an unprefixed secret detectable at all.

for a senior

Show the rollout discipline: samples, dry run, precision tuning, alert-only period, then push protection. Explain why a noisy pattern damages trust in the whole control.

for a principal

Push the question upstream: argue for prefixed, recognisable credential formats at issuance so detection and automated revocation become easy, and own where patterns are defined and reviewed across the organisation.

## Why built-in coverage is incomplete Secret scanning finds credentials whose format is *known*. That knowledge comes from the partner programme: providers publish the shape of their credentials so GitHub can match them, and in return get told when one leaks publicly. This works brilliantly for third-party credentials and not at all for yours. If your platform team issues tokens that look like a bare 32-character hex string, there is nothing distinctive to match and nothing published for GitHub to match against. The categories that most often need a custom pattern: - Internal service or API tokens minted by your own auth service. - Database connection strings that embed a password inline. - Private keys or certificates in a house format. - Credentials for a vendor small enough not to be in the partner programme. - Legacy credential formats that predate any convention. ## Anatomy of a custom pattern A custom pattern is more than one regex. You provide: - **Secret format** — the expression matching the credential itself. This is the part that will be reported. - **Before secret** — what must appear immediately before the match. Very useful for unprefixed secrets: requiring something like `password\s*=\s*` turns a hopeless generic match into a precise one. - **After secret** — what must follow, commonly a boundary such as end-of-line or a quote. - **Additional match requirements** — further constraints that must hold for a candidate to be reported. The supported regular-expression syntax is a restricted dialect optimised for scanning at scale; elaborate constructs are not available. Practically, this pushes you towards simple character classes, explicit lengths and good anchoring — which is what produces precise patterns anyway. ## The rollout sequence 1. **Write it against real examples.** Collect actual (already-rotated) samples of the credential and a set of near-misses you must not match. 2. **Dry run.** Custom patterns can be dry-run against selected repositories or the whole organisation before publishing. This is the step teams skip and regret: it tells you the true positive rate on your own codebase. 3. **Read the hits.** Look specifically at what it matched that is *not* a credential — test fixtures, sample documentation, base64 blobs, UUIDs. Tighten `before secret` and length constraints until the noise is gone. 4. **Publish for alerting only.** Let it run and generate alerts for a while. You now have a real backlog to triage, which is also a rough measure of how bad the problem was. 5. **Enable push protection for the pattern.** This is the payoff. Only do it once precision is high, because a pattern that blocks pushes wrongly trains an entire engineering organisation to reach for the bypass button — and that habit does not stay confined to your noisy pattern. ## Where to define it Patterns can be defined at repository, organisation or enterprise level. Define internal credential formats at the **organisation** level: the credential type does not belong to one repository, and per-repository definitions drift and go stale. A repository-level pattern is for something genuinely local, and even then it is usually a sign the credential format deserves organisation-wide attention. ## The upstream lesson The deepest version of this answer is that custom patterns are a workaround for credential formats that were not designed to be found. If you also control credential issuance, give your tokens a **distinctive, unambiguous prefix** — a short vendor-and-type marker followed by the random part. That single change makes the credential trivially matchable with a precise pattern, makes it recognisable in a log or a screenshot, and lets you build revocation tooling that can act on a found string. Interviewers notice when a candidate turns "how do I detect this" into "why is this hard to detect, and can I fix that instead". ## Interaction with exclusions If a pattern legitimately matches sample values in documentation or fixtures, the choices are: change the samples to be obviously fake, or exclude those paths from alerting via the repository's secret scanning configuration file. Prefer changing the samples — an exclusion is a permanent blind spot in a directory, and directories acquire new files.

  • Why not enable push protection for a new custom pattern immediately?
    Because a pattern with poor precision will block legitimate pushes, and the fastest resolution developers find is the bypass button. Once a team learns to bypass reflexively, the control loses value across every pattern, not just the noisy one. Run in alert-only mode until the false-positive rate is acceptable.
  • What makes a credential format easy or hard to detect?
    A distinctive fixed prefix plus a known length is trivially matchable with high precision. A bare random string of hex or base64 characters is indistinguishable from hashes, IDs and test data, so any pattern for it is either noisy or anchored on surrounding context such as an assignment to a password field.
  • Where should an internally issued token's pattern be defined?
    At organisation level, so every repository inherits it and there is one definition to maintain. The credential type belongs to the platform that issues it, not to whichever repository happened to leak it first. Repository-level patterns should be reserved for genuinely local formats.

saying these in an interview costs you the question

  • Writes a broad pattern and enables push protection immediately
  • Assumes built-in patterns cover internally issued tokens
  • Defines the same pattern separately in many repositories
  • Skips the dry run and triages the fallout later
  • Excludes whole directories rather than fixing fake-looking fixtures

context