skip to content

How should you read an OpenSSF Scorecard result when vetting a library for adoption?

level: middleimportance: nice to knowfreq 38%

answer

  1. checks, not the composite number
  2. process evidence, not code correctness
  3. small project, low score, little meaning
  4. mechanical checks are cheap to game
  5. input to a review, not a merge gate

basics

~20 s

Read the individual checks, not the composite number. Checks such as code review, branch protection, pinned dependencies and a published security policy are evidence about how a project is run — not about whether its code is correct.

solid answer

~50 s

OpenSSF Scorecard runs automated checks against a repository, scores each one, and rolls them into a single weighted number. The number is the least useful part. I read the check families instead: maintenance (is it still maintained, are there unfixed known vulnerabilities), review process (code review, branch protection), build hygiene (pinned dependencies, workflow token permissions, unreviewable binary artifacts in the repo), and disclosure (a security policy, signed releases). Each is evidence about process discipline, not about code correctness. Context sets the weight: a two-file utility with no CI can score badly and still be lower risk than a well-scored framework I will run with production credentials. And because the checks are mechanical and the result is public, some are cheap to satisfy without changing behaviour. I use Scorecard to generate the questions a human review should ask, never as a pass/fail gate on its own.

go deeper

for a junior

Know that this is an automated assessment of a project's development and release practices, that it reports individual named checks alongside one overall number, and that it does not read the code for bugs.

for a middle

Be ready to name real check families — code review, branch protection, pinned dependencies, security policy, unfixed known vulnerabilities — and to explain what threat each one speaks to.

for a senior

Demonstrate that you weight checks by what the dependency will be allowed to reach, and that you can explain why a low score on a tiny library is usually noise rather than a finding.

for a principal

Be able to argue against a naive score threshold as an organisational gate: it fails small healthy projects, rewards cheap settings changes, and moves the debate from risk onto the tool.

## What Scorecard is OpenSSF Scorecard is an automated assessment that inspects a project's repository and release practices and reports a set of named checks, each scored on a 0-10 scale, plus a weighted aggregate. Checks are grouped by the risk they speak to, and each carries its own risk weighting, so the aggregate is not a plain average. It is a *process* measurement performed from the outside, entirely without reading the project's logic for correctness. ## The check families and what they tell you **Is anyone maintaining it.** A maintenance check looks at recent commit and release activity; a vulnerabilities check looks for known, unfixed advisories affecting the project itself. These are the closest thing to "is this alive and is it currently carrying a known problem". **How changes get in.** Code review and branch protection checks ask whether commits reach the default branch without a second pair of eyes. This is the single most decision-relevant family for supply chain risk: if anything can be merged and released unilaterally, then compromising one account compromises the artifact, and no amount of downstream scanning changes that. **Build hygiene.** Pinned dependencies, restrictive workflow token permissions, absence of dangerous automation patterns, and absence of committed binary artifacts. A binary blob checked into a source repository is code you cannot review and cannot reproduce, which is why its presence is a genuine finding rather than an aesthetic one. **Disclosure and release integrity.** Whether a security policy exists, and whether releases are signed. The first tells you where to send a report; the second matters only if you actually intend to verify something on the way in. **Testing rigour.** Whether CI tests run on pull requests, whether static analysis runs, whether the project is fuzzed. Fuzzing in particular is meaningful for anything that parses input. ## Why the composite number is the weakest output Three reasons. *It has no notion of your blast radius.* The same score means different things for a two-file string utility and for a framework that terminates TLS in your service. Scorecard cannot know which one you are adopting or what it will be allowed to touch. *It penalises smallness and unusual hosting.* Many checks look for infrastructure that a tiny, stable project has no reason to have, and several depend on the project living somewhere the checks can inspect. A low score frequently means "there was nothing to measure", not "this is dangerous". *Mechanical checks are gameable.* Because the checks are public and deterministic, a project can raise several of them by flipping settings — adding a file, changing a repository option — without changing how it actually reviews or releases code. The score improves; the behaviour does not. Any metric that becomes a target stops being a measurement. ## Using it well Treat Scorecard as a question generator for a human review. A failing code-review check turns into "who reviews changes here, really?". A failing pinned-dependencies check turns into "can an upstream change of theirs alter what I receive?". A failing security-policy check turns into "where would I send a report?". Answers to those questions, weighted by what the dependency will be allowed to reach in your system, are the vetting decision. The common organisational mistake is turning the aggregate into a merge gate — for example, refusing anything under some threshold. That gate fails small, healthy libraries; passes projects that gamed the cheap checks; provides no way to express that this particular dependency will hold keys; and shifts every conversation into arguing with the tool rather than about the risk. If you want a mechanical rule, apply it to a *single* check whose failure genuinely maps to a threat you care about, and always with a documented exception path and an owner. ## What it never gives you Scorecard does not audit the code, does not vet the humans, does not tell you whether the published artifact matches the source you inspected, and does not tell you whether the library is a good fit for your problem. It measures the shape of the project's process. That is genuinely useful and genuinely limited, and a candidate who can say exactly where the line falls is showing more judgment than one who quotes a number.

  • Which individual checks would actually change your adoption decision?
    Code review and branch protection, because they decide whether one compromised account can publish arbitrary code. Unfixed known vulnerabilities in the project itself. Binary artifacts committed to the repository, which are unreviewable. Over-broad automation token permissions, because a hostile pull request could then reach publishing credentials. Signed releases matter to me only if I intend to verify them on the way in.
  • Why not simply require a minimum aggregate score before a dependency can be merged?
    Because it fails the wrong things and passes the wrong things. Small, stable libraries score low for lacking infrastructure they do not need; projects that flip a few settings score well without changing behaviour; and the number cannot express that this particular library will handle keys. If you gate, gate on a specific check tied to a specific threat, with an exception path and a named owner.
  • How do you vet a package whose source is not somewhere the checks can inspect?
    Ask the same families by hand. Who reviews changes, and is there evidence of it in the history? Do tagged sources correspond to published artifacts? Is there a disclosure address and any record of a fix turnaround? Are there unreviewable binaries in the tree? For a commercial supplier the same questions go to the vendor, in writing.

saying these in an interview costs you the question

  • Treats the composite score as a pass or fail verdict
  • Assumes a high score means the code was audited
  • Says a low score always means the project is unsafe
  • Gates every dependency on one number regardless of impact
  • Cannot name a single check or what threat it maps to

context