Your CI worker image's SBOM lists OS packages but no Ruby gems - how do you diagnose it?
answer
- two halves, two analyzers
- success exit, empty section
- evidence, not packages, went missing
- check the stage the lockfile lived in
- assert a non-zero count per ecosystem
basics
~20 sThe two halves come from different analyzers. The OS half read the distribution package database; the gem half found no evidence it recognises and emitted nothing without failing. Find where the gems really live, and generate there.
solid answer
~50 sTreat it as a coverage bug, not a tool bug. An SBOM generator runs independent analyzers: OS packages come from the `dpkg`/`apk`/`rpm` database, gems come from installed `.gemspec` files or a lockfile. If the gems went into a vendored path the analyzer never walked, or the lockfile was left behind in an earlier build stage, the gem analyzer finds nothing and the run still exits successfully - an empty ecosystem is indistinguishable from an absent one in the document. So: confirm which digest was scanned, look inside the image for where the gems really live, and read the tool's debug output to see which analyzers fired. Then fix it at the source by generating in the stage where the manifest and gemspecs exist, and fail the build when an expected ecosystem yields zero components. On a CI worker that holds deploy credentials, an inventory you cannot trust means a blast radius you cannot bound.
code
json · 12 lines{
"bomFormat": "CycloneDX",
"components": [
{ "type": "library", "name": "openssl",
"version": "3.0.11-1~deb12u2",
"purl": "pkg:deb/debian/[email protected]~deb12u2?arch=amd64" },
{ "type": "library", "name": "zlib1g",
"version": "1:1.2.13.dfsg-1",
"purl": "pkg:deb/debian/zlib1g@1:1.2.13.dfsg-1?arch=amd64" }
]
// ... 180 more pkg:deb components, no pkg:gem component anywhere
}go deeper
Recognise that an SBOM's language packages and OS packages come from different places, and that an empty section means the tool found no evidence rather than proving nothing is installed.
Explain the evidence each analyzer needs and name the common reasons it goes missing: vendored paths, version-manager directories, a multi-stage build that dropped the lockfile.
Walk a diagnosis in order - confirm the subject digest, look inside the image, read which analyzers fired, re-scope the run - and then fix generation at the stage where the evidence lives.
Argue for completeness assertions as a fleet-wide default, since an inventory nobody can trust is worse than none: it converts an unknown into a false assurance the organisation then relies on.
## Read the failure correctly The document is not corrupt and the tool did not error. It reported everything it had evidence for. Two analyzers ran; one found a package database full of entries, the other found nothing it recognised. **Absence of components and absence of evidence look identical in the output** — that is the whole lesson of this scenario, and it is why interviewers like it. ## Why the halves differ OS packages are easy: the distribution's package manager maintains a database at a well-known path, and every installed package has a row in it. It is a single, authoritative source. Language packages have no such registry. For Ruby, a generator looks for installed gem specifications (the `.gemspec` files a gem installation writes) or a resolved lockfile listing the pinned set. Both are conventions about *where files land*, and any of these break them: - **Vendored bundles.** Gems installed into a project-local vendor path may not sit anywhere the analyzer walks by default, especially if the scan was scoped to a subtree. - **Version managers.** A Ruby installed under a per-user version-manager tree puts its gems somewhere the tool has no reason to look. - **Multi-stage builds.** The `Gemfile.lock` and the gemspecs existed in the build stage; the final stage copied only the application directory, so the evidence never reached the image being scanned. - **Stripped images.** A cleanup step that removes documentation and specification directories to shrink layers takes the evidence with it. - **Wrong subject.** The SBOM was generated over the base image or an intermediate stage rather than the published one. ## A diagnosis order that works 1. **Confirm the subject.** Which digest was scanned? If the pipeline scanned a stage or a tag other than the artifact you ship, stop here — everything downstream is about the wrong thing. 2. **Look inside the image.** Start a shell in it and find where gems actually live and whether specification files survived. If they are gone, no generator will find them. 3. **Turn on the generator's verbose or debug output** and read which analyzers ran and what they matched. This distinguishes "analyzer never fired" from "analyzer fired and found nothing". 4. **Re-run scoped at the real path.** Point the generator at the filesystem directory that holds the bundle. If components appear, it is a path or scope problem; if not, the evidence is genuinely absent. 5. **Compare against a known-good count.** Whatever the application's resolved dependency set is, the SBOM should be in that neighbourhood. A count of zero for a Ruby application is a fault, not a result. ## Fix it where the evidence lives The durable fix is to generate the language half at the point where the manifest and lockfile exist, rather than trying to recover it from a stripped runtime image. That means generating during the build stage that installed the dependencies and carrying the document forward, and keeping the OS half generated over the final image. Merging the two views gives an inventory that matches what you ship. ## Make the silence loud The control that prevents a recurrence is a **completeness assertion**: after generation, fail the build unless each ecosystem you expect for this service yields at least a plausible number of components. A service you know is Ruby-based with an OS base should never publish a document with zero gems. This converts a silent gap into a red build, which is the only way a coverage regression gets noticed — nobody reads a passing SBOM. ## Why the asset matters here A CI worker image is not an ordinary workload. It holds the credentials that push artifacts and deploy them, so it is a high-value target for anyone who can influence the build. An inventory that quietly omits the entire application layer of that host means that when an advisory lands against one of those gems, your search across the estate returns nothing for the machine that can sign and deploy everything. The document was not merely incomplete; it was incomplete in exactly the place where being wrong costs the most. ## What a weak answer sounds like "The scanner is broken, try a different one." Swapping tools sometimes papers over the symptom because another generator happens to walk that path, but it leaves the underlying condition — evidence missing from the artifact — untouched, and the next ecosystem you add disappears the same way.
- Why is a silently empty ecosystem worse than a generator that exits with an error?An error stops the pipeline and someone looks at it. A successful run that emits a well-formed, schema-valid document creates confidence: policy checks pass, the document is published, and downstream consumers query it as if it were complete. The gap is only discovered when an advisory lands and your estate-wide search comes back empty for a host that is in fact affected.
- How would you stop this recurring across dozens of service images?Make completeness a build-time check rather than a review habit. Each service declares the ecosystems it expects, and the pipeline fails if the generated document has zero components for any of them. Pair that with generating the language inventory in the stage where manifests exist, so the check has something to find. Both are cheap and catch the regression the day it is introduced.
- The team suggests fixing it by keeping the lockfile in the final image. Is that a good idea?It works and is often reasonable, but it is a workaround: you are shipping build metadata into a runtime image so a scanner can read it later. Generating the inventory at the stage where that evidence naturally exists, and attaching the document to the artifact, gets the same coverage without changing what you ship. Choose deliberately rather than by accident.
saying these in an interview costs you the question
- Concludes the image really has no gems
- Blames the tool and proposes swapping generators
- Assumes a successful exit code means complete coverage
- Ignores which build stage was actually scanned
- Treats an empty section as a formatting problem