Your Kubernetes estate emits no telemetry for a technique your Windows fleet detects - how should that coverage cell read?
answer
- not every gap is a rule gap
- a fourth state, not a darker red
- different owner, different budget
- not-applicable is the flattering error
- API audit is not process telemetry
basics
~20 sAs a distinct no-visibility state, not a failed detection. Red says a rule is missing and detection engineering can fix it; no telemetry says the raw material is absent, a platform team owns it, and it costs money.
solid answer
~40 sGive it its own state. On most heatmaps red means "we could see this and have not built a rule" - an actionable detection-engineering item. "We collect nothing that would ever show this" is a different claim with a different owner, cost and lead time: somebody must deploy a runtime sensor or enable a channel and pay for the ingest. Colouring both red loses that routing, and Kubernetes is the sharp example - the API server audit log records requests to the API, so it shows an `exec` request against a pod, but never a process started inside a running container. Keep no-visibility strictly apart from not-applicable: not-applicable means the behaviour cannot occur there and legitimately leaves the denominator, while no-visibility means it can occur and you would never know.
go deeper
Know that a rule cannot exist without a log source underneath it, so 'no detection' and 'nothing to detect on' are different situations. Be able to say which one a platform with no sensor is in.
Explain the mechanics of the specific gap: what an API server audit log records and what it does not, and why in-container process telemetry needs a separate node-level mechanism. Then name the fourth state and why three colours cannot express it.
Show that you route each state to the owner who can close it, attach cost, lead time and a date to a visibility gap, and state the compensating position at other layers instead of leaving the cell blank. Refuse to average platforms into one figure.
Own the sequencing argument: a visibility gap on a platform carrying real workloads blocks every detection above it, so it outranks a stack of missing rules on a platform you can already see - even though the grid draws both as one square.
## Three colours are not enough states The usual heatmap offers covered, partial and not covered. That vocabulary has an assumption baked into it: that every cell is a detection problem, and the difference between cells is how much detection work has been done. For a large part of a modern estate that assumption is false, and the falsehood is flattering. A cell needs at least four states: - **Covered**, at a stated assertion level. - **Not detected, but observable** - the log source is collected, nothing is watching it. Detection engineering owns this; it costs days. - **No visibility** - the estate emits nothing that would ever show the behaviour. A platform, endpoint or infrastructure team owns this; it costs a deployment, an ingest bill, sometimes a performance budget and a change window. - **Not applicable** - the behaviour cannot occur on this platform at all. Only the first two are SOC work items. Collapsing the third into the second sends the gap to a team that cannot fix it, and it will sit unfixed while looking like a backlog item. ### Why Kubernetes is the sharp case The Kubernetes API server audit log is a record of *requests made to the API*: who called it, what resource, what verb, what the outcome was. It will show a `create` on a pod, and it will show an `exec` subresource request against a running pod. What it will never show is a process started *inside* a container - that is below the API entirely. Container runtime and in-container execution telemetry comes from a different mechanism: a node-level sensor, kernel-level instrumentation, or runtime security tooling deployed to the cluster. So an execution technique that is green on the Windows endpoints, where a process-creation record with the command line exists, is not merely un-ruled on the container platform. It is unobservable there until somebody deploys something new. No amount of query writing changes that, and a red cell implies otherwise. ### The state that flatters, and how it gets chosen The most common quiet dishonesty in this artefact is marking an un-instrumented platform *not applicable*. It reads as rigour - we scoped it out - and it removes the cell from the denominator, so the coverage percentage goes up rather than down. But not-applicable is a claim about the behaviour (it cannot happen here) and no-visibility is a claim about the sensors (it can happen and we would not know). One is a fact about the platform; the other is a fact about your budget. Getting them the wrong way round is how an estate ends up reporting ninety per cent coverage of an environment it cannot see into at all. ### What you owe a cell once you mark it no-visibility A state with no obligations attached becomes a permanent parking space. The minimum is: - **The missing source, named.** Not "container visibility" but the specific mechanism that would produce the record. - **An owner outside the SOC.** The team that can actually deploy or enable it. - **A cost and a lead time.** Ingest volume, licence, change window, performance impact. This is what turns the cell into a decision somebody can make. - **A date.** Either the date it will be fixed by, or the date the decision not to fix it gets reviewed. - **The compensating position, if any.** What you would still see: an outbound connection at the network layer, a control-plane action the workload took afterwards, an identity artefact. Usually something remains, and saying what it is turns a blank cell into a partial one at a different layer. At that point the gap has stopped being a colour and become an owned item, which is the only form in which it ever gets closed. ### One number over three platforms The same discipline kills the single estate-wide percentage. A technique that is validated on managed Windows endpoints, un-ruled on the Linux fleet and unobservable on the container platform is three different facts, and the coverage figure that averages them describes no environment that exists. Report per platform, or report the worst, and be explicit about which you did. If a reader wants one number, the number they need is the one for the environment their data actually lives in. ### What this changes about how you work the gap Separating the states also changes the sequencing. Detection gaps can be worked continuously by the team that owns the rules. Visibility gaps have to be argued for, budgeted, scheduled and rolled out, and they block every detection that would sit on top of them - so a visibility gap on a platform carrying real workloads outranks a dozen missing rules on a platform you can already see, even though the heatmap draws them the same size.
- Is 'not applicable' the same as 'no visibility'?No, and the difference is the most flattering error in the artefact. Not applicable is a claim about the behaviour - it cannot occur on this platform - and it legitimately leaves the denominator. No visibility is a claim about your sensors: the behaviour can occur and you would never know. Marking the second as the first raises the coverage percentage while making the blind spot invisible.
- The platform team says the ingest cost is unacceptable. Now what?The cell stays no-visibility, and the decision gets an owner, a date and a written rationale rather than a colour change. Say what you would still see at other layers - network egress, control-plane actions the workload takes, identity artefacts - so the residual position is explicit. An owned, dated blind spot is defensible later; a cell quietly recoloured to not-applicable is not.
- Does a no-visibility cell belong in the coverage denominator?It should be counted and shown, not silently removed, because removing it is exactly what makes the number climb while the estate gets darker. The usual compromise is to report coverage over applicable cells and report the no-visibility count alongside it as a second figure, so a reader can see how much of the environment the percentage was never computed over.
A camera that is switched off and a corridor with no camera at all are both unwatched. Only one of them is fixed by changing the recording schedule.
saying these in an interview costs you the question
- Colouring a blind spot red as if a rule would fix it
- Marking un-instrumented platforms not applicable to protect the score
- Assuming Kubernetes API audit records in-container process execution
- Leaving the missing log source unnamed and unowned
- Averaging three platforms into one coverage percentage