skip to content

Stack counting reads the rarest rows - what kind of intrusion hides in the common ones?

level: middleimportance: nice to knowfreq 34%

answer

  1. the method assumes rare equals hostile
  2. the head is not cleared
  3. which column has the adversary's freedom
  4. automation makes malice common
  5. pair fields, or stack inside the head

basics

~10 s

Anything the adversary made common or borrowed from something already common: an approved image whose entrypoint was overridden, persistence that automation deployed hundreds of times, or a value shared with heavy legitimate use.

solid answer

~50 s

Least-frequency analysis encodes one assumption - that hostile activity is infrequent - and an adversary who breaks it is invisible to it. Three ways that happens. First, they work in a column you did not stack: pods launched from the estate's approved, thousand-pull base image but with `command` overridden at creation sit inside the top row of an image stack, because the freedom they used lives in a different field. Second, they inherit your automation: persistence pushed through a controller or a shared chart gets deployed as many times as the workload it rides in, so it is common by the time you look. Third, they choose a deliberately unremarkable value, such as a shell invocation thousands of legitimate containers run. The counter is not a better cut-off; it is stacking the field where the adversary has freedom and the estate has convention, and stacking pairs so a common image with an uncommon entrypoint becomes a rare row.

go deeper

for a junior

Remember the assumption the technique rests on: it finds the infrequent. Something an adversary got deployed a thousand times through your own automation is not infrequent.

for a middle

Explain how to defeat the blind spot mechanically - stack the field the adversary actually controls, pair two conventionally linked fields, or stack within the largest row.

for a senior

Show judgment about cardinality: each field added to the grouping key grows the tail, so justify each pairing by the convention it exploits rather than pairing everything available.

for a principal

Be explicit with stakeholders that a hunt evidences effort over a field set, not coverage of an estate, and that reporting a quiet stack as 'we are clean' is a claim the method cannot support.

## The assumption inside the method Stack counting ranks by frequency and asks the hunter to read the bottom of the distribution. That is only useful under an assumption: **the hostile thing is infrequent in the field you counted**. It is a good assumption often enough to make the technique the workhorse of hunting, and it fails in specific, predictable ways worth being able to name. ## Failure one: the adversary is in a column you did not stack Consider a cluster where almost every workload runs from a small set of approved base images. An adversary who can create pods does not need an image nobody has seen; they can start a pod from the most-pulled image in the estate and override the container `command` at creation to run whatever they like. Stack the image field and their pods land inside the largest row in the table. Stack the entrypoint and they are a singleton. The general rule: **stack the field where the adversary has freedom and the estate has convention**. A field the estate standardises but the adversary can still control is where rarity carries the most information. A field the adversary must inherit from the estate carries none. ## Failure two: your automation makes them common If something hostile gets into a shared chart, a base layer, a DaemonSet or an admission-injected sidecar, the estate's own deployment machinery replicates it across hundreds of pods. By the time a hunter counts, the artefact has a count in the hundreds and sits comfortably in the head. This is the failure mode with the widest blast radius, and the one the method is least equipped to find: the more thoroughly something spread, the more invisible least-frequency analysis makes it. ## Failure three: the value is deliberately unremarkable An adversary who runs `/bin/sh -c` shares a value with thousands of legitimate containers. Nothing about their execution is rare in that column, even though the specific instance is entirely abnormal in context. Rarity is a property of a value, not of an event, and an event can be extraordinary while the value it produced is ordinary. ## What to do about it **Pair fields.** A joint value can be rare while both components are common. An image with tens of thousands of pulls and a command run by thousands of containers may occur together exactly once. Pairing costs cardinality - every field you add to the grouping key grows the tail - so pair fields that are *conventionally linked*, where the estate normally couples them: image with entrypoint, service account with verb, registry with namespace. **Stack within the head.** Rather than reading the tail across a heterogeneous estate, condition on the largest row and stack inside it. Take the single most-deployed image and count the entrypoints, service accounts, namespaces or node pools its pods used. Rarity within a large, homogeneous group is a far stronger signal than rarity across a mixed one, because the population you are comparing against is genuinely comparable. **Change the record source.** If a technique is common in every field of one source, another source may see it differently. What is unremarkable in creation records may be unusual in what the container then did on the network or on disk. ## The negative-evidence limit This is where the technique's honesty matters most in an interview. A stack that surfaced nothing constrains **only the field set, window, normalisation and record source you used**. It is not a statement about the estate. An adversary who is common in that field, or absent from that record source entirely, leaves the stack looking exactly the same as a clean cluster. Report a hunt as 'no rare values in this field set over this window', never as 'the cluster is clean'. ## Common mistakes - Believing malicious activity must be rare. - Clearing the head without inspecting it. - Reading an empty tail as proof of absence. - Pairing every available field at once, which multiplies cardinality until the entire data set is the tail.

  • How do you hunt inside the top row instead of the tail?
    Condition on it. Take the single most-deployed image and stack the fields that vary underneath it - entrypoint and arguments, namespace, service account, node pool, the destinations its pods reached. Rarity inside a large, homogeneous group is a much stronger signal than rarity across a mixed estate, because the comparison population is genuinely comparable.
  • Does an empty tail mean the cluster is clean?
    No. A stack constrains only the field set, normalisation, window and record source it used. An adversary who is common in that field, or who never appears in that source, leaves the distribution looking identical to a clean estate. Absence of a rare row is not evidence of absence.
  • Why does pairing two fields find things that stacking either alone misses?
    Because the joint value can be rare while both components are frequent: a heavily used base image is common, a shell invocation is common, and that image running that command may have happened once. The cost is cardinality - the tail grows - so pair fields the estate conventionally couples rather than pairing everything.

saying these in an interview costs you the question

  • Believes anything malicious must be rare
  • Clears the top rows without inspecting them
  • Treats an empty tail as proof the estate is clean
  • Pairs every field at once and drowns in a larger tail

context