skip to content

Your stack of Kubernetes pod-create records by container image is all count 1 - what went wrong?

level: middleimportance: should knowfreq 48%

answer

  1. the value must be stable
  2. per-build entropy in the tag
  3. normalise, then count
  4. controller identity, not the developer's
  5. you cannot stack what the policy dropped

basics

~20 s

The counted value carries per-build entropy: image tags and digests change on every build, so each deploy becomes a distinct value. Normalise before counting - strip tag and digest and group on registry host plus repository path.

solid answer

~50 s

Stack counting only works if the grouping value is stable while behaviour is unchanged. A full reference like `registry.internal/team-a/api:v1.4.2-9f3ac1` changes on every build, so a hundred deploys of one service produce a hundred singleton rows. I normalise first: drop the tag and digest and group on registry host plus repository path, collapsing those hundred rows into one. The same discipline applies to the other fields on that record - pod names carry a generated suffix, and per-developer namespaces are unique by design, so I collapse them to team. Two traps are specific to this source: the actor on a pod create is usually the ReplicaSet controller's service account rather than the person who deployed, so stacking on the user tells you almost nothing; and if the audit policy runs at `Metadata` level the pod spec is never recorded, so image and command are not fields you can stack at all.

code

json · 20 lines
json
{
  "kind": "Event",
  "apiVersion": "audit.k8s.io/v1",
  "level": "RequestResponse",
  "stage": "ResponseComplete",
  "verb": "create",
  "user": { "username": "system:serviceaccount:kube-system:replicaset-controller" },
  "objectRef": { "resource": "pods", "namespace": "team-a-dev" },
  "requestObject": {
    "spec": {
      "containers": [
        {
          "image": "registry.internal/team-a/api:v1.4.2-9f3ac1",
          "command": ["/bin/sh", "-c", "..."]
        }
      ]
    }
  },
  "responseStatus": { "code": 201 }
}

go deeper

for a junior

Know that the value you count has to stay the same across ordinary repeats. If it changes on every build, every row will read one and the stack teaches you nothing.

for a middle

Name the entropy in each field of the record - tag or digest, generated pod suffix, per-developer namespace, argument strings - and the normalisation you would apply to each.

for a senior

Show that you confirm the field exists at the audit level in force before building a hunt on it, and that you validate a normalisation against a change you know really happened.

for a principal

Own the standing ask: audit policy level, image tagging convention and namespace naming decide which hunts are possible at all, and those are platform decisions that need an owner outside the SOC.

## The requirement the field has to meet Stack counting produces signal only when the value you group on is **stable across ordinary repetition and changes when behaviour changes**. Every distinct value costs you a row; a field carrying per-record entropy therefore turns the whole data set into its own tail, and a table of four thousand rows all reading `1` contains exactly as much information as no table at all. So before counting anything, do an **entropy audit** of the record: for each candidate field, ask what makes it change when nothing meaningful has changed. ## The entropy in a Kubernetes pod-create record - **Container image reference.** `registry.internal/team-a/api:v1.4.2-9f3ac1` embeds a version and often a build or commit id; a digest pin (`@sha256:...`) changes on every rebuild by design. Normalise to registry host plus repository path. The hundred builds of one service collapse to one row, and a workload from a repository nobody else uses still stands out. - **Pod name.** Pods created by a Deployment carry a generated suffix, so the name is unique by construction. It is never a grouping field; it is an identifier you use *after* a row interests you. - **Namespace.** In an estate with per-developer or per-team namespaces, the namespace is close to a user id. Collapse it to a team or an environment class, or drop it from the key and use it as an enrichment column. - **Container command and args.** The executable is stable; the arguments usually are not, carrying job ids, timestamps and generated paths. Group on the executable path, or on the executable plus a normalised argument shape with numeric and hex tokens masked. If arguments still dominate, count the argument count and the flag set rather than the string. - **Actor.** Pods for a Deployment are created by the ReplicaSet controller, so `user.username` reads `system:serviceaccount:kube-system:replicaset-controller` across most of the estate - one enormous row and no information. The human identity sits on the *workload* create (the Deployment), which is the record to stack when you want to know who. ## The field you cannot stack because it was never written A Kubernetes audit policy sets a **level** per rule. At `Metadata` the event records who, what, when, from where, and the outcome - but no request or response body. Only at `Request` or `RequestResponse` does the pod spec, and therefore the container image and command, appear at all. Confirm the level for the resource you intend to hunt before designing the hunt around a field; discovering mid-hunt that the column is empty is a common and avoidable waste. ## How coarse to normalise Normalisation is a dial, and both ends are useless. Too fine and every row reads `1`. Too coarse - grouping every internal image into a single `registry.internal` row - and the thing you were hunting for is buried inside the head. The test to state out loud: **would a change that matters produce a new value, and would a change that does not matter leave the value alone?** Stripping the tag passes it, because a rebuild of the same service keeps the repository path while a workload from a genuinely new repository creates a new row. Stripping the whole path fails it, because a new repository under the same registry produces no new value. Verify rather than assume. Pick a change you know happened - a service that genuinely started deploying a new image last month - and confirm your normalisation still produced a new row for it. If it did not, the grouping is too coarse. ## A practical sequence 1. Confirm the record source actually carries the field, at the audit level in force. 2. Enumerate the entropy in each candidate field. 3. Apply the normalisation and count. 4. Look at the *shape* of the result: a healthy stack has a fat head and a thin tail. If the head is missing, normalise more coarsely; if the head swallows everything, less coarsely. 5. Keep the raw record joinable, so a tail row can be expanded back into full events for enrichment. ## Mistakes that show up in interviews - Counting the raw field and reporting the resulting unique-value count as a finding. - Stacking on pod name, or on any generated identifier. - Assuming every audit event contains the object's spec. - Naming the pod-create actor as the developer who ran the deploy. - Normalising so aggressively that the whole estate lands in a handful of rows.

  • How coarse should the normalisation go?
    As coarse as the estate has convention, and no coarser. Registry plus repository path groups builds of one service; the registry host alone groups everything internal into a single row and hides what you are hunting. Test it: a change that matters must create a new value, and a mere rebuild must not.
  • You stack the container command and every row is still unique. What now?
    The arguments carry the entropy - job ids, timestamps, generated paths. Group on the executable path alone, or on the executable plus an argument shape with numeric and hex tokens masked. If that is still unique, count the argument count and the set of flags rather than the literal string.
  • Why does stacking pod-create records on the user field usually fail?
    Pods for a Deployment are created by the ReplicaSet controller, so the audit user is that controller's service account across most of the estate: one giant row, no information. The human identity appears on the workload create, so stack Deployment creates if the question is who deployed something unusual.

saying these in an interview costs you the question

  • Counts the raw field without normalising it
  • Stacks on pod name and reports thousands of unique values
  • Assumes every audit event contains the object's spec
  • Names the pod-create actor as the deploying developer
  • Normalises so coarsely that the estate lands in three rows

context