skip to content

Scoping an unexplained authorized_keys entry, you find a config-management run deployed it — why not close as benign?

level: seniorimportance: should knowfreq 46%

answer

  1. an explanation is not an authorisation
  2. the first plausible answer is the trap
  3. both hypotheses predict the same run
  4. find the observation where they differ
  5. config-managed means more hosts, not fewer

basics

~20 s

Because the run explains the mechanism, not the authorisation. The next question is which commit put the key in the configuration repository, who wrote it and who reviewed it. That answer also widens the scope to every host in the group.

solid answer

~50 s

Finding the explanation you went looking for is the most dangerous moment in a triage, not the safest. Automation deploying the key tells me *how* it arrived; it says nothing about *who authorised it*. An intruder with commit access to the configuration repository produces evidence that looks exactly like this. So I pick an observation where the two hypotheses differ: the commit that added the key — its author, its review record, the identity that pushed it, when it landed — and whether that person recognises it. Note also what this finding does to scope: config-managed means the key is on **every host in that group**, so the population just grew, and the repository's git history outlives `auditd`, so "since when" may now be answerable from the commit date. The finding widens the case before it can close it.

go deeper

for a junior

Know that finding a plausible explanation is not the same as verifying it, and that automation deploying something does not mean somebody approved it. Say what you would check next.

for a middle

Explain where the deciding evidence lives — the repository commit, its author and its review record — and why the host-side facts are consistent with both the benign and the malicious story.

for a senior

Demonstrate the habit: state both hypotheses, identify the observation that separates them, and show that the finding widened the host, identity and time axes rather than closing the case.

for a principal

Be ready to discuss how your organisation prevents the automation path from being an unaudited persistence channel, and what evidence a triage analyst should be able to obtain without escalating to the platform team.

## The trap This leaf's characteristic failure is not missing evidence. It is the analyst who forms a benign hypothesis, goes looking, finds exactly what they expected, and stops. Confirmation bias plus premature closure, at the precise moment when widening was supposed to start. The asymmetry is what makes it expensive. Spending another thirty minutes on an authorised key costs thirty minutes. Closing a genuine intrusion as benign ends *all* further work on it — no second host is checked, no other account is reviewed, and the case label tells the next analyst there is nothing here. ## Why the config run is not an answer Configuration management deploying the key answers the mechanical question: the key arrived because a run applied a desired state that contained it. It leaves the security question untouched: - **Who put the key into that desired state?** - **Was that change reviewed, and by whom?** - **Does the change correspond to a request anyone can point to?** An adversary who obtained commit rights to the configuration repository — through a stolen developer credential, a compromised CI token, or a self-approved change — deploys persistence *through your automation*, and the resulting host state is indistinguishable from an authorised deployment. Automation is a distribution channel, not an authorisation. ## The technique: name the discriminating observation A disciplined widening step states both hypotheses and then asks which observation separates them. | | Benign hypothesis | Adversary hypothesis | | --- | --- | --- | | Config run applied the key | predicted | predicted | | Key present on all group hosts | predicted | predicted | | Reviewed commit by a named engineer who recognises the change | predicted | not predicted | | Commit authored by an identity whose owner denies it, or pushed outside the normal path | not predicted | predicted | The first two rows carry no information, which is precisely why finding them feels like progress and is not. The decisive evidence lives in the repository and the identity that wrote to it, not on the host. This works in the other direction too: an analyst who *expects* an intrusion will read a legitimate administrator's key as an implant and escalate on the same reasoning error. The cure is the same — state what would have to be true for the other hypothesis, then look for it. ## What this does to scope The finding should expand the investigation before it can end it: - **Hosts**: the population is now the membership of that configuration group, which you can enumerate exactly. One host became N, and you know N precisely rather than guessing. - **Accounts and identities**: the interesting identities are no longer only the service account on the host, but everyone with write access to the configuration repository and the pipeline that applies it. - **Time**: git history does not rotate the way host audit logs do, so the commit date may answer a "since when" that `auditd` could not. That is why this is a scoping question rather than a verdict question. Discovering the deployment path changes all three axes at once. ## Closing it properly when it *is* benign If the repository shows a reviewed commit by a named engineer, you confirm out of band — ask the engineer, not the commit message — and record the commit id, author and reviewer in the case. Then close it as a **benign true positive**: the detection was correct, the artefact was real, the activity was authorised. Closing it as a *false positive* would be wrong and harmful, because it feeds a tuning signal saying the rule misfires on something it actually caught correctly. ## What you are not doing here You are still scoping. Declaring an intrusion, isolating hosts and notifying anyone are separate decisions with separate owners, and reconstructing the intruder's full path is a different exercise again. The output of this step is a bigger, better-bounded picture and a named piece of evidence that decides the question — not the decision itself.

  • The commit author says they do not recognise the change — what changes now?
    The hypothesis moves from the server fleet to the configuration pipeline: a developer credential or CI token with write access is the likely path. The identity population becomes everyone who can write to that repository, and you have gained a firm date from the commit, which the host logs could not give you.
  • The repository shows a reviewed commit by a named engineer who confirms it. How do you label the case?
    As a benign true positive, not a false positive. The artefact was real and the detection was right; the activity was authorised. Record the commit id, author and reviewer so the next reader can see the close rested on evidence rather than on the run merely existing.
  • How does discovering config-management deployment change the host population?
    It defines it exactly. Instead of guessing which hosts might share exposure, you enumerate the configuration group's membership — every one of them has the key. The scope typically grows, but it becomes precisely bounded rather than open-ended.

Finding the delivery van that dropped off the parcel explains how it got to the door. It tells you nothing about who ordered it, and that is the only question that mattered.

saying these in an interview costs you the question

  • Accepts the config run as the explanation without checking the commit
  • Closes as a false positive when the key was genuinely deployed
  • Assumes automation implies authorisation
  • Narrows back to one host after learning config management pushed it
  • Looks only for evidence supporting the benign story

context