One unexplained entry in a Linux server's ~/.ssh/authorized_keys: what defines the scope of your investigation?
answer
- one host is not a scope
- three axes, not one
- how many, which accounts, since when
- population chosen by shared exposure
- cheapest query with a decisive answer
basics
~20 sScope has three axes: how many hosts carry the same artefact, which accounts it grants access to, and over what time window you can see. One host with no accounts and no time bound is a finding, not a scope.
solid answer
~40 sOne host is where the alert landed, not the size of the problem. I scope on three axes. **Hosts**: which other machines carry this same public key or were exposed the same way — the config-management group, the golden image, the key-distribution path — searched by key fingerprint first because that is one cheap query with a decisive answer. **Accounts**: which local account owns the file, what that account can reach or `sudo` to, and which identities the key could plausibly belong to. **Time**: how far back each log source can actually answer, which differs per source. I write the population and the window down before I start searching, so the result means something. Scoping is not declaring an incident — it produces the facts that decision is made on.
go deeper
Be ready to name the three axes out loud — how many hosts, which accounts, since when — and to say why the alerting host is a starting point rather than the answer.
Expect to explain how you pick the host population: shared config group, image or key-distribution path, and a fingerprint search as the cheap first query. Say what the account can reach.
Show that you sequence the searches by decisiveness rather than thoroughness, and that you write down the population and window before searching so the result is interpretable afterwards.
Own the separation between scoping and declaring. Be able to argue why the two must stay distinct steps with distinct owners, and what goes wrong operationally when analysts collapse them.
## The situation An alert (or a config-drift report) puts one line in front of you: a public key in `~/.ssh/authorized_keys` belonging to a service account on one Linux application server, and nobody can say where it came from. There is no execution behind it, no second host, no beacon, no ransom note. It is a *capability* — anyone holding the matching private key can authenticate as that account. The first thing interviewers listen for is whether you treat "one host" as the answer or as the starting point. ## Axis 1 — how many hosts The alerting host is where you happened to look. The right population is chosen by **shared exposure**, not by who alerted: - hosts in the same configuration-management group or role - hosts built from the same golden image or provisioning template - hosts fed by the same key-distribution mechanism - hosts reachable from this one with the same credential The cheapest discriminator is the key itself: a public key has a stable fingerprint, and searching the fleet's `authorized_keys` files (or your config-management inventory) for that fingerprint is one query. If it appears on twelve hosts, the picture changed completely; if it appears on one, you have bounded the problem for the price of a single search. That is the shape to aim for at every widening step — spend the next twenty minutes on the query with the most decisive possible answer, not the most thorough one. Jumping straight to "search all 900 servers' audit logs" is not thoroughness; it is a way to spend a day and learn nothing. ## Axis 2 — which accounts A key in `authorized_keys` grants login **as the account that owns the file**. So: which account is it, what can that account do on the host, what secrets sit on disk readable by it, what can it `sudo` to, and what does it use to reach other systems (database credentials, cloud instance role, deploy tokens)? The blast radius of the artefact is the reach of that account, not the importance of the host. The trailing comment on the key line (`deploy@build01`) is a lead, never evidence. It is free text on the line; `ssh-keygen` sets it to `user@host` by default, which is exactly why people over-trust it, and anyone who can write the line can write anything into it. ## Axis 3 — since when Every finding needs a time window, and the window is bounded by what each source can answer, not by what you would like to know. `auditd`, centrally shipped `sshd` authentication logs, configuration-management run history and host backups all have different horizons. "Since when" is often the axis that turns out to be partly unanswerable, and the honest output is a stated range rather than a guess. ## What scope is not Scoping answers *how big and how far back*. It does not declare an intrusion, it does not contain anything, and it does not reconstruct how the intruder first got in — those are separate steps with separate owners. Keeping them separate matters, because the pressure to skip from "unexplained key" to "we are compromised" (or to "it is fine") is exactly what scoping exists to resist. ## Direction of the claims Watch what each fact actually proves. The key's presence proves someone with write access to that file put it there — not that it has ever been used. A successful public-key authentication record proves *the credential was accepted*, not that a particular person was at a keyboard. And the absence of a matching record proves only that the source you searched holds none, over the window it covers. ## What a good answer sounds like "One host, one account, no time bound is not a scope. I would fingerprint the key and search the fleet for it, list what that service account can reach, and set a per-source lookback window — then decide whether this is one stray key or twelve hosts' worth of standing access."
- Why search the fleet by key fingerprint before anything else?Because it is one cheap query with a decisive answer either way. A public key has a stable fingerprint, so a fleet-wide search of authorized_keys files or the config-management inventory either bounds the problem to this host or immediately multiplies it. Expensive searches come after you know the population.
- What does the comment at the end of the key line tell you?Almost nothing on its own. It is free text on the same line as the key; ssh-keygen defaults it to user@hostname, which is why it looks authoritative, but whoever wrote the line chose it. Use it to decide who to ask, never as evidence of origin.
- How is scoping different from declaring an incident?Scoping produces the facts: how many hosts, which accounts, over what window. Declaring is a separate decision made on those facts by whoever owns that call. Merging them is how one ambiguous artefact becomes either an unnecessary estate-wide response or a premature close.
A single unexplained key is like finding one copied door key on the floor of an office. The question is never just that door — it is how many locks it opens, who it lets in as, and how long copies have been circulating.
saying these in an interview costs you the question
- Says the scope is one host because only one host alerted
- Treats the key's trailing comment as proof of who owns it
- Widens to every host in the fleet before trying a cheap discriminator
- Omits the time axis entirely and reports only hosts
- Assumes an accepted key means a specific person logged in