skip to content

How do you sweep 300 Linux servers for persistence when a third have no EDR agent?

level: middleimportance: should knowfreq 47%

answer

  1. no sensor, but there is a baseline
  2. declared state versus actual state
  3. drift: added key, unknown unit, package hook
  4. no check-in is not a pass
  5. the host answers about itself

basics

~20 s

Compare every host against its declared configuration and its package database instead of against a sensor. Planted persistence usually surfaces as drift: an added SSH key, an undeclared systemd timer, a new package hook. The gap is unmanaged paths and hosts whose agent stopped reporting.

solid answer

~50 s

On a configuration-managed fleet you already have a baseline that does not depend on EDR: the declared state. A no-op or check-mode run reports, per host, where actual state differs from what the catalogue says, and package-database verification compares on-disk files against the distribution's own hashes. Persistence tends to appear as drift — an extra key in `authorized_keys`, a `.timer` and `.service` pair in `/etc/systemd/system` that no catalogue declares, an `/etc/apt/apt.conf.d/` drop-in with a `DPkg::Post-Invoke` line, an `/etc/cron.d` entry, an `ld.so.preload` addition. On a fleet built from one image, an artefact present on one host out of 300 is the signal. Know the limits and say them: the tool only reports on resources it manages, so an unmanaged directory is invisible; a compromised host can lie through its own agent; and a host that has not checked in produces no report, which is a hole, not a pass.

code

text · 8 lines
text
drift sweep - day 3 - 300 hosts in scope

web-041  /root/.ssh/authorized_keys       +1 key, comment 'backup@ops'   [path not managed]
db-07    /etc/systemd/system/sysstat-collect.timer (+ .service)         [not in catalogue]
app-112  /etc/apt/apt.conf.d/99-local     DPkg::Post-Invoke ...          [not in catalogue]
...
113 hosts  0 differences
187 hosts  no report - agent last check-in > 14 days

go deeper

for a junior

Know that persistence on Linux lives in a small set of predictable places — SSH authorized keys, cron, systemd units, package hooks — and that a fleet under configuration management already records what each host is supposed to look like.

for a middle

Explain the mechanics: a no-op or check-mode run yields per-resource actual-versus-declared differences, package verification compares files to recorded hashes, and each method has a defined blind spot you can name.

for a senior

Show how you would run this at fleet scale inside a live incident — prioritising rare artefacts, handling hosts with no data, and deciding when a host's self-report is too weak to accept.

for a principal

Frame the gap for the organisation: a third of the fleet without telemetry is a decision someone made, and eradication cost is the bill for it. Turning that into funded coverage is the durable outcome.

## The problem: partial sensor coverage during an eradication sweep You are inside an incident. You know the mechanisms this intruder plants because you recovered them from the hosts you did investigate. You now need to answer a fleet-wide question — where else does this exist — across 300 servers, a third of which have no endpoint agent to query. Waiting to install agents everywhere is not an answer while an intruder still holds access. ## The substitute source: declared state A configuration-managed fleet carries a second, independent description of what each host should look like. A Puppet no-op run or an Ansible check-mode run with diff reports, per resource, where the machine differs from the catalogue. Package-manager verification adds a second baseline: files installed from packages have known hashes, sizes, modes and owners recorded in the package database, and verification reports every file that no longer matches. Neither source needs security tooling on the host. That is the point — during eradication you are looking for a *difference from intended state*, and configuration management is a machine-readable statement of intended state that already exists. ## What you sweep for Translate the mechanisms you recovered into locations, and sweep for the class rather than the exact string, because a competent adversary varies file names: - `~/.ssh/authorized_keys` for **every** account with a shell, including root and service accounts, and the `AuthorizedKeysFile` path if `sshd_config` was changed to point somewhere else. - `/etc/systemd/system` and the user unit directories: a `.timer` unit paired with a `.service` unit gives an adversary cron-like execution that looks like ordinary platform plumbing. - `/etc/cron.d`, `/etc/cron.*` and per-user crontabs. - Package hooks: a drop-in under `/etc/apt/apt.conf.d/` carrying a `DPkg::Post-Invoke` command runs on every package operation on that host, which is a durable and low-traffic trigger. - `/etc/ld.so.preload` and additions to PAM configuration, both of which give execution inside legitimate processes. - Shell profile and rc files for accounts an operator actually logs into. On the Windows jump host in the same estate the equivalent lives somewhere a file sweep cannot reach: a WMI event subscription is a `__EventFilter`, a consumer and a `__FilterToConsumerBinding` registered in the CIM repository under `root\subscription`. You enumerate that namespace; no file diff will show it. ## Triaging the output Drift reports are noisy — operators make emergency changes, and every fleet has them. Two filters cut the volume fast. First, rarity: on hosts built from one image and one catalogue, an artefact present on one host and absent on the other 299 identical ones deserves attention that a difference present everywhere does not. Second, the mechanism class: an unexpected file in `/var/log` is drift, while an unexpected timer unit or a new key in `authorized_keys` is drift in a location whose only purpose is execution or access. ## What this method structurally cannot see Say these before an interviewer asks, because the honest boundary is what the question is testing: - **Unmanaged paths.** Configuration management reports on resources in the catalogue. A file dropped into a directory the tool does not manage, or into a managed directory without recursive purging enabled, produces no drift at all. - **Content it does not check byte-for-byte.** A resource declared only to exist can be modified without the tool noticing. - **The agent itself.** You are asking software running on a possibly-compromised host to describe that host. Root-level access is enough to make the answer whatever the adversary wants; the same is true of package verification, since the package database is a file on the same disk. - **Anything not on disk.** An implant living only in the memory of a running process, or persistence achieved through an interface the tool does not model — a scheduler, an orchestrator, an identity provider — leaves no file to differ. ## Hosts with no report are not clean hosts The most dangerous line in a sweep is the summary. If 113 hosts report zero differences and 187 produce no report because their agent has not checked in for two weeks, you have swept 113 hosts. Those 187 go on the unswept list, get collected out of band, and are treated as unknown in the eradication decision. An agent that stopped reporting during the intrusion window is itself a finding worth chasing rather than a scheduling annoyance.

  • Your report shows zero differences for 113 hosts and no report for 187 whose agent last checked in two weeks ago. What is your conclusion?
    That I have swept 113 hosts. The other 187 produced no evidence in either direction and go on the unswept list, to be collected out of band or rebuilt. I would also treat the stalled check-in as a lead rather than an operations chore: an agent that went silent during the intrusion window is either a coverage failure that hid the adversary or something they did deliberately.
  • Why would a file-level drift sweep miss the WMI event subscription on the Windows jump host?
    Because it is not a file. A WMI event subscription is three objects registered in the CIM repository under `root\subscription` — an `__EventFilter` that defines the trigger, a consumer that defines what runs, and a `__FilterToConsumerBinding` that ties them together. Nothing appears in a directory listing or a package verification, so the sweep for that host has to enumerate the namespace instead.
  • How much do you trust a clean drift report from a host where the adversary had root?
    Not much on its own. The agent, the package database and the file metadata all live on a disk the adversary controlled, so a root-level intruder can make the host describe itself however they like. A clean report from such a host is weak evidence; comparing from outside — a mounted image, network-side records, or the configuration server's own history — is what makes it stronger.

saying these in an interview costs you the question

  • Reads no drift report as a clean result
  • Trusts an agent running on a host the adversary held root on
  • Sweeps only the hosts that produced alerts
  • Believes a file diff would reveal WMI subscriptions or memory-only implants
  • Searches for the exact file names recovered rather than the mechanism class

context