skip to content

Which shared observables justify merging two intrusions into one activity cluster?

level: middleimportance: should knowfreq 54%

answer

  1. cheap to change, cheap to share
  2. who else could plausibly have this?
  3. library defaults are not tradecraft
  4. two independent non-commodity overlaps
  5. record what carried the merge

basics

~20 s

Merge on overlaps the operators control and could not cheaply replace: a non-public implant with the same embedded configuration, a reused key or certificate, distinctive hands-on tradecraft. Shared hosting space, default TLS fingerprints and commodity tooling prove nothing alone.

solid answer

~50 s

I ask two questions of every shared observable: who else could plausibly have it, and what would it cost the operators to change it. A hosting provider's autonomous system is shared by thousands of unrelated tenants. A JA3 value is a hash over TLS ClientHello fields, so every tool built on the same library version produces the same one. An imphash is a hash of a PE import table, and unrelated binaries built from the same compiler and imports collide. Those are leads, not glue. Strong glue is exclusive and expensive: a private backdoor carrying the same embedded key or campaign identifier, a code-signing certificate, a distinctive multi-step command sequence with the same staging path and the same typo. My working rule is at least two independent non-commodity overlaps before merging, and I record which ones they were so the merge can be unwound later.

go deeper

for a junior

Know that some shared artefacts are meaningful and many are not, and that a shared hosting provider or a common tool is the classic false link. Be able to name one strong and one weak example.

for a middle

Explain the mechanics: what a JA3 or an import hash is actually computed over, and why that makes them shared by unrelated software. This is the tier where interviewers expect precision about the artefacts.

for a senior

Demonstrate a merge standard you actually apply - independent non-commodity overlaps, written records of what carried each merge, and a habit of re-testing old clusters for circular reasoning.

for a principal

Frame it as a quality bar the team is held to: what evidence is required before a linkage leaves the building, who reviews it, and what the cost of a wrong merge is to the people who consume your reporting.

## What you are actually joining on Clustering joins closed cases on observables drawn from your own telemetry. They fall into a few families: - **Infrastructure** - domains, IP addresses, hosting address space, TLS certificates, registration patterns, the fingerprint a server presents. - **Tooling** - implant families, build artefacts, embedded configuration, mutex names, protocol quirks, hashes of a file or of parts of it. - **Tradecraft** - the sequence and shape of hands-on activity: which directory files are staged in, how archives are named, which native tools are chained in which order, the hour of day the operator works. - **Victimology and timing** - the sector and size of target, and whether cases fall in a coherent window. ## The two tests that decide the merge **Exclusivity: who else could have this?** An observable that thousands of unrelated parties share carries no information about a common operator. **Cost to change: what does the operator lose by rotating it?** Something an operator can replace in ten minutes tells you little; something baked into a build pipeline, a signing key or a person's habits tells you a lot. These two are why cheap network artefacts sit at the bottom and human tradecraft sits near the top. ## The commodity traps, precisely - **A shared hosting autonomous system.** A cheap VPS reseller's address space hosts thousands of tenants, many of them criminal and unrelated to each other. "Same ASN" is close to no evidence at all. - **A TLS client fingerprint (JA3).** JA3 is a hash over fields of the TLS ClientHello - version, cipher suites, extensions, curves. It fingerprints the *TLS stack*, not the actor. Two entirely unrelated tools compiled against the same Go or Python version emit the same value, and a benign application can too. - **A server-side fingerprint (JARM).** JARM is produced by actively probing a server with a series of crafted ClientHellos and hashing the responses. Every default installation of the same C2 framework on the same platform answers alike, so a JARM match finds *the framework*, not the crew. - **An import hash (imphash).** A hash over a PE file's ordered import table. Unrelated programs built with the same compiler and the same imports collide, and a single recompile with one changed import breaks the match. Useful as a pivot, weak as glue. - **Commodity capability.** A leaked or cracked commercial C2, a public exploit for a widely deployed appliance, remote-management software abused by a dozen different crews, a ransomware affiliate programme's toolkit. All shared by design. ## What strong glue looks like - A **non-public implant** whose configuration blob carries the same key, campaign identifier or hard-coded fallback host. - A **reused cryptographic artefact** - a code-signing certificate, an SSH host key, a self-signed certificate with an identical unusual subject appearing on infrastructure in both cases. - **Operator habits**: the same staging directory, the same archive naming scheme, the same misspelled command, the same order of enumeration commands typed within seconds of each other. - **An unusual chain of techniques** performed the same way, where each step alone is common but the specific composition is not. ## Two independent overlaps, and why "independent" matters A useful bar is **two non-commodity overlaps that do not derive from each other**. A domain and the certificate on that domain are one fact, not two. A sample's hash and its import hash are one fact. Independence means the overlaps would have to coincide by chance twice. ## Circularity is the failure mode The characteristic clustering error is self-reinforcing: you build a cluster on a weak observable, then admit every new case that shows the same weak observable, then cite the size of the cluster as evidence it is real. Each admission looks justified, and the whole structure rests on one commodity artefact. The defence is mechanical: record for every merge which specific observables carried it, and periodically re-read those records asking whether any single one of them, if it turned out to be commodity, would collapse the cluster. ## Write the merge down A merge is a hypothesis with a date, an author and a list of observables. If the record does not exist, a future analyst cannot tell whether the cluster was joined on a private implant key or on a shared VPS - and cannot split it responsibly when the evidence shifts.

  • Why is victimology a weak basis for a merge on its own?
    Because a sector is a shared target, not a shared operator. Regional banks are attacked by many unrelated crews for many reasons, and the overlap you observe is mostly a property of your own visibility - you see banks because you defend one. Victimology is useful as corroboration once a technical overlap exists, and as a hypothesis for a hunt, but it never carries a merge alone.
  • Your cluster was built entirely on one imphash and now has nine cases. How do you sanity-check it?
    Test the observable's exclusivity: how many unrelated samples in public and internal corpora share that import hash, and would a routine recompile change it? If it is common, treat the cluster as unproven, look for a second independent overlap within each case, and be prepared to demote it back to nine separate cases rather than defending its size. Cluster size is not evidence.

saying these in an interview costs you the question

  • Merges on a shared hosting provider or ASN
  • Treats a JA3 match as an actor fingerprint
  • Cites cluster size as evidence the cluster is real
  • Cannot say which observables justified an existing merge
  • Counts a domain and its certificate as two overlaps

context