skip to content

What does a C2 beacon look like in TLS-encrypted flow and proxy metadata?

level: juniorimportance: must knowfreq 72%

answer

  1. rhythm and volume, never content
  2. same pair, same gap, same size
  3. session duration barely varies
  4. upload-to-download ratio, not payload
  5. updaters look identical

basics

~20 s

Repeated connections from one host to one destination at a near-constant interval, each session short and roughly the same size, often with more bytes out than in. Metadata carries the rhythm and the volume, never the content.

solid answer

~50 s

Even with the payload encrypted, the connection records still carry when, where, how long and how much. A beacon shows up as a repeating rhythm: the same source reaching the same destination or SNI over and over at a gap that clusters around one value, each session lasting a second or less, and each exchange moving a small, remarkably consistent number of bytes — commonly a slightly larger request than response while the implant is only asking "any tasking?". So you hunt three features together: check-in periodicity, session duration, and the upload-to-download ratio. What none of it proves is malice or content — a flow record proves bytes moved in each direction at a time, and nothing about what those bytes were. Polling clients, licensed telemetry agents and heartbeat health checks produce exactly the same shape, which is why the shape is a hypothesis and not a verdict.

go deeper

for a junior

Be ready to name the three metadata features that expose a check-in loop — interval regularity, short and uniform session length, small and symmetric byte counts — and to say plainly that metadata shows when and how much, never what.

for a middle

Explain why the ratio for an idle beacon differs from web browsing, and why the same signature is produced by updaters and management agents, so the first pass yields candidates rather than findings.

for a senior

Show how you turn the raw window into per-pair features and rank them, and how you keep the claim honest: bytes moved is provable, malice is not. Expect to defend the volume of benign traffic your pass returns.

for a principal

Frame what this analytic is worth on a segment where it is the only telemetry: it buys coverage where no agent can be installed, at the cost of verdicts that are always argued rather than confirmed. Be ready to say what you would fund to close that gap.

## The premise: you cannot read it, so read around it When a channel is TLS-encrypted, a passive observer on the wire or a proxy sees the envelope, not the letter. That envelope is still rich: the source, the destination address, the hostname in the TLS SNI or the proxy `CONNECT` target, the start time, the duration, and byte and packet counters in each direction. Hunting encrypted egress means treating those fields as a **time series and a set of ratios** rather than as individual records. ## The three features that betray a check-in loop **1. Periodicity.** An implant that has no inbound path to it must ask its controller for work. That produces a loop: connect, ask, disconnect, sleep, repeat. Sorted by time, the gaps between one host's connections to one destination cluster tightly around a base sleep value — 30 s, 60 s, 300 s — instead of following the ragged, bursty spacing of a human driving a browser. Jitter (randomising the sleep by some percentage) widens that cluster but does not remove it; the gaps still concentrate in a band rather than spreading uniformly across the day. **2. Session length and shape.** A check-in is a transaction, not a conversation. Sessions are short and their durations barely vary, because the same code path runs each time. Human-driven or content-driven traffic to the same destination varies wildly: a page load, an idle tab, a video, a large download. **3. Byte asymmetry.** Ordinary web browsing is download-heavy — you ask for a little and receive a lot. An idle beacon inverts or flattens that: a small request carrying host identity and status, and a small response that is usually "nothing for you". So a destination where a host uploads about as much as it downloads, in near-identical amounts every time, is anomalous *for a web-shaped destination*. When tasking does arrive, or when data is staged out, the ratio swings hard the other way — a single long session with a large outbound total and a trivial inbound one is exfiltration-shaped rather than beacon-shaped, and the two look different in the same data. ## What the data can and cannot support The discipline that separates a good hunter from a noisy one is being precise about the direction of each claim: | Observation | Supports | Does **not** support | |---|---|---| | Fixed-interval connections to one host | A programmatic client is running | That the program is malicious | | Small, symmetric byte counts | A control-style exchange, not content delivery | What was in the request | | Long session, heavy upload | Bulk data moved outbound | That the data was sensitive, or that it left the company | | Destination is a well-known CDN name | The name resolves into shared infrastructure | That the destination is trustworthy — domain fronting and abuse of sanctioned SaaS put attacker infrastructure behind exactly those names | And the negative that catches people out: an encrypted session's counters are *not* hidden by the encryption. Encryption hides the payload; the collector still counts bytes and measures time. Conversely, no amount of metadata analysis recovers the payload — if the question is "what did it send", metadata cannot answer it and you need a different source. ## Why the shape is a hypothesis, not a finding The uncomfortable truth of this hunt is that the benign population produces the identical signature at enormous scale. Software updaters, licence checks, EDR and management agents, MDM check-ins, printer and building-management telemetry, monitoring heartbeats, mail clients polling, and half the mobile applications on the estate all connect on a timer, hold the session briefly, and exchange a couple of hundred bytes. On a mature estate the periodic-egress population is *mostly* legitimate. That is why the correct output of the first pass is a candidate list, and why every candidate needs a discriminator — who owns the destination, whether the behaviour is shared by an entire device population, whether it started at an explainable moment — before anything is escalated. ## How you actually run the first pass Aggregate the window by `(source, destination-or-SNI)` pair and derive per-pair features: connection count, the median gap and its spread, median session duration, median bytes out and in, and the out/in ratio. Filter to pairs with enough connections to have a rhythm at all, then sort by tightness of the gap distribution. Read the top of that list by hand. Most of it will be software you can name in seconds; what remains — a destination nobody can attribute, on a host that has no business talking to it — is where the hunt actually begins.

  • How would the picture change if the same host were staging data out rather than checking in?
    The rhythm collapses into volume. Instead of many tiny, evenly spaced sessions, you see one or a few long sessions with a large outbound byte total against a trivial inbound one — an inverted ratio and a duration measured in minutes or hours. Some tooling deliberately chunks the transfer to stay beacon-shaped, so a beacon whose outbound bytes are steadily creeping up over days is worth reading as staging rather than tasking.
  • The destination resolves to a large content-delivery network. Does that make it safer?
    No. It makes attribution harder. A CDN name tells you which shared infrastructure answered, not who owns the content behind it, and adversaries deliberately front control channels on CDN and sanctioned SaaS names precisely because defenders read those as safe and because blocking the name would break legitimate traffic. Treat a CDN destination as an unresolved attribution problem, not as evidence of benignity.
  • Your only telemetry for this segment is network metadata — no endpoint agent. What does that cost you?
    You lose corroboration and you lose the process. Metadata can tell you a device is beaconing but not which binary opened the socket, what its parent was, or how it got there. Every verdict then has to be argued from the network pattern plus provenance you gather elsewhere — vendor documentation, firmware history, the device owner — and you have to state the residual uncertainty rather than pretend to a confirmation you cannot reach.

You are watching a letterbox, not the letters. You can see that an envelope of the same size goes out every minute and a slim one comes back — enough to say someone is running a routine, never enough to say what it says.

saying these in an interview costs you the question

  • Claims flow records reveal what the encrypted traffic contained
  • Says encrypted traffic carries no usable hunting signal at all
  • Treats any fixed-interval connection as confirmed command and control
  • Assumes a CDN or well-known SaaS destination is automatically benign
  • Confuses a flow or proxy record with a full packet capture

context