skip to content

Depth of Inspection

How much a device can learn about a flow it is already forwarding — the application behind the port, the user behind the address, the plaintext behind TLS — and what that sight costs.

on this pageshow

explore

questions

16

Why must an application-aware firewall pass a flow's first packets before naming the application, and what does that hand an adversary?

level: juniorimportance: must knowfreq 72%

answer

  1. the label is a verdict, not a header field
  2. evidence arrives with the traffic
  3. reassembly before matching
  4. teardown, not a refused connection
  5. a small payload fits the window

basics

~20 s

Identifying an application needs evidence, and the evidence is the traffic itself. The device reassembles and inspects the opening bytes before it can commit to a name, so those packets have already reached the far end when a deny finally fires.

solid answer

~50 s

Addresses are known at the first packet because the sender wrote them in the header. The application name is not written anywhere — it is a verdict the device derives from the reassembled byte stream: a protocol decoder recognising an opening exchange, handshake fields such as TLS SNI and ALPN, or behavioural heuristics on sizes and timing. So the device forwards a small opening window, then re-evaluates policy once a label exists. When that label is a denied application, the action is a mid-session teardown, not a refusal to connect. Practically: this control reliably kills a *session* — a tunnel, a bulk transfer, a long-lived channel — and does not reliably stop a small payload that fits inside the pre-verdict window. An adversary who only needs a few hundred bytes out gets them. Where that matters, pair it with a control that denies before the connection forms.

go deeper

for a junior

Be ready to say plainly that the application name is worked out from the traffic, so some traffic must pass first. Know that the resulting deny ends a session rather than preventing a connection.

for a middle

Explain the mechanics: reassembly before matching, decoders and handshake fields as evidence, heuristics as a fallback, and policy re-evaluation once a label exists. Say why buffering the flow instead is expensive.

for a senior

Show you design around the gap. State which threats this control genuinely stops (persistent sessions) and which it does not (a payload that fits the opening window), and name the coarser control you pair with it.

for a principal

Own the framing that this is a probabilistic control with a measurable exposure. Be able to describe that exposure to a customer or an auditor honestly rather than claiming the boundary blocks the application outright.

## The claim the device is making A filter that reads only addresses and ports has everything it needs in the first packet, because everything it needs was written into the header by the sender. An application-aware firewall makes a stronger claim — that the traffic in this flow *is* a particular application — and that claim is a **verdict derived from evidence**, not a field it can read. The evidence is the traffic itself. The device therefore cannot have it until some traffic has already arrived, and "arrived" at a forwarding device means "was on its way onward". ## Where the evidence comes from - **Stream reassembly.** Segments arrive split, out of order, sometimes retransmitted. A classifier that matched on individual packets could be defeated by splitting a distinguishing string across two of them, so it works on the reassembled stream — which means holding state and waiting for enough of the stream to exist. - **Protocol decoders.** Many applications announce themselves in their opening exchange: a request line, a banner, a binding request, a version negotiation. These are the cheap wins and land within a packet or two. - **Handshake metadata for encrypted flows.** A TLS handshake exposes the requested name and the negotiated next protocol before any application data flows, and the server certificate names a service. This names *who is being talked to*, never *what is being said*. - **Behavioural heuristics.** For anything with no readable header, the classifier falls back on shape: packet sizes, inter-packet timing, direction ratios, session length. Shape needs a stretch of traffic before it means anything, and it is the weakest evidence of the four. ## Why the verdict is late, and how late It varies by application, and that variability is the point. Something that identifies itself in its first client message can be labelled almost immediately. Something distinguishable only by how the *server* answers needs a full round trip. Something carried inside an encrypted session may need the handshake plus a run of data packets before a heuristic will commit — and some flows never resolve at all, which is how the unknown bucket fills. ## What the device does while it does not know The normal design is **forward, then re-evaluate**. The alternative — hold the flow until the classifier commits — is not free: it adds latency to every session, it costs per-flow buffer memory at exactly the aggregation point where flow counts are highest, and it breaks protocols where the server speaks first or where a client gives up quickly. Devices accept a small forwarded exposure instead of a large capacity and latency bill. When the verdict lands and the matching rule changes to a deny, what happens is a **session teardown** — the flow is dropped, often with a reset toward one or both ends. The connection existed. Bytes crossed. The session record shows byte counters that are not zero, and reading only the `deny` action understates what reached the server. ## The honest limitation, and it is the interview answer Application identification is a **strong control over sessions and a weak control over first payloads**. Anything that must persist to be useful — a tunnel, an interactive channel, a large transfer — dies at the verdict. A single small request and response that completes inside the pre-verdict window is finished before the label exists. If what you are defending against fits in a few hundred bytes, this is not the control that stops it; you need something that refuses before the connection forms, such as an explicit deny toward that peer at the zone boundary, or no path at all. The adversary side of this is not exotic. Traffic dressed to look like an ordinary permitted application is trying to be labelled as something allowed; traffic that only needs the opening window does not even need to win that argument, because it is done before the argument starts. ## What it costs the defender Three things, and you should be able to name all three: reassembly latency and memory on every flow; a forwarded exposure per flow that you cannot describe precisely to anyone who asks "how much gets through"; and, at a shared boundary that classifies several customers' traffic through one engine, an exposure that exists identically for every one of them while none of them has ever been quoted a number for it.

  • If the verdict is late, why not simply buffer the flow until the classifier commits?
    Because buffering is paid on every session, not just the bad ones. It adds latency to all traffic, needs per-flow memory at the busiest point in the path, and breaks protocols where the server speaks first or where the client times out quickly. Forwarding a small window and re-evaluating trades a bounded exposure for a much cheaper device.
  • Does the number of packets needed to identify an application vary?
    Substantially. A protocol that announces itself in its first message can be named in a packet or two; one whose distinguishing evidence is in the server's reply needs a round trip; one carried inside encryption may need the handshake plus a stretch of data before a shape heuristic commits, and some never commit at all and land in the unclassified bucket.
  • What does the session record show for a flow denied after the verdict landed?
    Both facts: the flow was permitted while unlabelled or labelled generically, then re-evaluated and terminated. The byte counters show what crossed before the teardown. If you report only the deny action, you are reporting an intent, not an outcome.

A doorman who must let you start talking before he can tell whether you are a courier or a salesman. He can stop the conversation, but not the sentence you already said.

saying these in an interview costs you the question

  • Says the firewall blocks the connection before any packets reach the server
  • Treats the application name as a field carried in the packet
  • Assumes the application name is read from the port number
  • Assumes a denied application means nothing was transferred
  • Thinks buffering every flow until identification is free

context

open as a page

Your inspection proxy's private root sits in every managed laptop's trust store: what can a key holder do, and what does it commit you to?

level: juniorimportance: must knowfreq 65%

basics

~20 s

A root your devices trust covers any hostname, not only the sites you inspect. Whoever holds its key can impersonate any site to your fleet, so you now run a certificate authority and must guard the key like one.

open as a page

An unmanaged client uses encrypted client hello through your border — what survives as evidence of its destination, and what does that fallback cost you?

level: juniorimportance: must knowfreq 60%

basics

~20 s

You keep the destination address and port, packet sizes, direction and timing, plus an outer name identifying a shared front rather than the real site. The true server name is gone, and TLS 1.3 encrypts the certificate too.

open as a page

What does a firewall's user-to-address mapping actually prove about who sent a packet?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Only that some identity source associated that address with a user earlier, and the association has not yet expired. The packet itself carries no identity, so whoever is using that address next inherits the name.

open as a page

How do you choose the timeout on a user-to-address mapping, and what breaks in each direction?

level: middleimportance: must knowfreq 62%

basics

~20 s

Set it below the shortest interval in which an address can change hands. Too long and the next holder inherits the departed user's rules and name. Too short and you evict live users into the unknown-user rule.

open as a page

An application-aware firewall re-labels a live flow from permitted web traffic to a tunnel mid-session — what has it already cost you?

level: middleimportance: should knowfreq 55%

basics

~20 s

Everything transferred under the first label. A re-label re-runs policy on a flow already in progress: the device can tear the session down from that moment, but the bytes that moved while it was called ordinary traffic are delivered and unrecoverable.

open as a page

Your TLS decryption exemption list has grown to two hundred destinations — what has that cost, and how would an intruder use it?

level: middleimportance: should knowfreq 52%

basics

~20 s

Each entry is a permanent unread path, and the list only grows because removal means proving nothing breaks. An intruder just picks a destination already covered: the bypass verdict uses the name the client claims.

open as a page

A tenant's unclassified-traffic bucket grows weekly and nobody owns an application inventory — how do you name what is in it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Rank by volume and endpoint pair, then read host role, who owns the far end, byte direction and session shape. That narrows the candidates but never names them: a flow record carries no payload. The application's owner names it; you record it.

open as a page

You add name constraints to your private inspection root so a stolen key cannot mint certificates for banking sites: where does that assurance hold, and what does maintaining it cost?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Name constraints bind the validating client, not the CA. A compliant validator rejects excluded names, and a key thief cannot strip the constraint from the copy already in your trust stores. Non-enforcing clients and unconstrained name types get nothing.

open as a page

Your inspection root's private key may have left with a departing engineer: what can they now do, and what does replacing that root cost?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They can impersonate sites to any managed device, on any network, with no warning. A root cannot be revoked, since clients never check revocation for a trust anchor, so recovery means removing it from every trust store.

open as a page

A business application moved to QUIC on UDP/443, and so did a channel you did not authorise — what are your options and what does each cost?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Three options, each with a bill: deny UDP/443 and force fallback, which breaks clients that cannot; allow it unread, making that path the preferred one; or inspect QUIC, paying capacity and state broken by connection migration.

open as a page

The user-mapping feed stalls mid-shift and entries expire: what should enforcement do?

level: seniorimportance: should knowfreq 45%

basics

~10 s

Neither blanket allow nor blanket deny. Detect the stall, freeze existing mappings instead of letting them expire, and send unattributed traffic to a designed restricted policy rather than to a catch-all nobody chose.

open as a page

Your inspection root can mint any name: what custody evidence shows a customer's auditor that no insider mints one silently, and what does it cost?

level: principalimportance: should knowfreq 30%

basics

~20 s

Controls whose failure a party outside your team would notice: a non-exportable hardware key, a witnessed generation ceremony, an issuance log held append-only by another owner, and name constraints the auditor can read for themselves. They buy detectability, not impossibility.

open as a page

You report that eighty-five percent of egress is inspected — what denominator makes that honest, and what can an adversary keep out of it?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

The denominator must cover every egress path and come from a source independent of the inspection device, or you divide what the device saw by what the device saw. Paths crossing no control are uncounted, not unread.

open as a page

On a shared terminal-server address, how do you attribute one flow to one user?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

An agent on the host gives each session its own source-port range and reports which range belongs to which user. The firewall attributes a flow by source port, not address; anything outside a range maps to nobody.

open as a page

One inspection stack serves ten tenants: what does your unclassified-traffic rule say, who signs for it, and what hides behind an allow?

level: principalimportance: nice to knowfreq 34%

basics

~20 s

The default is a risk acceptance, not a technical preference, and a provider cannot accept risk for a customer. Aim at per-tenant defaults: deny for new tenants from onboarding, a dated migration with a log-only phase for the rest, and a named tenant signature on every allow.

open as a page