A tenant's unclassified-traffic bucket grows weekly and nobody owns an application inventory — how do you name what is in it?
answer
- rank by bytes, not by arrival
- endpoint pair is the unit of work
- direction and time shape narrow it
- counts move, content does not
- the owner names it, you record it
basics
~20 sRank by volume and endpoint pair, then read host role, who owns the far end, byte direction and session shape. That narrows the candidates but never names them: a flow record carries no payload. The application's owner names it; you record it.
solid answer
~50 sDo not work the bucket alphabetically or by alert time — rank it by bytes and by distinct endpoint pair, because a handful of pairs will be most of it. For each group ask four things of the flow record: the internal host's role, who owns the external endpoint, which direction the bytes go and how asymmetric, and whether the session is long, short or periodic. That narrows it: a nightly outbound bulk transfer to an unowned endpoint is a small set of candidates. It does not name it, because a flow record carries the five-tuple, counts and timestamps and no payload at all, so it cannot separate a backup job from an exfiltration channel. The name comes from the tenant's application owner. Then record it as a scoped definition or an allow keyed to that endpoint pair, with an owner and a review date, so the bucket shrinks permanently.
code
text · 10 linessrcAddr 10.42.7.19 (tenant B, finance segment)
dstAddr 203.0.113.44 (external, no owner on record)
proto/dstPort TCP/9021
app-label unclassified-tcp <- classifier never committed
octets c->s 412,900,112
octets s->c 1,204,338
packets 318,004
flowStart 02:14:07 flowEnd 04:51:52 (repeats nightly)
...
payload -- not carried by this record --go deeper
Know that traffic no classifier can name lands in an unclassified bucket, and that a flow record tells you who talked, how much and when, but never what was said.
Explain the triage: group by endpoint pair, rank by volume, read direction and time shape, and understand why a custom protocol is unclassified by construction rather than by mistake.
Demonstrate the whole loop, including the part engineers skip: narrow with evidence, get the name from the application owner, and write it down as a scoped definition with an owner and an expiry so the bucket actually shrinks.
Own the fact that this queue has no service rate. Argue for bounding the inflow, reporting the bucket as a metric with a name against it, and treating an unanswered ownership question as a contract issue rather than a backlog item.
## What the bucket actually is Every classification engine has a residual label for traffic no decoder and no heuristic would commit to — unclassified TCP, unclassified UDP, unknown. It fills from three quite different sources, and the whole exercise is telling them apart: - **Bespoke and internal applications** that no signature was ever written for. On a shared boundary in front of customers whose software you did not build, this is most of it. - **Ordinary applications seen from an unhelpful vantage** — a flow that was already in progress when the device saw it, or one whose distinguishing exchange happened elsewhere. - **Something written by an adversary.** This is the awkward part: anything with a custom protocol is unclassified *by construction*, so a permissive unknown rule is a hole shaped exactly like a purpose-built channel. The defender's price here is not capacity — it is a review burden that grows every week and that, at a shared boundary, has an owner on neither side of the contract. ## Triage order Rank by bytes and by distinct endpoint pair. Buckets like this obey a steep distribution: a small number of pairs are most of the volume, and clearing those is most of the work done. Ranking by first-seen time, or working an alert queue, spends the same effort on a one-off connection as on a nightly transfer. ## What flow evidence can tell you For each candidate group, four questions: 1. **What is the internal host?** A database server, a build agent, a finance workstation and a printer produce very different priors for the same flow. 2. **Who owns the far end?** Registration, hosting, whether other tenants also talk to it, whether it is internal at all. 3. **Which way do the bytes go, and how asymmetric?** A heavily outbound-skewed transfer means something different from a symmetric interactive session. 4. **What is the time shape?** Long and nightly, short and constant, or beaconing at a fixed interval are three different stories. | Evidence in a flow record | What it can establish | What it cannot | |---|---|---| | Five-tuple | which parties talked | who the user was | | Byte and packet counts | that data moved, and which way | what the data was | | Start, end, duration | the time shape and periodicity | whether it was authorised | ## What flow evidence cannot do, and you must say this A flow record carries no payload. It can show you a 400 MB outbound transfer to an endpoint with no owner you recognise, every night at 02:14, and that record is **equally consistent with a backup job somebody configured last month and with an exfiltration channel**. There is no amount of staring at counters that resolves it. The direction of the claim matters: the record proves bytes moved, not what they were. ## Where the name comes from The tenant's application owner. The provider does not own the inventory and cannot invent it, so the process is: narrow with flow evidence, produce a short, specific question — *"host X talks to endpoint Y nightly, 400 MB outbound, is this yours?"* — and route it to the one person who can answer. A question this specific gets answered; "please send us your application inventory" does not. ## Make the answer stick An answer you do not record is one you will re-derive next month. Each resolved group becomes either a per-tenant definition for that traffic, or an allow keyed to that specific endpoint pair — never a blanket allow for unclassified traffic on that host, which converts a narrow answer into a wide hole. Attach a named owner and a review date to each, because these are exactly the entries that outlive the application they were written for. ## Cap the inflow, not just the backlog The bucket is a queue with no service rate, so the honest engineering answer includes bounding it: report *new* unclassified endpoint pairs per tenant rather than every flow, publish the bucket's age and volume as a metric so it has a visible owner, and make onboarding of a new tenant include a first pass over their unknown traffic while somebody on their side is still paying attention. Without that, the leftover is the part an adversary can sit in, because nobody is reading it.
- Looking at that nightly 400 MB outbound flow, what can you conclude and what can you not?You can conclude that a finance-segment host moves a large, one-directional volume to a single external endpoint on a nightly schedule, and that the classifier could not name the protocol. You cannot conclude what the bytes were, whether anyone authorised it, or which user or process produced it — the record carries counts and timestamps and no payload.
- Why not just allow unclassified traffic from hosts whose owners have confirmed one flow is legitimate?Because that converts a specific answer into a general one. The owner confirmed a destination and a schedule, not every future unnamed protocol that host might emit. Key the allow to the endpoint pair; a host-wide allow for unclassified traffic is the same hole, just with a smaller sign on it.
- The tenant will not answer your questions. What then?Narrow further and make the ask cheaper — one line naming the host, the endpoint, the schedule and the volume, addressed to the owner of that host rather than to the account. If it still goes unanswered, the flow stays in the bucket with the non-response recorded, and that record is what turns an engineering problem into a contract conversation about who owns unclassified traffic.
saying these in an interview costs you the question
- Claims a flow record can show what the transferred data was
- Works the bucket by arrival order instead of by volume and endpoint
- Calls a large nightly outbound transfer malicious with no owner check
- Resolves a flow verbally and never records a definition or an allow
- Fixes one flow with a host-wide allow for unclassified traffic