A proxy shows 40 GB to a sanctioned file-sync host and the provider's audit trail names the files — what does that prove?
answer
- bytes without payload
- names without contents
- the originals are still there
- reconcile the sizes against the volume
- an account is not a person
basics
~20 sTogether they prove a named account moved specific file objects to that host, and how many bytes left. Neither record carries contents; what personal data was involved comes from the originals still in the source repository.
solid answer
~50 sThe proxy record proves volume and destination: bytes left for that host, in that direction, on that session, under a given account or source address. Because the session is TLS, the proxy sees the connect target and the byte counts and no payload at all — it cannot tell you which files or what was in them. The provider's audit trail closes half of that gap: it names the objects retrieved, the acting account, the device and the source address, but it stores no contents either. So the join gives you `which objects` and `how much`; `what personal data` comes from a third place — the originals still in the source repository, inspected by file id. And note the direction of the actor claim: an audit event proves the service accepted a request authenticated as that account, not that the named employee was at the keyboard.
code
json · 11 lines// egress proxy - TLS session, no payload recorded
{ "ts": "2026-04-12T21:14:03Z", "user": "j.mercer", "src_ip": "10.4.18.77",
"dest_host": "sync.example-cloud.com", "method": "CONNECT",
"bytes_out": 412839551, "bytes_in": 1904722, "duration_s": 2711 }
// sync provider audit trail - names objects, stores no contents
{ "ts": "2026-04-12T21:15:40Z", "event": "file.downloaded", "actor": "j.mercer@corp",
"object_id": "f-8841203", "object_path": "/HR/onboarding/2025-cohort.xlsx",
"size_bytes": 2214880, "client": "desktop-sync/5.2", "device_id": "unenrolled-9f21",
"src_ip": "10.4.18.77" }
// ... 1,398 further file.downloaded events, same actor, 12 Mar - 24 Aprgo deeper
Be ready to say what each source records: the proxy has hosts, direction and byte counts with no payload, and the SaaS audit trail has object names and actors with no contents. Knowing what a record cannot contain is half the answer.
Explain the join across three sources and the reconciliation between summed object sizes and proxy volume, including why the numbers never match exactly and what a large mismatch actually implies.
Show that you turn the object list into a content determination from the surviving originals, that you state the retention edge rather than quietly scoping around it, and that you name the specific facts making the destination unauthorised.
Own the collection consequence: if a six-week transfer to a sanctioned host cannot be reconstructed, the retention and the object-level audit coverage on that path are the deliverable, and someone has to fund and own them before the next case.
## Three sources, three different facts A slow exfiltration through a sanctioned enterprise file-sync client is deliberately quiet: a few hundred megabytes a day into a linked account the organisation does not control, for six weeks, over a protocol and to a hostname that thousands of legitimate users hit every day. Nothing spikes, nothing is unfamiliar, and the two sources you have each answer only part of the question. **The egress proxy** records the connection, not the conversation. For a TLS session it has the connect target or the TLS server name, the client address, the authenticated user where proxy authentication is in place, timestamps, and bytes sent and received. It has **no payload**. That is not a configuration gap you can close after the fact; it is what the record is. So the proxy establishes that bytes moved, in which direction, to which host, over which period — a volume-and-destination fact. It cannot name a file and cannot describe a single byte of content. **The provider's SaaS audit trail** records the operations the service performed on named objects: the event type, the object name or id, the acting account, the client or device identifier, the source address, the timestamp, and often the size. It has no contents either — a provider logging every byte of every file would be storing a second copy of your data. So the audit trail establishes *which objects* an account touched and *when*. **The source repository** is the third source and the one candidates forget. A sync client copies; it does not move. The originals are almost always still there. Once the audit trail has given you object names or ids, you can go to the originals and determine what those specific documents contained, and therefore which categories of personal data and roughly which data subjects are in play. That is the whole shape of the determination: volume from the proxy, object identity from the SaaS trail, content from the originals. Lose any one of the three and the picture has a hole in it that the other two cannot fill. ## Reconciling the two record sets, and what a mismatch means The first analytic step is a reconciliation. Sum the sizes in the audit trail across the window and compare with the proxy's byte counts for that host. They will not match exactly — TLS and protocol overhead, chunking, retries, compression, and other legitimate traffic to the same hostname all move the number — but the *shape* should agree. A material excess of proxy bytes over accounted-for objects means something moved that the audit trail did not record, which is a scoping problem, not a rounding error. A material shortfall means the audit trail is capturing activity that did not traverse this path, which usually means another egress route or another client. Interviewers like this step because it is where a candidate shows they treat two sources as things to be cross-checked rather than two views of a single truth. ## What the records do not let you say - **They do not prove a person acted.** An audit event proves the service accepted a request bearing that account's credential from that client. Whether it was the named employee, someone with their session token, or an automation using their credential is a separate question answered from authentication and device evidence. - **They do not prove intent.** The same records are produced by a legitimate backup, a departing employee taking their own project folder, a sanctioned migration and a deliberate staging operation. - **They do not prove the contents left intact.** If the originals show the documents were encrypted, or if the sync happened at a layer where the client encrypted before upload, what reached the destination differs from what the originals say — and the location of the key becomes part of the finding. - **They do not extend past retention.** A six-week drip against a two-week proxy retention means the earliest and most important part of the window is simply gone. Say so; do not silently scope the incident to the window you happen to still have. ## The facts that move this from unusual to unauthorised Movement alone determines nothing on a sanctioned client. What upgrades the finding is context: the destination account is outside the corporate tenant; the client is linked to a personal identity; the device is unenrolled; the objects fall well outside the account's normal working set; the retrievals continue after the employee's last working day; or the source address belongs to a session the user cannot account for. Each of those is a separate check with its own source, and together they turn *data moved* into *data was taken by a party who should not have had it*. ## Where this leaves the determination At the end of the join you can normally say: these named objects were retrieved by this account to a destination we do not control, over this window, totalling roughly this volume; the originals show these objects contained these categories of personal data relating to approximately this many people; and here is the part of the window our retention no longer covers. That is a determination made on an incomplete picture — which is the normal case — and it is defensible precisely because each clause is tied to the source that actually supports it.
- The proxy shows 40 GB but the audit events only account for 26 GB. What do you do with that gap?Treat it as unaccounted movement, not as noise. Overhead and retries explain a modest excess, not 14 GB. Check whether other traffic to that hostname is being aggregated into the same records, whether the audit trail has gaps or a retention edge, and whether a second client or account used the same route. Until it is explained, the scope statement has to acknowledge that some transferred objects are unidentified.
- Could you get content evidence from the proxy by decrypting the sessions?Not retrospectively. The proxy stored byte counts and connection metadata, not payload, so there is nothing to decrypt. Even where TLS interception is deployed, proxies normally log metadata rather than retain bodies. Content evidence in this case comes from the originals in the source repository, matched by the object ids the provider's audit trail gave you.
- Why is the device identifier in the audit events worth as much as the file names?Because it separates an authorised pattern from an unauthorised one. Retrievals from an unenrolled or personal device, or from a client linked to a non-corporate account, mean the copies landed somewhere the organisation cannot reach or recall. That is what converts a movement of data into a disclosure to a party who should not have it.
A weighbridge ticket says the lorry left heavy; the loading manifest says which crates were on it. Neither says what is inside the crates — for that you open the identical ones still in the warehouse.
saying these in an interview costs you the question
- Claims the proxy can reveal file names or contents
- Assumes the SaaS audit trail stores the transferred data
- Forgets the originals still exist in the source repository
- Reads an audit actor field as proof a person acted
- Scopes the incident only to the window retention still covers