skip to content

Why does an operator index and stage documents on the victim's own hosts before exfiltrating?

level: middleimportance: should knowfreq 45%

answer

  1. access is not the same as knowing
  2. search is cheaper inside than outside
  3. candidate list, then one archive
  4. volumes sized to the destination
  5. three collection steps, one transfer

basics

~20 s

Holding read access is not the same as knowing what is worth taking. Enumerating and indexing inside the estate builds a candidate list cheaply, so only the small set the objective needs is archived and moved once.

solid answer

~40 s

A document tenant may hold millions of files, and broad read access says nothing about which ones matter. So the work happens where it is cheap: the operator uses directory listings and the estate's own search service to answer 'which documents mention this programme', assembles the hits into one staging location on a host that already holds documents, archives them, and often splits the archive into volumes sized to what the destination accepts in a single request. Only then does anything move. Staging turns thousands of scattered reads into one predictable transfer, and turns terabytes of repository into tens of megabytes of objective. ATT&CK separates these steps for exactly this reason - `T1074` Data Staged and `T1560` Archive Collected Data are collection-tactic decisions, distinct from the exfiltration step that follows.

code

text · 5 lines
text
Collection    T1213     Data from Information Repositories   -> tenant's own search builds the candidate list
Collection    T1074.001 Local Data Staging                   -> hits copied into one folder on a document host
Collection    T1560.001 Archive Collected Data: via Utility   -> archived, split into carrier-sized volumes
Exfiltration  T1567.002 Exfiltration to Cloud Storage         -> a single upload to a permitted destination
...

go deeper

for a junior

Recall that a theft has steps before the upload: finding what exists, deciding what matters, and collecting it in one place. Be able to say why that ordering exists.

for a middle

Explain the mechanics: enumeration and content search inside the estate, a single staging location chosen so the copy is unremarkable, archiving, and splitting into volumes sized to what the destination accepts.

for a senior

Reason about cost. Show that because the expensive step is search, the controls that bite are read-scope and search-scope, not another outbound restriction, and be able to defend that ordering.

for a principal

Be ready to argue for constraining an estate-wide search feature that people value, and to say who owns that decision and what the productivity cost buys in reduced value per compromised identity.

## The gap between access and objective An operator who has taken over an account with broad read rights, or an insider who already has them, is in the same position: they can read an enormous amount and know almost none of it. A mid-sized document estate holds millions of files under folder names invented by dozens of teams over a decade. Nothing about read access tells you which five documents describe the pricing model, the acquisition, or the design under negotiation. Closing that gap is the real work, and it happens *inside* the estate because that is where it is cheap. ## Enumerate, then index The first move is enumeration: list what exists. Directory listings, repository indexes, site and drive inventories - the platform hands these out to anyone entitled to read, because that is what a document system is for. The second move is the one that surprises people: the operator uses the estate's **own search service**. A document platform provides full-text search across everything an identity may read, and it is fast, complete and entirely ordinary to use. Content search does in a few queries what would otherwise take weeks of reading, and it produces exactly what the operator wants - a candidate list, ranked, with paths. ## Stage in one place The hits are then copied into a single staging location: one folder on a host that already holds documents, or one container in storage the estate already uses. Two reasons: 1. **It makes the transfer one event instead of thousands.** Pulling a thousand files individually from a thousand paths is slow, fragile and interruptible; a single archive is not. 2. **It puts the copy somewhere it does not look strange.** Documents accumulating on a file server are unremarkable; the same accumulation on an infrastructure host is not. ## Archive, and size the volumes to the carrier The staged set is archived - which compresses it, collapses it into one object, and incidentally makes the content opaque to anything that reads bytes in flight. Then it is frequently **split**: into volumes sized to what the chosen destination accepts in one request, or to whatever size is ordinary for that path. ATT&CK names this decision on its own (`T1030`, Data Transfer Size Limits) precisely because it is a deliberate choice about the carrier and not an accident of the archiving software. ## The point: exfiltration is the last and smallest step ``` Collection T1213 Data from Information Repositories -> the tenant's own search returns the candidate list Collection T1074.001 Local Data Staging -> hits copied into one folder on a document host Collection T1560.001 Archive Collected Data: via Utility -> archived, split into volumes the carrier accepts Exfiltration T1567.002 Exfiltration to Cloud Storage -> one upload to a permitted destination ... ``` Three of those four steps happen before a byte leaves. That is what makes the common mental picture - a huge outbound transfer as the whole of the attack - so misleading. By the time anything moves, the interesting decisions have already been made, and what moves is small because the staging step made it small. ## What raises the operator's cost here Since the expensive work is search, the controls that bite are the ones that make search expensive: narrowing what one identity may read (a smaller corpus means a smaller candidate list and more identities to obtain), and constraining estate-wide content search so a single account cannot enumerate everything it nominally may read. Neither of those is a control on the outbound path, which is the usual instinct. ## Common mistakes Saying 'they just copy everything' describes one objective - the extortion-scale sweep - and misses the espionage case entirely. Saying 'the archive is encrypted so nothing can be done' confuses making content opaque with making the operator's job cheap; the archive is a convenience, not the defence. And treating staging as an implementation detail hides the fact that it is where the operator spends most of their time on target.

  • What decides how the archive is split into volumes?
    What the destination accepts in a single request, and what size is ordinary on that path. It is a deliberate choice about the carrier - ATT&CK gives it its own identifier, `T1030` Data Transfer Size Limits - not a default of the archiving step, and it is why a transfer can appear as several ordinary-sized uploads rather than one outsized one.
  • Why stage on a host that already holds documents rather than on a jump host?
    Because the read pattern and the disk consumption are unremarkable there. A file server accumulating documents is doing its job; the same growth on a management or infrastructure host is out of character for the role, and the operator gains nothing from putting it somewhere the placement itself is odd.
  • What if the operator skips staging and pulls files directly?
    They pay for it in time on target and in fragility: thousands of individual reads from many paths, over a longer window, any part of which can be interrupted by an account expiring or a password change. Staging exists to compress that exposure into one short transfer.

saying these in an interview costs you the question

  • Assumes the operator simply copies everything
  • Treats staging as an implementation detail, not a decision
  • Thinks the archive's encryption is what defeats controls
  • Overlooks the estate's own search as the enumeration tool

context