In a file-summarising pipeline, why does an attacker pick structure over wording?
answer
- what does the splitter count?
- unit count is set before any model reads
- nothing for a content screen to match
- shape decides how many calls, not phrasing
basics
~20 sStructure decides the unit count, wording does not. Part, sheet and embedded-object counts and nesting depth set how many calls one submission becomes, and a structural file carries no directive span for a content screen to score.
solid answer
~50 sThe multiplier lives in the ingest stage, and the ingest stage reads structure. How many parts an extractor emits, how many sections or sheets it finds, how deeply objects nest, how many embedded artefacts it must open — those properties decide how many units of work a submission becomes, and each unit is at least one model call before aggregation and any verification pass adds more. Wording buys nothing on that axis: there is no instruction for a model to obey here, and the objective is spend rather than behaviour. Choosing structure also gets past a different obstacle for free. A screen that scores text for directive or unsafe content has nothing to match on a file whose only unusual property is its shape, and a per-request token cap sees each derived call as compliant because each one is.
code
json · 13 lines{
"submission_id": "s-4417",
"accepted_bytes": 812443,
"derived_calls": [
{"stage": "extract", "unit": "part-001", "input_tokens": 3910, "output_tokens": 512, "cap_exceeded": false},
{"stage": "summarise", "unit": "part-001", "input_tokens": 4102, "output_tokens": 498, "cap_exceeded": false},
{"stage": "summarise", "unit": "part-001", "attempt": 2, "input_tokens": 4102, "output_tokens": 501, "cap_exceeded": false},
"... 213 further rows ...",
{"stage": "aggregate", "unit": "whole", "input_tokens": 7788, "output_tokens": 903, "cap_exceeded": false}
],
"calls_total": 217,
"tokens_total": 1046118
}go deeper
Know that a summarising pipeline splits a submission into parts and calls a model per part, so the number of parts is where the cost comes from.
Be ready to trace unit count from artefact properties through extraction, per-unit calls, aggregation and retries, and to say why a text-scoring screen has nothing to match here.
Show how you would establish the claim from usage records joined to a submission, and state clearly what a single run's total does and does not prove about the method.
Be able to say which stage owns the multiplier and what a change there actually achieves — a bound at ingest relocates the exposure to submission count rather than ending it.
## Two different levers on the same pipeline A pipeline that summarises uploaded documents has a wording surface and a structure surface, and they are attacked for different payoffs. The wording surface is where somebody writes text intended to be read as instruction — that is the injection family, aimed at changing what the assistant does. The structure surface is where somebody chooses an artefact's shape so that the pipeline does more work. The second is the lever when the objective is consumption, and it is a better lever for that objective for three separate reasons. ## Reason one: structure is what the unit count is a function of An ingest stage has to decide how many pieces a submission becomes before any model sees anything. That decision is mechanical and it reads properties of the artefact: the number of pages, sheets, slides, sections or records; the number of embedded objects an extractor must open and process in turn; how deeply containers nest, when each level of nesting produces its own parts; how many distinct extraction routes a mixed-format artefact triggers. Each emitted unit becomes at least one model call. Then the loop's own stages multiply what ingest produced: an aggregation pass over the partial summaries, a further pass that checks the aggregate against its sources, and a regeneration whenever a stage's output does not satisfy the next stage's expectations. A submission whose parts reliably produce output the next stage rejects therefore costs more than one call per unit, and the extra calls are indistinguishable from ordinary work. Wording changes none of this. The same sentence repeated across a thousand parts and a thousand different sentences cost the same, because the count came from the shape. ## Reason two: there is nothing to score An input screen that classifies submitted text — for directive content, for unsafe content, for anything the deployment cares about — is looking at what the text says. A submission chosen for its shape can be entirely mundane prose, or near-empty parts, and score exactly as an honest document does. The property that makes it expensive is not expressible as a phrase, so a content screen is not the obstacle it gets past by accident; it is simply orthogonal to it. The per-request token cap is orthogonal in the same way. Every derived call is genuinely small. Compliance per call is the pipeline working as designed. ## Reason three: it does not depend on a model's behaviour Injection-style constructions depend on a model choosing to follow a span, which varies between deployments, between versions and between runs. A structural submission's cost is decided by deterministic code — the extractor, the splitter, the scheduler. That makes the method more stable across deployments and easier to reason about when someone asks what it will still be worth later. ## Reading the usage record The evidence for this class is a set of usage rows for one submission rather than any single row. What the rows show is a count and a total; what no individual row shows is anything unusual at all. If you are asked in an interview what you would look at, the honest answer is the join key — every derived call carrying the identity of the submission that caused it — because without it the calls are just calls. Be careful with two directions of inference. Rows all under the cap prove size compliance and nothing about count. And a high total for one submission proves what that run cost, not what the method costs, because retry counts vary between runs of a probabilistic pipeline. ## Where the structural lever stops It stops at the stage that fixes the unit count independently of the submission. If the extractor emits a bounded number of parts whatever it is handed, the multiplier is gone from that stage. What it does not do is disappear: a bound per submission relocates the question to the number of submissions, which for an anonymous submitter is a different and often cheaper axis. This is why an honest answer distinguishes removing a multiplier from moving it, and why the interesting follow-up is always which stage the fan-out ends up living in next.
- Where do retries fit into the multiplier, and why are they hard to see in the numbers?A retry is an ordinary call with ordinary token counts, so it looks like work rather than waste. Content whose parts reliably produce output the next stage cannot use turns one unit into two or three calls, and only an attempt counter joined to the submission distinguishes that from a document that simply had more parts.
- If the extractor is changed to emit a bounded number of parts per submission, is the method finished?That stage's multiplier is finished. The work per submission is now bounded, so the axis moves to how many submissions can be made, which for an anonymous submitter is often the cheaper axis anyway. The honest claim is that the multiplier moved rather than disappeared.
- Why is a structural submission more stable across deployments than a wording-based one?Its cost is produced by deterministic code — an extractor, a splitter, a scheduler — rather than by a model choosing to follow a span. Model behaviour varies between versions and between runs; a page count does not.
saying these in an interview costs you the question
- Assumes the file's byte size determines the work
- Looks for a directive span that is not there
- Says a content screen would catch it
- Confuses this with injection because both use uploads
- Treats a bound per submission as removing the multiplier