An Allure 2 report is generated from a results directory in which a test's `-result.json` names an attachment `source` with no matching file. What does the generated report show for that attachment?
answer
- the pointer is resolved at generation
- the entry outlives the bytes
- generation degrades, it does not fail
- a size of zero is the tell
- usually a collection step globbing JSON
basics
~20 sThe entry survives and the bytes do not. Generation records an error for the unresolvable source and keeps an attachment carrying only its declared name and type, with a size of zero and nothing behind it.
solid answer
~50 sGeneration is where the pointer is cashed in. For each `Attachment` entry, the reader resolves `source` against the results directory; on the healthy path it mints a report-side identifier, probes the file's real content type, measures the file to fill in `size`, and the copy is placed under `data/attachments/` keyed by that source. When the file is not there — or the name fails validation, escapes the directory, or still carries a temporary suffix — the reader records an error and returns an attachment holding only the declared `name` and `type`, with `size` set to zero and no usable source. The report therefore still lists an attachment with its label; clicking it gets you nothing. **A zero size is the tell**, and the usual cause is a collection step that archived the JSON without the payload files.
go deeper
Know that the result JSON only points at payload files, so both have to be copied together. If only the JSON travels, the report still lists attachments that lead nowhere.
Explain what the reader does at generation: validate and resolve source, measure the file to fill in size, and copy the bytes into the report. Then explain which of those steps a missing file skips.
Diagnose from the symptom. A run of zero-size attachments points at the collection step, not the harness, and you should be able to name the usual culprits and the arithmetic check that confirms it.
Own the degradation policy itself: argue when a generator should render a partial report over failing loudly, and what a pipeline must surface so a silently evidence-free report is never mistaken for a complete one.
## Generation is where the pointer is cashed in Writing an attachment is cheap and local: the bytes go to a file, and an entry naming that file is appended to the test's result JSON. Nothing checks anything. The check happens later, in a different process, possibly on a different machine — at report generation, when the reader has to turn each `source` string into actual bytes. On the healthy path the reader does more than open a file: 1. It validates the `source` name against a pattern and resolves it **inside** the results directory, so a crafted `source` cannot reach out to an arbitrary path. 2. It confirms the resolved path is a regular file, and that the name does not still carry the temporary suffix a half-finished write leaves behind. 3. It probes the file for its real content type and measures it, which is where the `size` the report displays actually comes from — the writer never set it. 4. It lets the **declared** `type` from the JSON override the probed one, and mints a report-side name for the copy. 5. The bytes are copied into the generated report under `data/attachments/`, keyed by that source, so the report is self-contained and the results directory is no longer needed to read it. ## What happens when the file is not there Any of the guards failing takes the same branch. The reader records an error saying it could not find the attachment in that directory, and returns an attachment object built from what it does have: - `name` — the label the test declared. Still present. - `type` — the media type the test declared. Still present. - `size` — **zero**. - the pointer to any bytes — gone. So generation does **not** fail, and the test result is **not** dropped. The report renders an attachment with a perfectly ordinary-looking label, and there is nothing behind it. This is the design choice worth arguing about at senior level: the generator prefers a partial report over no report, which is right for a build that is already failing and wrong for anyone who assumes a rendered attachment implies retrievable evidence. ## How results directories lose their payloads This is not a rare corruption case; it has a handful of ordinary causes, and every one of them is a pipeline mistake rather than a bug in the tooling: - **A collection step that globs only JSON.** Archiving `*.json` out of a results directory takes every result file and no payload file, because payload files are named `…-attachment` plus whatever extension the media type implied. This is the most common cause by a wide margin. - **Copying files individually rather than the directory.** Any per-pattern copy that enumerates result files leaves the payloads behind. - **A run killed mid-write.** A payload file that never finished being written keeps its temporary suffix, and the reader deliberately refuses to treat such a name as a real payload. - **Sharded runs merged unevenly.** Results from several workers are combined, but one worker's payload files were never uploaded, so its JSON arrives with pointers to files that exist only on a machine that has since been destroyed. - **A cleanup that ran between the run and the generation step**, removing large binaries from a working directory that the generator had not read yet. In every case the JSON survives the trip and the evidence does not, because the JSON is small and text-shaped and the payloads are large and binary — exactly the split that any size-conscious or extension-based copy rule will fall through. ## How to catch it, and how to prevent it **Catching it after the fact:** - A rendered attachment with a **zero size** is the signature. One is a curiosity; a whole run of them is a collection bug. - Generation output is not decorative. The reader emits an error per unresolvable source, so a job that discards the generator's output throws away the only warning it was going to get. - The cheap pre-flight is arithmetic: count the `source` values across the results JSON and count the payload files present. They should match. **Preventing it:** 1. Move the **whole results directory** as one unit — archive it, upload it, restore it — rather than selecting files from inside it by pattern. 2. Generate the report as close to the run as you can, before any cleanup can touch the working directory, and move the generated report rather than the raw results when the report is what people will read. 3. If you must filter, filter in the direction that keeps evidence: match the payload glob explicitly alongside the JSON globs, remembering it needs a wildcard at both ends because some payload files have no extension at all. The underlying lesson generalises past this one format. Any results model that stores metadata and payloads as separate artefacts has a moment where the reference is resolved, and everything between the write and that moment is an opportunity to deliver half of the pair.
- Why does the reader refuse a `source` that resolves outside the results directory?Because `source` is untrusted input: it arrives in a JSON file that any harness, plugin or hand-edit can produce. Resolving it unchecked would let a crafted name pull an arbitrary file off the generating machine into a published report. The reader validates the name and requires the resolved path to stay under the results directory.
- How would you check a results directory for dangling attachments before generating a report?Count both sides. Extract every `source` value from the result and globals JSON, list the payload files actually present, and compare the two sets. A non-empty difference is a collection bug, and it is far cheaper to find there than in a report someone is trying to read during an incident.
saying these in an interview costs you the question
- Says generation fails when a payload is missing
- Assumes a visible attachment means retrievable bytes
- Thinks the missing test result is dropped entirely
- Blames the writer rather than the collection step