In Allure 2, a suite attaches a WEBM video to every case and the report is generated with `--single-file`. What does the generator do with those video bytes, and what does that cost every reader of the report?
answer
- one document, no lazy fetch
- binary must become text
- base64 grows it by a third
- generator holds it all in memory
- every reader pays for every video
basics
~10 sEvery video is read whole, base64-encoded and embedded in the one HTML document. Base64 inflates binary by about a third, and every reader downloads and parses all of it before seeing the first test.
solid answer
~40 sIn `--single-file` mode the generator does not write a report directory at all. It swaps its on-disk storage for an in-memory one that base64-encodes every data file it is handed, and attachments are data files like any other — so each video is read into memory whole, expanded by roughly a third by base64, and embedded in the one HTML document. The consequences compound with case count. Generation must hold the encoded set in memory at once. The output file is larger than the sum of the payloads. And because there is nothing to fetch lazily, **every reader pays for every video on open**, whether or not they watch one — where a directory report requests a payload under `data/attachments/` only when someone clicks it.
go deeper
Know that a single-file report contains the attachments themselves rather than links to them, so its size grows with everything the run attached, not just with how many tests ran.
Explain the mechanism: binary cannot sit in an HTML document, so payloads are base64-encoded and embedded, which expands them by about a third and removes any chance of fetching one lazily.
Reason about who pays and when. Every reader downloads and parses every payload before the first test name renders, and generation peak memory tracks the encoded total — so the failure mode changes with suite size, not with one file.
Own the packaging decision: say which runs may emit one file at all, what payloads such runs are allowed to carry, and how you keep a convenient distribution format from quietly becoming unusable as a suite grows.
## What single-file output changes about payloads A directory report and a single-file report describe the same run, but they answer the question "where are the bytes?" in opposite ways. In the directory form, each payload is copied out as a file of its own under `data/attachments/`, named by the attachment's source. The page fetches one when a reader opens that attachment, and never otherwise. Payload delivery is **lazy and per-reader-action**. In single-file mode there is nowhere to fetch from — the whole point is that there is one document. The generator therefore replaces its report storage with an in-memory one whose every write **base64-encodes** the incoming bytes and keeps the encoded string in a map. Attachments are handed to that storage exactly like every other data file, so a video is read into memory in full, encoded, and embedded in the document. Payload delivery becomes **eager and total**. ## Why base64 specifically, and what it costs An HTML document is text. Arbitrary binary cannot be dropped into it, so anything binary is carried as a `data:` URI whose payload is base64. That encoding is the price of admission, and it is not free: - **It expands.** Base64 represents three bytes as four characters, so encoded binary is about a third larger than the original before any markup is added. - **It cannot be skipped.** There is no branch where a payload is included by reference instead; a single-file report that omitted the bytes would not be single-file. - **It is paid three times** — once in the generator's memory, once in the file's size on disk, and once in every browser that opens it. ## Why per-case video is the pathological input Run the arithmetic in the direction of case count rather than of one file. A per-case video means the payload total scales linearly with the size of the suite, and each of those payloads is among the largest artefacts a test run produces. Single-file mode then multiplies that total by the base64 overhead and puts the product in one document. The result behaves qualitatively differently from a directory report of the same run: | | directory report | single-file report | |---|---|---| | where payload bytes live | one file each, under `data/attachments/` | inline in the one document | | encoding | none — raw bytes | base64, roughly a third larger | | when a reader pays | on clicking that attachment | on opening the report | | generator memory | streams file to file | holds the encoded set at once | | what a reader who watches nothing pays | nothing | everything | That last row is the one that matters in practice. A reader who wants to know *which* tests failed — by far the commonest reason anyone opens a report — has to download and parse every video first. The cost is borne by every reader on every open, for evidence almost none of them will look at. ## Diagnosing it The symptoms are recognisable and often misattributed: 1. A single HTML file whose size is far larger than the run's raw payloads, by roughly the base64 ratio plus markup. 2. A browser tab that stalls or runs out of memory before the report renders anything, with no partial view — the document must be parsed before the first test name appears. 3. Generation that succeeds on a small suite and fails on the full one, because peak memory tracks the encoded total rather than the largest single payload. A useful reflex when someone reports "the report will not open": ask whether it is one file or a directory before looking at anything else. The two forms fail in completely different ways, and the fix for one is not the fix for the other. ## Choosing between the forms Single-file output earns its keep when the report has to travel as one attachment or artefact and the payloads are small and text-shaped — a stack trace, a request body, a short log. It stops being viable as soon as the payload total is large, because the format has no mechanism to defer any of it. So the decision is about **what the report carries**, not just how it is packaged: - If the report must be one file, keep heavy media out of it and let the small text payloads travel inline. - If heavy media is the point, keep the directory form, where the payload is fetched only by the person who wants to watch it. - Deciding a run needs a video at all is a harness question. What is settled here is only what the report does with those bytes once they exist — and one-file output does the most expensive possible thing with them. One version note, because it changes what you are measuring: single-file output exists in both Allure majors, but the in-memory base64 storage described above was measured on Allure 2. Allure 3 reaches one-file output through its own report plugins, so establish which major produced a file before reasoning about why it is the size it is.
- Why does generation memory become a problem before the file size does?Because the in-memory storage holds the base64 of every data file at once before anything is written out. Peak memory therefore tracks the encoded total for the whole run, not the largest single payload, so a suite that generated fine at half its size can fail outright once the payload total crosses what the generating process has available.
- Would attaching the same video once at run level instead of per case fix a single-file report?It fixes the multiplication, not the mechanism. One video is one payload rather than hundreds, so the document shrinks by orders of magnitude. But that single video is still read whole, base64-encoded and embedded, and every reader still pays for it on open. The eager, total delivery is a property of the format, not of where the pointer lives.
saying these in an interview costs you the question
- Thinks single-file mode links out to payload files
- Assumes base64 is roughly size-neutral
- Says only readers who watch a video pay
- Believes the browser can render before parsing everything