Your JMeter plan ticks Retrieve All Embedded Resources on every page sampler. What does the injector pay?
answer
- Parsing happens where the sampling happens
- One page sample is many results
- Bodies are kept unless you say otherwise
- A digest keeps the timing, drops the payload
basics
~20 sEvery HTML response is parsed on the sampling thread, and each asset becomes another request whose result hangs off a container result JMeter builds for the page. That container can carry dozens of results, and every listener sees all of them.
solid answer
~60 sThe tick turns one request into a small crawl, and the crawl runs on the injector. The response bytes are handed to a link-extractor parser on the **sampling thread**, so every page response costs HTML parsing before the sample completes. Each URL found is filtered against *URLs must match* / *URLs must not match* if you set them, then fetched - serially on the same thread, or through a JVM-wide downloader whose pool grows on demand when **Use concurrent pool** is ticked. Each asset's result is attached as a **sub-result** of a *container* result JMeter builds for the page - `httpsampler.separate.container` defaults to `true`, so the page's own result becomes the container's first sub-result and the assets hang off the container beside it - and a twenty-asset page is twenty-two results flowing to every listener and writer. Bodies are kept: `httpsampler.max_bytes_to_store_per_request` defaults to `0`, meaning no truncation. Setting `httpsampler.embedded_resources_use_md5=true` stores a 32-character digest for embedded results instead - but that property is read in exactly one place, the `ASyncSample` constructor, so it only bites on the concurrent-download path: **Use concurrent pool** ticked with a size above 1. The serial default never consults it.
code
properties · 7 lines# bin/jmeter.properties
# embedded-resource results store the 32-char MD5 of the body, not the body
# read only on the concurrent-download path - Use concurrent pool, size > 1
httpsampler.embedded_resources_use_md5=true
# shipped default: 0 means responses are not truncated at all
#httpsampler.max_bytes_to_store_per_request=0go deeper
Know that the option makes JMeter fetch the page's assets too, and that this multiplies both the requests made and the results produced. Do not tick it by reflex on every sampler.
Explain the sequence - parse, filter, fetch, attach as sub-results - and say where each step runs, so it is clear the parsing is on the sampling thread rather than somewhere free.
Diagnose an injector that saturates before the target does: check this box alongside retaining listeners, then reach for URL filters and MD5 storage rather than turning the feature off blindly.
Decide when modelling a full page load is worth the generator capacity, and set the convention for which hosts a load plan is allowed to fetch from at all.
## What the tick actually makes the thread do **Retrieve All Embedded Resources from HTML Files** turns one request into a small crawl, and every step of that crawl is work on the injector, inside the sampling thread's own loop. 1. **Parse.** When the response comes back, the sampler hands the response bytes to a link-extractor parser, which walks the HTML looking for images, scripts, stylesheets, applets and frames. This is CPU work on the same thread that just made the request, and it happens for every page response, every iteration, on every thread. 2. **Filter.** Each discovered URL is tested against the *URLs must match* and *URLs must not match* regular expressions, if you set them. No expression means no filter: everything the parser found gets fetched. 3. **Fetch.** Each surviving URL becomes another real HTTP request. With **Use concurrent pool** ticked, they are submitted to a JVM-wide downloader whose maximum pool size is effectively unbounded — the pool grows on demand and the threads it creates are named `ResDownload-...`. With it unticked, they are fetched one after another on the sampling thread itself. 4. **Attach.** Because `httpsampler.separate.container` defaults to `true`, JMeter wraps the page in a new *container* sample result, makes the page's own result the container's first sub-result, and adds every asset's result as a further sub-result of that container. (Set the property to `false` and you get the pre-51939 shape, where the assets hang directly off the page sample.) ## The results it manufactures Step 4 is the one people miss. The container now carries a child result per asset alongside the page's own result, and the whole object travels together: to every listener in scope, to the result writer, into whatever a listener retains. A twenty-asset page is not one sample result flowing through the plan; it is twenty-two — the container, the page itself, and the twenty assets. That has a specific consequence for anything that keeps samples. The View Results Tree's entry cap counts *main* samples, so 500 buffered entries on a twenty-asset page plan means roughly ten thousand retained results, each holding its own response bytes. ## The memory the bodies occupy By default JMeter does not truncate response data at all: the HTTP sampler's `httpsampler.max_bytes_to_store_per_request` defaults to `0`, meaning *do not truncate*. So the bytes of every fetched font, image and bundle are held in the sample result for as long as anything references it. There are three cheap levers — two that shrink what a result stores and one that stops the fetch happening at all: | Lever | Where it lives | Effect | |---|---|---| | `httpsampler.embedded_resources_use_md5` | `bin/jmeter.properties`, default `false` | Embedded-resource results store the 32-character MD5 of the body instead of the body — but only on the concurrent-download path; the serial default, and a pool sized 1, never read this property | | **Save response as MD5 hash** | A checkbox on the HTTP Request sampler | Same trade for that sampler's own response — the manual calls it *"intended for testing large amounts of data"*. Not a substitute here: on a page sampler it digests the HTML the parser needs, so nothing is discovered to fetch | | **URLs must not match** | A field on the HTTP Request sampler | Never fetch the assets in the first place | The MD5 options are the interesting ones because they keep the request, the timing and the byte count — the sampler still records the size of the original response — while dropping the payload you were never going to read. ## When to leave it on, and when not Leave it on when the thing you are measuring is genuinely a page load, and when the assets are served by the system under test rather than by a CDN you have no interest in loading. Turn it off, or fence it with *URLs must not match*, when: - the assets are static and cached in production, so fetching them every iteration models nothing real; - they are served from third-party hosts you should not be load testing at all; - the injector, not the system under test, is the thing running out of room — parsing HTML at rate is real CPU, and it is CPU spent on the generator. The diagnosis that matters in a review is simple: if a plan ticks this box on every sampler and the injector saturates before the target does, this is one of the first three things to look at, alongside retaining listeners and per-sample scripting. It is also the cheapest to fix, because switching the embedded-resource results to MD5 storage changes no timings and no counts.
- What do you lose by switching embedded-resource results to MD5 storage?The bodies, and nothing else - provided the plan is on the concurrent-download path, which is the only one that reads the property. The request is still made, the timing is still recorded, and the byte count still reflects the size of the original response - the sampler substitutes a 32-character digest for the payload. You lose the ability to inspect what an asset returned, which for images, fonts and bundles is rarely what you were going to look at anyway.
- How would you cut the fetch set without turning embedded-resource retrieval off entirely?Use the URLs must match and URLs must not match fields on the HTTP Request sampler. Excluding third-party hosts you have no business load testing, and asset types you do not care about, removes both the requests and the results they would have produced. That is cheaper than any retention setting, because the work is never done.
- Why does a retaining listener get much more expensive once embedded resources are on?Because the sub-results travel inside the parent. The View Results Tree's entry cap counts main samples, so a 500-entry buffer on a twenty-asset page plan is holding roughly ten thousand results, each with its own response bytes. The listener looks bounded and is not, in the dimension that matters.
saying these in an interview costs you the question
- Assumes the asset fetches are free because the server is fast
- Thinks the HTML parsing happens off the sampling thread
- Believes only the parent sample reaches the listeners
- Says the concurrent pool Size caps how much is stored
- Leaves it on for third-party hosts nobody is testing