In JMeter's HTTP Request sampler, what does ticking Retrieve All Embedded Resources do?
answer
- A checkbox that multiplies your requests
- Advanced tab, embedded resources panel
- Parser picked by response media type
- Two regex fields narrow the list
basics
~20 sJMeter parses the response for referenced images, stylesheets, scripts and frames, then issues an extra GET for every one it finds. The sampler's result becomes a container holding the page plus one sub-sample per asset.
solid answer
~50 sThe checkbox lives in the **Embedded Resources from HTML Files** panel on the HTTP Request's Advanced tab and is stored as `HTTPSampler.image_parser`. After the main response arrives, JMeter picks a parser by media type — `HTTPResponse.parsers` ships as `htmlParser wmlParser cssParser` — extracts every referenced URL and sends a plain `GET` for each one. Parsing happens only when the main sample succeeded *and* its data type is text, so a binary response never fans out — and a JSON response, though JMeter does mark it text, fans out nothing either, because no parser is registered for its media type. The extra requests become sub-samples of a container result whose byte count is the sum of them all and whose elapsed time runs to the last one. **URLs must match** and **URLs must not match** are regular expressions applied to the extracted URLs before any of them is fetched.
code
xml · 10 lines<HTTPSamplerProxy guiclass="HttpTestSampleGui" testclass="HTTPSamplerProxy" testname="Home Page">
<stringProp name="HTTPSampler.domain">shop.example.invalid</stringProp>
<stringProp name="HTTPSampler.path">/</stringProp>
<stringProp name="HTTPSampler.method">GET</stringProp>
<boolProp name="HTTPSampler.image_parser">true</boolProp>
<boolProp name="HTTPSampler.concurrentDwn">true</boolProp>
<stringProp name="HTTPSampler.concurrentPool">6</stringProp>
<stringProp name="HTTPSampler.embedded_url_re">https://shop\.example\.invalid/.*</stringProp>
<stringProp name="HTTPSampler.embedded_url_exclude_re">.*\.(?i:svg|png)</stringProp>
</HTTPSamplerProxy>go deeper
Be able to point at the checkbox and say what it costs: one sampler stops being one request. Know that the extra calls are GETs and that they appear underneath the sampler, not beside it.
Explain the gating conditions — successful sample, text data type, a parser registered for that media type — and name the two regex fields that filter the extracted URLs.
Talk about what the container sample's timings and byte counts then mean, and about the failure semantics: one bad asset fails the whole container unless you flip httpsampler.ignore_failed_embedded_resources.
Own the plan-wide call: whether every page sampler carries this flag, whether the allow and deny regexes are set once on HTTP Request Defaults, and how you keep the resulting sample counts comparable between runs.
JMeter's **HTTP Request** sampler normally sends exactly one request. Ticking **Retrieve All Embedded Resources**, in the *Embedded Resources from HTML Files* panel on the sampler's Advanced tab, turns it into a small crawler: once the main response is in, JMeter reads the body, pulls out every URL a browser would have had to fetch to render the page, and requests each of those too. In the saved `.jmx` the flag is the boolean property `HTTPSampler.image_parser`. ## Three conditions before anything is parsed JMeter only looks at the body when all of these hold: 1. the flag is on for that sampler (or inherited from an **HTTP Request Defaults** element); 2. the main sample is **successful** — a page that failed is never scanned; 3. the result's data type is **text**, which JMeter derives from the response's `Content-Type`. Then it looks up a parser for the response's media type. The mapping shipped in `bin/jmeter.properties` is: ```properties HTTPResponse.parsers=htmlParser wmlParser cssParser htmlParser.className=org.apache.jmeter.protocol.http.parser.LagartoBasedHtmlParser htmlParser.types=text/html application/xhtml+xml application/xml text/xml cssParser.types=text/css wmlParser.types=text/vnd.wap.wml ``` If the media type is in none of those lists there is no parser and nothing extra is fetched, box ticked or not. That is why a JSON API response never fans out, however many URLs its payload contains. `htmlParser.className` can be pointed at the JTidy, Jsoup or older regexp parsers instead of the default Lagarto-based one. ## How the extra requests are issued Every extracted URL is requested with **GET**, whatever method the parent sampler used, and each response is attached to the parent as a sub-sample. Downloads are **serial by default**: the same JMeter thread that ran the page runs each asset one after another. Ticking **Parallel downloads. Number:** hands them to a shared executor instead; the number field defaults to `6` (`HTTPSampler.concurrentPool`). Nesting is followed as well — a frame or iframe whose own response is HTML gets parsed in turn — but only to the depth in `httpsampler.max_frame_depth`, default `5`. Past that JMeter attaches an error sub-sample reading *Maximum frame/iframe nesting depth exceeded*. ## Narrowing what gets fetched Two fields on the same panel take regular expressions and are applied to each extracted URL before it is requested: - **URLs must match** — an allow list. Left empty, everything is allowed. - **URLs must not match** — a deny list. Left empty, nothing is denied. Worth knowing: a syntactically broken pattern is *not* an error. JMeter logs `Ignoring embedded URL allow string: …` and falls back to the field's default answer, so a broken allow expression quietly downloads everything and a broken deny expression quietly excludes nothing. ## What lands in your results Instead of one result you get a **container** sample carrying the page result plus one sub-sample per asset. The container's byte count is the sum of all of them and its elapsed time runs from the page request's start to the last asset's end. With the default `jmeter.save.saveservice.subresults=true`, each sub-sample is written to the JTL as its own row, so a page with twenty assets contributes twenty-two rows, not one. By default a failed asset drags the container down with it: JMeter sets the container unsuccessful and rewrites its response message to `Embedded resource download error: …`. Setting `httpsampler.ignore_failed_embedded_resources=true` keeps the container green and leaves the failure visible only on the child. ## Traps worth naming in an interview | Belief | What actually happens | |---|---| | "The assets use the sampler's method" | Every embedded request is a `GET`. | | "A failed page still lists its assets" | Parsing is skipped entirely when the sample failed. | | "Assets are downloaded in parallel" | Serial unless **Parallel downloads** is ticked. | | "One tick, one sample" | One tick, one container plus N children. | Two related switches sit beside it and are frequently confused with it: **Save response as MD5 hash** stores a 32-character hash instead of the body, and `httpsampler.embedded_resources_use_md5=true` does the same for embedded responses only, keeping the size and hash rather than the bytes. Neither changes which requests are sent — only what is kept.
- Which response content types does JMeter actually scan for embedded resources?Only those listed against a configured parser. The shipped `HTTPResponse.parsers` value is `htmlParser wmlParser cssParser`, covering `text/html`, `application/xhtml+xml`, `application/xml`, `text/xml`, `text/css` and `text/vnd.wap.wml`. Anything else — JSON, images, PDFs — has no parser, so no embedded requests are generated no matter what the body contains.
- What stops JMeter looping forever on a page whose frames reference each other?The `httpsampler.max_frame_depth` property, default `5`. JMeter parses nested HTML documents recursively and increments the depth each time; past the limit it stops and attaches an error sub-sample saying the maximum frame/iframe nesting depth was exceeded, rather than continuing to descend.
saying these in an interview costs you the question
- Says the embedded requests reuse the parent sampler's method
- Believes a JSON response is also scanned for URLs
- Assumes the assets are fetched in parallel by default
- Thinks each asset shows up as its own top-level sampler
- Expects a failed page to still fetch its assets