In Selenium 4, what does WebElement.getScreenshotAs capture compared with calling it on the driver?
answer
- Two receivers, two different endpoints
- The receiver decides the region
- No cast needed on an element
- Take Element Screenshot, bounding rectangle
- OutputType.FILE is only temporary
basics
~10 sCalling getScreenshotAs on a WebElement returns a PNG cropped to that element's bounding rectangle. The driver-level call captures the page's whole visible area instead, so the receiver you choose decides the region.
solid answer
~40 sIn Selenium 4 the `WebElement` interface extends `TakesScreenshot`, so you call `getScreenshotAs` straight on an element with no cast; only the driver-level call needs `((TakesScreenshot) driver)`. The element call issues the W3C `Take Element Screenshot` command, `GET /session/{session id}/element/{element id}/screenshot`, and the remote end answers with a Base64 PNG of the visible region inside that element's bounding rectangle, after scrolling it into view. The `OutputType` you pass decides what the client hands back: `BASE64` for the raw string, `BYTES` for decoded PNG bytes, `FILE` for a temporary file that is deleted when the JVM exits, so copy it if you want to keep it. On a customs declaration form, cropping to the duty summary panel gives a reviewer one readable artefact instead of a full page they have to hunt through.
code
java · 15 linesimport java.nio.file.Files;
import java.nio.file.Path;
import org.openqa.selenium.By;
import org.openqa.selenium.OutputType;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
class DutySummaryEvidence {
static Path capture(WebDriver driver, Path target) throws Exception {
WebElement dutySummary = driver.findElement(By.id("duty-summary"));
byte[] png = dutySummary.getScreenshotAs(OutputType.BYTES);
return Files.write(target, png);
}
}go deeper
Be ready to name the method and say which region each receiver captures. Knowing that an element needs no cast, and that OutputType.FILE hands back a temporary file, is enough at this stage.
Explain which endpoint each receiver hits and what the remote end does before it paints: it resolves the reference, scrolls the element into view and reads its bounding rectangle. Say why BYTES is usually the cleaner output.
Show that you crop because a human has to read the artefact. Be ready to say what a crop still contains, such as overlays and partial content, and how the bytes get stored somewhere that outlives the JVM.
Own the shape of the shared capture helper: which receiver it accepts, which OutputType it returns and how it names files, so artefacts across suites are readable without a private convention per team.
## The two commands behind one method name Selenium's `TakesScreenshot` interface declares a single method, `getScreenshotAs(OutputType<X>)`, and **two different receivers implement it**. When the receiver is the driver, the client sends the W3C WebDriver **Take Screenshot** command, `GET /session/{session id}/screenshot`. When the receiver is a `WebElement`, it sends **Take Element Screenshot**, `GET /session/{session id}/element/{element id}/screenshot`. Same method name, different endpoint, different picture. In Selenium 4 the `WebElement` interface is declared as extending `SearchContext` and `TakesScreenshot`, so the element call needs **no cast**: - `panel.getScreenshotAs(OutputType.BYTES)` compiles directly on any `WebElement`. - `((TakesScreenshot) driver).getScreenshotAs(OutputType.BYTES)` needs the cast, because the `WebDriver` interface itself does not declare the method. That asymmetry is the most common surprise for someone meeting the API for the first time, and it is a cheap thing for an interviewer to check. ## What each call puts in the file | | driver receiver | `WebElement` receiver | |---|---|---| | Command | Take Screenshot | Take Element Screenshot | | Endpoint | `/session/{id}/screenshot` | `/session/{id}/element/{id}/screenshot` | | Region | the top-level browsing context's visible area | the visible region inside the element's bounding rectangle | | Scrolls first | no | yes, the remote end scrolls the element into view | | Java receiver | `(TakesScreenshot) driver` | the `WebElement` itself | Both return a **lossless PNG**, Base64-encoded on the wire; the client decodes it for you before you ever see it. ## The wire sequence, step by step 1. You already hold an element reference produced by a `findElement` call. 2. `RemoteWebElement.getScreenshotAs` issues the `elementScreenshot` command carrying that element id. 3. The remote end resolves the reference, scrolls the element into view, waits for the next animation frame, reads the element's rectangle and paints that region of the framebuffer. 4. It encodes the result as a Base64 PNG string and returns it in the response value. 5. The `OutputType` you passed converts that string into the shape you asked for. Nothing in that sequence is done on the client. The crop is decided by the remote end; the Java side only chooses the container. ## Choosing an `OutputType` - **`OutputType.BASE64`** returns the Base64 string exactly as the remote end sent it. Best when you are about to embed the image in an HTML report. - **`OutputType.BYTES`** returns the decoded PNG bytes. Best when you will write them yourself with `Files.write`, or hand them straight to an image library. - **`OutputType.FILE`** returns a `java.io.File` created in the system temp directory and registered with `deleteOnExit`. It is **not** a durable artefact: copy it somewhere you own before the JVM ends, or the report links to a file that no longer exists. A crop of a customs declaration form's duty summary panel is a few kilobytes. The full page at the same moment is far larger and still leaves a reviewer scrolling to find the panel that mattered, which is the whole reason the element-level call exists. ## What the element call does not do for you - It does **not** locate the element. You need a live reference before you can crop to it, and if that reference is dead the command fails rather than returning an empty image. - It does **not** guarantee the element has finished rendering. The paint happens on the next animation frame, which is not the same thing as content having settled. - It does **not** ignore what sits on top. The pixels come from the framebuffer, so a modal, a sticky header or a consent banner overlapping the commodity-code row appears in the crop exactly as a user would see it. - It does **not** reach outside the viewport. The painted region is clamped to the visible area, so an element taller than the window comes back truncated rather than stitched. - It does **not** turn the image into a reference to compare against. Here the crop is evidence of a failure that a human reads; deciding whether it matches an approved image is a separate discipline with its own tooling. ## Why interviewers ask it The question separates a candidate who has only ever pasted a driver-level screenshot helper from one who has thought about the artefact's reader. On a customs declaration form with forty fields, a page image proves the run reached the page; a crop of the invalid commodity-code row proves what went wrong. Knowing that the receiver, and only the receiver, decides which of those you get is the point of the question.
- Why does the driver-level call need a cast in Java when the element-level call does not?The `WebDriver` interface does not declare `getScreenshotAs`; the method lives on `TakesScreenshot`, which the concrete drivers implement, so the compiler needs `((TakesScreenshot) driver)`. `WebElement`, by contrast, is declared as extending `TakesScreenshot`, so the method is already on the static type of every element reference you hold.
- What happens to the file returned by OutputType.FILE if you never copy it?It is created in the system temp directory and registered with `deleteOnExit`, so it survives only until the JVM that took it ends. A run that records the path in a report and then exits leaves a link to a file that no longer exists, so write the bytes into a directory you own instead.
saying these in an interview costs you the question
- Says the element call returns the page and the client crops it afterwards
- Believes a WebElement must be cast to TakesScreenshot before capturing
- Thinks OutputType.FILE writes a permanent file the run can keep
- Assumes an overlapping banner is excluded because the element was the receiver
- Claims both receivers hit the same WebDriver screenshot endpoint