skip to content

In Selenium 4, why does driver.getScreenshotAs() capture only the visible viewport rather than the whole page?

level: middleimportance: should knowfreq 72%

answer

  1. Look at what the spec samples
  2. Framebuffer, not document
  3. The rectangle gets clamped somewhere
  4. Paint width comes from the viewport
  5. Top-level context, one screen, PNG

basics

~20 s

Because the W3C Take Screenshot command dumps the visual viewport's framebuffer. Its algorithm clamps the painted rectangle to the viewport's width and height, so however tall the document is, the returned PNG is one screen.

solid answer

~40 s

`getScreenshotAs` sends the W3C **Take Screenshot** command, `GET /session/{session id}/screenshot`, which the specification defines as dumping a snapshot of the visual viewport's framebuffer as a lossless PNG returned base64-encoded. The remote-end steps do start from the document element's rectangle, but they then run "draw a bounding box from the framebuffer", and that algorithm derives paint width and paint height from the *visual viewport's* dimensions — so the canvas can never be taller than the window. It is also bound to the current **top-level** browsing context, which means switching into an iframe does not narrow or change the shot. In Selenium 4 there is no standard full-page mode: the routes are Firefox's `getFullPageScreenshotAs`, scroll-and-stitch, or `PrintsPage.print` for a PDF.

go deeper

for a junior

Recall the headline: the standard screenshot is one screen, not the whole page. Knowing that much stops you shipping evidence that misses the element your test was actually asserting on.

for a middle

Explain the mechanism, not just the outcome: Take Screenshot samples the visual viewport's framebuffer, and the bounding-box algorithm derives paint width and height from the viewport, clamping the document rectangle.

for a senior

Demonstrate that you have debugged this. Talk about capturing at the moment of failure with the element on screen, the top-level-context surprise inside frames, and device-pixel scaling breaking crop maths.

for a principal

Own the policy question: what a run's images have to prove, and whether the cost of whole-page capture across the suite is justified compared with targeted viewport evidence.

## The command behind `getScreenshotAs` When a Selenium 4 test calls `((TakesScreenshot) driver).getScreenshotAs(...)`, `RemoteWebDriver` sends `DriverCommand.SCREENSHOT`, which is the W3C WebDriver **Take Screenshot** command: ``` GET /session/{session id}/screenshot ``` The specification describes screenshots as a diagnostic mechanism that works by **dumping a snapshot of the visual viewport's framebuffer as a lossless PNG image**, returned to the local end as a base64-encoded string. That one sentence is the whole answer: the source of the pixels is the framebuffer — what the compositor has actually painted — not the document. ## Why the whole document never arrives The remote-end steps look, at first glance, as if they should capture everything. They compute **root rect** as the current top-level browsing context's *document element's* rectangle, which for a 40-page merger agreement in a redaction tool is thousands of CSS pixels tall. Then they call the algorithm named **draw a bounding box from the framebuffer**, and that algorithm clamps: 1. if the visual viewport's width or height is `0`, it fails with the error code `unable to capture screen`; 2. **paint width** becomes the visual viewport's width minus `min(rect x, rect x + rect width)`; 3. **paint height** becomes the visual viewport's height minus `min(rect y, rect y + rect height)`; 4. a canvas of exactly `paint width` × `paint height` is created and the framebuffer region is drawn onto it; 5. the canvas is serialised to PNG and base64-encoded. Steps 2 and 3 start from the *viewport's* dimensions, not the rectangle's. However tall the document rect is, the canvas can never exceed the viewport. A request that asks for the whole page comes back cropped to the window. ## What a redaction reviewer actually sees | the assumption | what the spec does | |---|---| | captures the full document top to bottom | clamps the paint rectangle to the visual viewport | | captures whatever context the driver switched into | always the current **top-level** browsing context | | the format is negotiable | always a lossless PNG, base64-encoded | | scroll position is irrelevant | the framebuffer is whatever is painted right now | Concretely, on a page showing a long agreement with a sticky redaction toolbar: - the image contains the toolbar and roughly one screen of clauses, and nothing else; - if the assertion that failed was about clause 47, three screens down, **the evidence PNG proves nothing** — it shows the top of the document; - because the framebuffer is sampled as painted, the current scroll offset silently decides the content of your evidence. ## The frame surprise The spec binds Take Screenshot to the session's *current top-level browsing context*. If the redaction tool renders the document inside an `iframe` and the test has switched into that frame, `getScreenshotAs` still returns the top-level viewport — the switch does not narrow the shot. Candidates routinely expect the opposite, and it is the fastest way to waste an hour wondering why the image looks unchanged. Two error codes are worth recognising: - **`no such window`** — the top-level browsing context is no longer open, typically because a `close()` ran before the capture; - **`unable to capture screen`** — a zero-dimension viewport, or a bitmap the browser refuses to serialise. ## Getting more than the viewport The spec deliberately does not offer a full-page mode, so every route is a workaround: - **A vendor extension.** Firefox exposes a non-standard full-page command through `HasFullPageScreenshot.getFullPageScreenshotAs(OutputType)`; there is no equivalent on Chromium drivers in Selenium 4. - **Scroll and stitch.** Take one shot per viewport-height scroll step and join the slices yourself, accepting the seams that produces. - **Print Page.** `PrintsPage.print(PrintOptions)` is a W3C command (`POST /session/{session id}/print`) that paginates the whole document — but it returns a `Pdf`, not an image. - **A bigger window.** Enlarging the browser enlarges the viewport and therefore the shot; how you size the window is a separate concern from how you capture it. ## The practical discipline Because scope is decided by the command and not by your `OutputType`, an evidence screenshot is only useful if the thing you want to prove is on screen when you take it. In a redaction suite that usually means capturing immediately at the point of failure, in the same run, with the element of interest already scrolled into view — and being explicit in the report that the PNG is one viewport, not the document.

  • A test switches into an iframe and then takes a screenshot. What does the image show?
    The top-level browsing context's viewport, unchanged by the switch. The W3C Take Screenshot command is defined against the session's current top-level browsing context, not the frame the driver is switched into, so the frame switch has no effect on the pixels you get back.
  • What errors can Take Screenshot return, and what does each mean?
    `no such window` when the current top-level browsing context is no longer open — usually a close() ran before the capture. `unable to capture screen` when the visual viewport has zero width or height, or when the canvas bitmap cannot be serialised, for example because its origin-clean flag is false.
  • Does the returned image size equal the CSS viewport size?
    Not necessarily. The spec samples the framebuffer, so on a high-DPI display or a zoomed browser the PNG's pixel dimensions are a multiple of the CSS viewport's width and height. Code that assumes image width equals window.innerWidth will miscalculate crops and stitch offsets.

It is a photo of the screen, not a scan of the document: the camera only records what the glass is currently showing, no matter how many pages are underneath.

saying these in an interview costs you the question

  • Saying the standard command captures the entire scrollable document
  • Claiming an OutputType or capability switches on full-page capture
  • Expecting a frame switch to scope the screenshot to that frame
  • Assuming the shot includes browser chrome like tabs and the address bar
  • Believing image pixel dimensions always equal CSS viewport dimensions