skip to content

In Selenium, where on an element does WebElement.click() actually land, and how is that point computed?

level: seniorimportance: should knowfreq 41%

answer

  1. One coordinate, not a whole element
  2. Only part of the element may count
  3. A rectangle list, not a bounding box
  4. First getClientRects entry, clipped to viewport
  5. Floored midpoints of the clamped box

basics

~10 s

Selenium clicks the in-view centre point: the centre of the intersection between the element's first client rectangle and the viewport, floored to whole pixels. Whatever is painted at that one coordinate receives the click.

solid answer

~40 s

The Element Click command targets a single coordinate, not an element. It takes the **first** `DOMRect` returned by `getClientRects()` on the element, intersects that rectangle with the viewport, and floors the midpoints of the result - that pair is the in-view centre point. It then dispatches a `pointerMove`, `pointerDown` and `pointerUp` there. The remote end also asks which elements are painted at that coordinate: if the top-most one is not the element or a descendant of it, the element counts as obscured. So a wrapped link is clicked in the middle of its first line, a card taller than the viewport is clicked in the middle of its visible part, and "the element is visible" is never the same claim as "the element is the top-most thing at its centre pixel".

go deeper

for a junior

Recall that a Selenium click goes to the middle of the element rather than to a corner, and that the driver scrolls the element into view first. Being able to say "the centre" out loud is enough here.

for a middle

Be ready to explain how that centre is derived: the first client rectangle, clipped to the viewport, with the midpoints floored. Say why that is not the same as the element's bounding box.

for a senior

Show that you have used this to diagnose a real misdirected click - a wrapped link, a card taller than the screen, something painted over the centre pixel - and that you reason about the coordinate before you blame the application.

for a principal

Own the consequence for a suite: teams should click the smallest element carrying the affordance, and treat "visible" as weaker evidence than "top-most at its centre point". Set that expectation in review rather than after every flaky run.

## What the click command is actually asked to do A Java call such as `driver.findElement(By.cssSelector("[data-article='printer-jam']")).click()` sends one **Element Click** command to the remote end. The remote end does not "click the element". It scrolls the element's container into view if it is not already in view, computes a single coordinate called the **in-view centre point**, and dispatches a pointer sequence at that coordinate: a `pointerMove` whose origin is the element, then a `pointerDown` with `button` 0, then a `pointerUp`. Every surprise in this area follows from one fact: the target is **a pixel, not an element**. ## How the in-view centre point is computed 1. Call `getClientRects()` on the element and take **only the first** `DOMRect` in the collection. 2. Clamp that rectangle to the viewport: `left = max(0, ...)`, `right = min(innerWidth, ...)`, and the same treatment for `top` and `bottom` against `innerHeight`. 3. Take the midpoints of the clamped box and floor them: `x = floor((left + right) / 2)` and `y = floor((top + bottom) / 2)`. Two clauses in there do almost all of the damage in real suites: *first* rectangle, and *clamped to the viewport*. ## Why the first rectangle, and what it costs you On a helpdesk knowledge-base results list, an article title such as "Resolving a printer jam on the shared floor printers" is an inline `<a>` that wraps onto two lines. `getClientRects()` returns one rectangle per line box, and the spec takes the first. The click therefore lands in the middle of the **first line**, which is nearly always on the text. If instead you reason about the element's bounding box, you get a different answer: for a wide first line and a short second line, the union box's centre can fall in the empty space to the right of the second line, which belongs to the surrounding row and not to the link at all. The rule protects you here, but only if you know it exists: - Coordinates you compute yourself from a bounding box will **not** agree with where Selenium clicks. - An element whose `getClientRects()` list has length zero has no centre point at all. - The rule is per-element, so a wrapped link inside a card behaves differently from the card itself. ## Why clipping to the viewport matters The clamp is why a tall element behaves oddly. A knowledge-base article card taller than the viewport can never be fully on screen, so its centre point is the centre of **the part that overlaps the viewport** and moves as the page scrolls. That is usually what you want. It is also why "the element is visible" and "the element is clickable at its centre" are different statements. ## The paint tree at that one pixel The specification defines an element's **pointer-interactable paint tree** as the result of asking the document which elements are at the centre point, and calls an element **obscured** when that list is empty, or when the first entry in it is not the element itself or one of its descendants. So the real question the remote end asks is never "is this link visible?" but "is this link the top-most thing painted at exactly this coordinate?". Anything drawn over that pixel - a sticky filter bar above the results list, a floating chat launcher, a transparent scrim - answers it for you. | Common assumption | What the Element Click command does | |---|---| | It clicks the element | It clicks one coordinate and reaches whatever is painted there | | It uses the element's bounding box | It uses the **first** `DOMRect` from `getClientRects()` | | It uses the whole box even off-screen | It intersects that rectangle with the viewport first | | The centre is a fractional point | Both coordinates are floored to integers | | A visible element is clickable | Only visibility **at that one pixel** counts | ## What to do with this in a knowledge-base suite - Prefer clicking the smallest element that carries the affordance - the `<a>` inside the result row, not the row - so the pixel and the intent coincide. - Treat "the element is on screen" as insufficient evidence; the relevant property is what occupies its centre point. - Remember that `WebElement.click()` performs this computation **once** and dispatches immediately; it contains no retry of its own, so a point that is free a moment later is of no help to the command. - When a click reliably lands somewhere strange, reason about the first client rectangle and the viewport clamp before you reason about the application.

  • Why does the specification use the first client rectangle rather than the element's bounding box?
    Because an inline element that wraps produces one rectangle per line box, and the union of those boxes can have its centre in empty space beside a short final line - space that belongs to the surrounding container. Taking the first rectangle keeps the coordinate on rendered content.
  • An element is taller than the viewport and can never fit on screen. Does it have a centre point?
    Yes. The rectangle is clamped to the viewport before the midpoints are taken, so the centre point is the centre of the visible portion and it moves as the page scrolls. Only an element whose client rectangle list is empty has no centre point at all.
  • If the centre point is momentarily covered, will the click command wait for it to clear?
    No. The Element Click command computes the coordinate and dispatches the pointer sequence once; it has no polling or retry step of its own. Whether the page was ready is a separate concern the command does not own.

It is a dart, not a hand. The driver throws one dart at a single pixel, and whatever is painted over that pixel catches it, however large and obvious the element looks to you.

saying these in an interview costs you the question

  • Says click targets the element rather than one coordinate
  • Reasons about the centre using the element's bounding box
  • Assumes a visible element is always clickable at its centre
  • Thinks the full box is used even when partly off-screen
  • Believes click retries by itself until the pixel is free