skip to content

Pointer and Key Sequences

The Actions builder assembles hover, drag, right-click, held modifiers, wheel scrolling and pauses into one sequence sent in a single request. Sticky modifiers and HTML5 drag are the traps.

on this pageshow

explore

questions

4

In Selenium, which class performs a hover or a right-click, and what call actually runs it?

level: juniorimportance: must knowfreq 76%

answer

  1. One builder class for pointer gestures
  2. Constructed from the driver, then chained
  3. moveToElement hovers, contextClick right-clicks
  4. The chain is inert until one final call
  5. perform() is what sends the sequence

basics

~10 s

Selenium's Actions class builds them: moveToElement for a hover, contextClick for a right-click, plus doubleClick, clickAndHold, release and dragAndDrop. Nothing reaches the browser until you call perform() on the chain.

solid answer

~40 s

You build the gesture with `org.openqa.selenium.interactions.Actions`, constructed from the driver, and run it with `perform()`. `moveToElement(el)` hovers by moving the pointer to the element's in-view centre; `contextClick(el)` right-clicks it; `doubleClick(el)` double-clicks; `clickAndHold(el)` presses without releasing and `release(el)` lets go; `dragAndDrop(source, target)` and `dragAndDropBy(source, x, y)` wrap hold-move-release for you. Each method returns the same builder, so the calls chain: `new Actions(driver).moveToElement(row).contextClick().perform()`. The classic beginner bug is omitting `perform()` - the chain is collected in memory and silently never sent, so the test fails later on the assertion rather than at the gesture. A second trap is reusing one `Actions` object for two gestures: `build()` empties the builder, so the second `perform()` sends nothing.

code

java · 19 lines
java
WebDriver driver = new ChromeDriver();
driver.get("https://reports.example.com/approvals");

WebElement row = driver.findElement(By.cssSelector("[data-report-id='EXP-4471']"));
WebElement approvedBucket = driver.findElement(By.id("approved-bucket"));
Actions actions = new Actions(driver);

actions.moveToElement(row).perform();

actions.contextClick(row).perform();

actions.doubleClick(row).perform();

actions.clickAndHold(row)
        .moveToElement(approvedBucket)
        .release()
        .perform();

driver.quit();

go deeper

for a junior

Name the Actions class, the method for each gesture, and the fact that perform() runs the chain. Being able to write a hover then a right-click from memory is the bar here.

for a middle

Explain that the convenience methods expand into pointer primitives, and that a manual hold, move and release is the same thing as dragAndDrop. Know why a missing perform() fails silently.

for a senior

Show judgment about when a hand-built sequence is warranted at all, and how you diagnose a drag that does nothing rather than blindly adding offsets and pauses to the chain.

for a principal

Own the cost question: pointer sequences are the least portable part of a suite, so decide where the team accepts them and where the application should expose a simpler path.

## The class, and the shape of a chain Every pointer or keyboard gesture that is more than a plain click lives on `org.openqa.selenium.interactions.Actions` in Selenium 4. You construct it with the driver, chain the steps you want, and call `perform()`: ```java new Actions(driver).moveToElement(reportRow).contextClick().perform(); ``` On an expense-report approval list that one line hovers a report row so its inline toolbar renders, then opens the row's right-click menu. Every method on the chain returns the same `Actions` object, which is what makes the chaining read the way it does. ## The convenience methods You almost never build raw pointer primitives yourself. `Actions` wraps them: | Method | What it does | |---|---| | `moveToElement(el)` | Moves the pointer to the in-view centre of `el` - a hover | | `moveToElement(el, x, y)` | Moves to an offset from that centre point | | `contextClick(el)` | Moves to `el`, then presses and releases the right button | | `doubleClick(el)` | Moves to `el`, then presses and releases the left button twice | | `clickAndHold(el)` | Moves to `el` and presses the left button without releasing it | | `release(el)` | Moves to `el` and releases the left button | | `dragAndDrop(src, tgt)` | Hold on `src`, move to `tgt`, release | | `dragAndDropBy(src, x, y)` | Hold on `src`, move by an offset, release | | `keyDown(key)` / `keyUp(key)` | Presses or releases one modifier key | | `pause(Duration)` | Waits inside the sequence | The no-argument overloads - `contextClick()`, `doubleClick()`, `clickAndHold()`, `release()` - act **at the pointer's current position** rather than on an element. That is what lets you write a hand-built drag as `clickAndHold(row).moveToElement(bucket).release().perform()`, which is exactly what `dragAndDrop(row, bucket)` does for you. A few details are worth carrying around: - Every one of these methods expands into **pointer primitives** - a `pointerMove`, a `pointerDown`, a `pointerUp` - so `doubleClick(el)` is one move followed by two press-and-release pairs. - `moveToElement(el, x, y)` measures its offset from the element's **in-view centre point**, not from its top left corner, so `(0, 0)` is the centre. - `dragAndDropBy(src, x, y)` moves by an offset **from where the pointer already is** after the press, so the numbers are relative, not page coordinates. - `keyDown(el, key)` and `keyUp(el, key)` click the element first to focus it, then press or release the key on the keyboard device. - `pause(long)` and `pause(Duration)` both exist and wait **inside** the sequence, so the driver paces the gesture rather than your test thread. ## perform() is not optional The single most common beginner mistake with this API is building a chain and never running it: ```java Actions actions = new Actions(driver); actions.moveToElement(reportRow).contextClick(); // nothing happened ``` `Actions` collects steps in memory; nothing reaches the browser until `perform()` (or `build()` followed by performing the returned `Action`). Because no request was ever made, the test does not throw - it simply carries on and fails later, on the assertion that expected the context menu. So when a gesture "does nothing", work through this order: 1. Check that the chain ends in `perform()`. 2. Check that this `Actions` object has not already been performed once. 3. Only then start suspecting the page. There is a second, quieter version of the same trap. `build()` empties the builder as it snapshots the sequences, and `perform()` calls `build()`, so running the same `Actions` object twice does **not** repeat the gesture: the second call sends an empty action set. Build a fresh chain for each gesture. ## Drag and drop, and where it stops working `dragAndDrop(source, target)` is a convenience for four steps on the virtual pointer: move to the source, press, move to the target, release. That is a **pointer-level** gesture, and pages that implement dragging by listening to pointer or mouse movement respond to it normally. A page that instead implements dragging through the browser's native HTML5 drag-and-drop API is a different mechanism, and a synthesized pointer sequence may not drive it. The event model behind that API is a browser-platform subject rather than a Selenium one; the practical Selenium-side point is only that `dragAndDrop` failing on a list that a human can clearly drag is a known shape of problem, not a mistake in your chain. Confirm which mechanism the approval list uses before you spend an afternoon tuning offsets. ## Element-targeted versus current-position overloads Choosing between `contextClick(row)` and `moveToElement(row).contextClick()` is usually taste - they build the same interactions. The difference appears once you need something in between: - `moveToElement(row).pause(Duration.ofMillis(300)).contextClick().perform()` gives a hover-triggered menu time to render before the right-click lands. - `clickAndHold(row).moveByOffset(0, 40).moveToElement(bucket).release().perform()` nudges the pointer first, which some drag implementations need before they treat the gesture as a drag at all. Both are ordinary chains. What makes them work is that each step is queued, the whole queue is sent once, and the browser replays it in order.

  • What is the difference between contextClick(element) and moveToElement(element).contextClick()?
    Almost nothing - the first is a convenience for the second, and both build a pointer move followed by a right button press and release. The difference only matters when you want something between the two, such as a `pause` that lets a hover-triggered menu render before the right-click lands.
  • Your dragAndDrop call does nothing on a list a human can clearly drag. What do you check first?
    Which drag mechanism the page uses. `dragAndDrop` synthesizes a pointer press, move and release, which drives pages that listen to pointer movement; a page built on the browser's native HTML5 drag-and-drop API is a different event model and may not respond. Confirm that before tuning offsets or durations.
  • Why can reusing the same Actions object for two gestures run only the first?
    `perform()` calls `build()`, and `build()` clears the builder's sequences as it snapshots them into a composite action. The second `perform()` therefore finds an empty set and sends an effectively empty command. Construct a new `Actions` for each gesture, or re-chain the steps before performing again.

saying these in an interview costs you the question

  • Builds an Actions chain and never calls perform()
  • Thinks hovering needs injected script rather than moveToElement
  • Reuses one Actions object for several gestures and expects each to run
  • Assumes dragAndDrop drives every draggable list on every page
  • Confuses contextClick with a plain left click on the element
open as a page

After a Ctrl+click Actions chain in Selenium, later steps in the same session misbehave. Why, and what clears it?

level: seniorimportance: must knowfreq 46%

basics

~20 s

Selenium's Actions never releases a modifier for you. A keyDown with no matching keyUp leaves the key depressed in the driver's input state, which outlives the chain and any new Actions object. Release it with keyUp, sendKeys(Keys.NULL), or resetInputState.

open as a page

In Selenium, what does calling perform() on an Actions chain send, and how are its ticks ordered?

level: middleimportance: should knowfreq 54%

basics

~20 s

Actions collects each step into a per-device sequence, and perform() posts them all as one WebDriver actions command. Steps sharing a tick run together across devices, and the browser waits for that tick's longest duration before the next one.

open as a page

In Selenium, how do Actions.scrollToElement() and Actions.scrollByAmount() differ, and which device do they drive?

level: juniorimportance: nice to knowfreq 31%

basics

~20 s

Both drive Selenium's virtual wheel device, added in Selenium 4.2. scrollToElement takes an element and scrolls it into view, letting the driver work out the distance; scrollByAmount takes pixel deltas from the viewport's top left corner.

open as a page