skip to content

How do you decide when a Playwright suite may drive page.mouse and page.keyboard instead of locator verbs?

level: principalimportance: should knowfreq 33%

answer

  1. Coordinates instead of a named element
  2. Focus decides where keys land
  3. Nothing re-resolves after a re-render
  4. Reserve it for gestures without a target
  5. Pin the viewport, wrap the helper

basics

~10 s

Treat the raw devices as a documented exception. page.mouse takes viewport coordinates and page.keyboard goes wherever focus already is, so neither re-resolves an element. Reserve them for gestures with no element to name.

solid answer

~40 s

`page.mouse` and `page.keyboard` are the untargeted layer: `mouse.move(x, y)` takes CSS pixels relative to the main frame's viewport and remembers where it left the pointer, and `keyboard.press()` goes wherever focus happens to be. Nothing is re-resolved, nothing is scrolled into view, and a mis-aimed click does not fail -- it succeeds somewhere else and something later goes red. The locator verbs -- `click()`, `hover()`, `press()`, `dragTo()` -- name their target, so they survive re-renders and read as intent. Keep the raw devices for gestures with no element to name: drawing on a canvas, `mouse.wheel()` scrolling, a marquee across empty board space, or holding a modifier across several actions with `keyboard.down()` and `keyboard.up()`. Wrap each use in a named helper, pin the viewport so coordinates reproduce, and assert an outcome rather than the gesture.

code

typescript · 11 lines
typescript
// No element to name: drag a selection marquee across empty board space.
await page.mouse.move(240, 180);
await page.mouse.down();
await page.mouse.move(620, 430, { steps: 12 });
await page.mouse.up();

// Hold a modifier across several locator actions, then release it.
await page.keyboard.down('Shift');
await page.getByRole('listitem', { name: 'Fix login redirect' }).click();
await page.getByRole('listitem', { name: 'Board drag drops nothing' }).click();
await page.keyboard.up('Shift');

go deeper

for a junior

Default to the locator verbs. Know that page.mouse works in viewport coordinates and page.keyboard sends keys wherever focus already is, so neither one knows which element you meant.

for a middle

Explain what you give up: no re-resolution, no scrolling into view, and no failure message naming an element. Name the cases with no element to target, such as canvas drawing or wheel scrolling.

for a senior

Show how you contain it. One named helper per gesture, a pinned viewport so coordinates reproduce, an assertion on the outcome rather than the gesture, and a habit of asking in review why no locator fits.

for a principal

Set the policy and the exception path. Decide what earns a coordinate-driven test, who reviews it, and how the suite avoids the slow drift where raw input becomes the standing workaround for weak locators.

## Two layers of input Playwright exposes pointer and keyboard input at two altitudes, and choosing between them is a design decision the whole suite lives with. The **locator verbs** -- `click()`, `dblclick()`, `hover()`, `press()`, `tap()`, `dragTo()` -- name their target. Each call re-resolves the locator against the live DOM, scrolls the element into view when it needs to, and reports failures in terms of the element you asked for. The test reads as intent: click the card called "Fix login redirect". The **raw devices** -- `page.mouse` and `page.keyboard` -- name nothing. `mouse.move(x, y)` takes CSS pixels relative to the main frame's viewport, and the pointer position is stateful: it stays where you left it between calls. `keyboard.press(key)` goes to whatever element currently holds focus, whether or not you put it there. `mouse.wheel(deltaX, deltaY)` scrolls from wherever the pointer is. `page.touchscreen.tap(x, y)` is the touch equivalent, and it, like `locator.tap()`, needs the context created with `hasTouch` enabled. ## What you give up by going raw - **Re-resolution.** A locator verb finds the element again on each attempt, so a board that re-renders between actions is handled. A coordinate is computed once and is stale immediately. - **Scrolling.** Locator verbs bring the element into view. `page.mouse.move()` clicks whatever is at those viewport pixels, which after a scroll is a different element. - **Diagnosability.** A failing locator verb names the element and shows what it matched. A raw click that lands on the wrong thing does not fail at all -- it succeeds, and something three assertions later goes red, far from the cause. - **Readability.** `page.mouse.click(412, 268)` tells a future reader nothing about what was clicked, and nothing about what should still be clicked after a redesign. - **Viewport independence.** Coordinates are only meaningful against a fixed viewport, so a new banner, a different window size or a device descriptor with another screen silently moves every target. ## Where the raw devices are the right answer They are not a smell in themselves -- they are the only option when there is genuinely no element to name: - drawing or gesturing on a `canvas`, where the interesting positions are pixels rather than nodes - a marquee selection dragged across empty space on the issue board, which starts nowhere in particular - `mouse.wheel()` scrolling, when the test is about what a wheel event does rather than about reaching an element - holding a modifier **across** several actions with `keyboard.down()` and `keyboard.up()`, which the per-call `modifiers` option cannot express because it restores the previous state after each click - sending a key to whatever the application focused itself, when the point of the test is that the application moved focus correctly ## Containing the exception The judgment call is not "never" -- it is how to keep a legitimate exception from becoming the path of least resistance for anyone who finds a locator hard to write. A workable policy: 1. **Wrap every raw gesture in a named helper.** `drawSelectionBox(page, from, to)` carries the intent that the coordinates destroy, and it gives you one place to fix the numbers. 2. **Pin the viewport** for any test that uses coordinates, so the numbers reproduce across machines and CI. 3. **Assert an outcome, never the gesture.** After a marquee, assert that three cards are selected. The gesture succeeding proves nothing. 4. **Require a reason in review.** "No element to name" is a reason. "The locator was flaky" is a bug report, and swapping in coordinates buries it. 5. **Derive coordinates rather than hard-coding them** where you can -- from a bounding box you measured in the same test -- so a layout change moves the gesture with it. | | locator verbs | `page.mouse` / `page.keyboard` | |---|---|---| | target | a named element | viewport coordinates, or current focus | | survives a re-render | yes, re-resolved | no | | scrolls into view | yes | no | | failure message | names the element | none; it silently lands elsewhere | | right for | almost everything | canvas, marquee, wheel, held modifiers | ## The strategic point Raw input is cheap to write and expensive to own. Each coordinate is a small undocumented assumption about layout, and a suite accumulates them faster than it removes them, because the failure mode is a later assertion rather than the gesture itself. Treating the raw devices as a documented, reviewed exception -- rather than banning them or leaving the choice to whoever is closest to a deadline -- is what keeps the suite's input layer honest as the product's layout keeps moving.

  • Why does keyboard.down('Shift') around two clicks behave differently from passing modifiers to each click?
    `keyboard.down()` leaves `Shift` genuinely held until you release it, so everything in between happens under that modifier, including the page's own focus changes. The `modifiers` option is scoped to one call: Playwright holds exactly those keys for that click and restores the previous state afterwards.
  • What breaks first when a suite leans on page.mouse coordinates?
    The viewport. Coordinates are CSS pixels in the main frame's viewport, so a different window size, a new banner, or a device descriptor with another screen shifts every target. The click still succeeds, it just lands somewhere else, which fails later and far from the cause.
  • A teammate wants coordinates because the locator was flaky. How do you respond?
    Treat it as a bug report, not a reason. A flaky locator usually means an ambiguous match, a re-render mid-action, or a genuine product defect, and coordinates bury all three while silently coupling the test to layout. Fix the locator, and reserve raw input for gestures with no element to name.

saying these in an interview costs you the question

  • Uses coordinates because a locator was hard to write
  • Assumes mouse coordinates are relative to the element
  • Thinks page.keyboard focuses the element for you
  • Believes raw mouse calls scroll the target into view
  • Expects the pointer position to reset between actions