Your Appium rooftop-solar quote test pauses three minutes for a backend quote, then every later command fails — why?
answer
- look at the gap, not the locator
- server counts silence per session
- any command restarts the countdown
- size it to the longest legitimate pause
basics
~10 sThe Appium server ended the session during the quiet stretch: appium:newCommandTimeout counts silence between commands, and three minutes of it exceeded the limit. Every later command then addresses a session that no longer exists.
solid answer
~50 sThe session was killed by its own idle timer, not by the app and not by a locator. `appium:newCommandTimeout` counts **seconds of silence between commands** on a session; three minutes of the test talking to a backend and not to the Appium server exceeded it, so the server tore the session down. The next command names a session that no longer exists, which is why everything after the gap fails identically instead of one step failing on an element. Three fixes suit different cases: size `appium:newCommandTimeout` to the longest legitimate gap; keep the session warm by issuing a cheap command on it from inside the waiting loop, since any command restarts the countdown; or move the backend wait outside the session entirely. Raising it suite-wide is the blunt option, because the same timer is what reclaims a device when a run wedges.
go deeper
Recognise the shape of the failure: after one long quiet step, every remaining command fails on the session itself rather than on an element it could not find.
Explain the mechanism — the Appium server counts seconds of silence per session, any command restarts the count, and expiry tears the session down rather than waiting.
Show the fix menu and the choice between the options: size the idle timeout, keep the session warm with a cheap command, or move the long wait outside the session entirely.
Own the fleet trade-off. A generous idle timeout rescues legitimate long gaps but lets a wedged run sit on a device, so pair it with a run-level cap and a policy on where long waits may live.
## Reading the symptom before touching a timeout The shape of this failure is diagnostic on its own. One step in the rooftop-solar quote test goes quiet for three minutes while a backend prices the roof survey. The step after it fails. So does every step after that, in the same way, and the failure is about **the session** rather than about any element — the server is being asked to act on a session id it no longer knows. A locator problem does not do that; it fails one step and leaves the session usable. A crash in the app under test does not do that either; the session survives an app crash and later commands still reach the server. What did happen is that `appium:newCommandTimeout` ran out. It is a **session-idle timer** measured in **seconds**, it counts silence between commands rather than anything on the device, and when it expires the Appium server tears the session down instead of waiting. Three minutes of a test talking to a backend and not to the server is exactly the silence it exists to end. ## Why the gap looked like a dead client - The timer lives in Appium's core, above the platform drivers, so it behaves the same for a UiAutomator2 session on Android and an XCUITest session on an Apple device. - It restarts on **any** command received for that session, whether that command succeeds or fails. - Nothing on the device restarts it: not the app computing a quote, not an animation, not a spinner, not a person watching the screen. - A wait that lives entirely inside the test process — polling your own backend, waiting on a queue, sitting on a breakpoint — sends no request at all, and that is the case that bites here. - When the timer fires there is no error to deliver, because the test is not asking the server anything; the news arrives on the next command instead. ## The levers, and when each is the right one 1. **Size the idle timeout to the longest legitimate gap.** Set `appium:newCommandTimeout` from the biggest silence the suite honestly has, not from the longest test. This is the right fix when the gap is real, bounded and unavoidable. 2. **Keep the session warm.** Any command received restarts the countdown, so a cheap periodic call on the session from inside the waiting loop — `GET /session/:sessionId/contexts`, for instance — carries one long wait across without loosening the timer for the whole suite. This is the right fix when a single step is the outlier. 3. **Move the wait out of the session.** If the three minutes are spent entirely on a backend, do that work before the session starts or after it ends, and keep the device session to the part that actually drives the device. This is the right fix when the gap has nothing to do with the phone. 4. **Change the timing surface on a running session.** `POST /session/:sessionId/timeouts` is how a live session's timeouts are updated after start, and it carries a caution of its own. ## A caution about the mid-session route The W3C parameters on `POST /session/:sessionId/timeouts` are `script`, `pageLoad` and `implicit`. Appium's base driver additionally still declares the legacy `type` and `ms` pair on that route and still implements the legacy branch behind it, including an Appium-specific `command` type — even though Appium 3's own migration guide describes the legacy pair as no longer accepted. Both of those statements are true of the code as it stands. Treat the legacy form as present but contested: verify it against the server you actually run before a suite leans on it, and prefer the capability at session start for something as structural as the idle timer. ## What losing the session actually costs Sessions are not free to replace, and the replacement cost is one of the genuine divergences between the two platforms: - On **Android**, a fresh session means the driver pushing and starting its helper server on the device again before your first command runs. - On **Apple platforms**, a fresh session means WebDriverAgent running again — which, unless you supply a pre-built agent or point the driver at one already running, means `xcodebuild` doing a build first. So a suite that loses sessions to the idle timer does not merely fail; it pays the startup price again for every retry, and it pays a different price on each platform. That is worth saying out loud in a review, because it turns retrying the test from a free move into a measurable one. ## The judgment call A generous idle timeout is not free either. The same timer is what reclaims a device when a run wedges with a session still open, so every second added to it is a second a stuck run can keep hardware busy. The defensible position has three parts: a value sized to the suite's real gaps, an explicit decision about where long waits are allowed to live at all, and a run-level cap outside Appium, so that nothing depends on a session timer to notice that a run has stopped making progress.
- How would you prove the idle timeout, rather than the app, killed the run?Line the timings up. The failure starts on the first command after the quiet step, every later step fails the same way, and each failure is about the session rather than any element. The Appium server log then shows the shutdown happening inside the gap, roughly `newCommandTimeout` seconds after the last command the server received.
- Is raising `appium:newCommandTimeout` for the whole suite a safe fix?It is safe for this failure but not free. The same timer is what reclaims a device when a run wedges with the session still open, so a very large value means a stuck run holds its hardware far longer. Size it to the longest legitimate gap you actually have, and put a cap on the run itself outside Appium.
- Which of the fixes would you reach for first on a suite that already runs in parallel?Keeping the session warm, or moving the wait out of the session. Both are local to the step that misbehaves. Raising the idle timeout is global, and on a parallel run it lengthens the window in which any wedged worker keeps a device tied up, so it is the change with the widest blast radius for the narrowest cause.
saying these in an interview costs you the question
- Blames the locator or the app when the whole rest of the test fails
- Thinks device activity during the gap keeps the session alive
- Reaches for element wait timeouts, which do not affect session silence
- Assumes a wait inside the test process still sends traffic to the server
- Believes only the Android drivers or only the XCUITest driver behave this way
- Treats a huge idle timeout as free rather than as a device-holding trade