You inherit a REST Assured API suite of several hundred tests that runs against a deployed service: it takes 40 minutes, roughly one run in three fails for reasons nobody trusts, and any API change turns dozens of tests red. How would you approach making it trustworthy?
answer
- classify failures before fixing anything
- per-test data with unique markers, no shared fixtures
- parallel requires dropping RestAssured statics
- poll with a deadline, never sleep
- narrow assertions + one schema; quarantine lane + owner
basics
~20 sMeasure first: classify failures into real bugs, shared-environment/data collisions, timing, and over-assertion. Then fix causes — per-test data created and cleaned up via the API, no shared mutable state or ordering, parallel execution with per-request specs, assertions narrowed to behaviour plus one schema check, and failure-only logging with correlation ids for traceability.
solid answer
~60 sI would not start rewriting; I would start measuring. Collect a few weeks of results, tag each failure, and see the distribution: real defects, environment/data collisions, timing/async, or brittle assertions. The usual causes and their fixes: - **Shared data.** Tests that assume a fixed user or a specific record count collide with each other and with anyone else using that environment. Each test should create its own data with unique values and clean up, so it is order- and neighbour-independent. - **Serial execution.** 40 minutes is mostly waiting on the network. Once tests are data-independent, run in parallel — which forces you off `RestAssured` static globals onto per-request specifications. - **Timing.** Async side effects asserted immediately produce intermittent failures; poll with a bounded timeout instead of sleeping. - **Over-assertion.** Pinning every field means every harmless API change breaks dozens of tests. Assert the behaviour-carrying fields plus one JSON Schema check for shape. - **Diagnosability.** Failure-only logging, a correlation id per request, so a red build can be traced into server logs. And I would give the suite an owner and a flake budget — an untrusted suite that nobody is accountable for degrades again within a quarter.
code
java · 21 linesString marker = "qa-" + java.util.UUID.randomUUID();
Integer id = null;
try {
id = given().spec(apiSpec)
.body(java.util.Map.of("name", marker, "email", marker + "@example.test"))
.when().post("/users")
.then().statusCode(201)
.extract().path("id");
given().spec(apiSpec)
.queryParam("q", marker)
.when().get("/users")
.then().statusCode(200)
.body("findAll { it.name == '" + marker + "' }.size()", org.hamcrest.Matchers.is(1));
} finally {
if (id != null) {
given().spec(apiSpec).when().delete("/users/{id}", id)
.then().statusCode(org.hamcrest.Matchers.anyOf(
org.hamcrest.Matchers.is(204), org.hamcrest.Matchers.is(404)));
}
}go deeper
Focus on the concrete habits: each test creates and removes its own data, no reliance on records another test made, no sleeps.
Add the mechanics — unique markers, GPath filters instead of absolute counts, specs instead of statics, bounded polling for async.
Lead with triage and measurement, then sequence data independence, parallelism, assertion narrowing and diagnosability, with concrete before/after numbers.
Frame it as restoring trust in a signal: classify failures, decide which tests should exist at this level at all, put the endpoint surface behind a client layer, and establish ownership, a flake budget and a quarantine policy so the fix holds.
## Start with evidence, not with a rewrite A suite that fails one run in three has already lost its function: people re-run until green, and a real regression slips through unnoticed. The first move is diagnostic, not corrective. Instrument the runs — most CI systems can retain per-test history — and for two or three weeks classify every failure into buckets: 1. genuine product defects, 2. data/environment collisions, 3. timing and asynchrony, 4. brittle or over-specified assertions, 5. infrastructure (deploy in progress, DNS, TLS, token expiry). The distribution tells you where the effort belongs, and it converts "the suite is flaky" into a list you can work. It also gives you a baseline to prove improvement, which matters because this work competes with feature work for time. ## Test data is usually the dominant cause Against a deployed, shared environment, the single biggest source of untrustworthy failures is **shared mutable data**. Symptoms: tests that assert a record count, tests that depend on a fixed seeded user, tests that pass alone and fail in the suite, tests that fail only when someone is manually clicking around the same environment. The cure is that each test owns its data: - create the entities the test needs at the start, via the API, with unique values (a UUID suffix in names/emails); - never assert on absolute counts of a shared collection — assert that *your* record appears, using a GPath filter on your unique marker; - clean up afterwards, and make cleanup tolerant of a failed test (delete-if-exists); - accept that setup calls cost time — that is bought back by parallelism. Where the API cannot create a precondition (an account in an exotic state, a third-party callback), that is a signal about the product's testability, and often the honest answer is that this case belongs at a lower level, not in the HTTP suite. ## Then parallelize Forty minutes of a network-bound suite is mostly idle waiting. Once tests are data-independent they can run concurrently, and wall-clock time drops close to linearly with worker count until the environment becomes the bottleneck. Parallelism has a hard prerequisite in REST Assured specifically: **abandon the static globals**. `RestAssured.baseURI`, `RestAssured.port`, `RestAssured.authentication`, `RestAssured.requestSpecification` and `RestAssured.filters` are static, shared, not thread-scoped. Concurrent tests mutating them produce failures that are genuinely impossible to reason about. Every request must carry its own `RequestSpecification`, built once and passed explicitly. This is the point where the earlier convenience becomes a blocker, and it is worth converting the whole suite in one mechanical pass rather than test by test. Watch for other shared state too: static fields caching an id created by one test and consumed by another, and any implicit ordering. ## Timing and asynchrony An API call that triggers asynchronous work (an event, a projection update, an email) cannot be asserted synchronously. `Thread.sleep` is the wrong fix — too short and it flakes under load, too long and it is pure runtime tax. Poll with a bounded timeout: retry the read endpoint until the condition holds or a deadline passes. REST Assured has no polling of its own, so this comes from an awaiting utility or a small loop helper; either way keep the timeout explicit and the failure message state what was still not true when time ran out. Response-time assertions on a shared environment are another quiet flake source. If they exist, they should be generous ceilings meant to catch pathological regressions, not performance measurements — real performance testing does not belong in a correctness suite. ## Fragility to API change When one field rename breaks fifty tests, the suite has been written as a snapshot of the payload rather than an assertion about behaviour. Two structural moves: - **Narrow the assertions.** Each test asserts the status code plus the two or three fields that carry the behaviour under test; volatile values (ids, timestamps) get `notNullValue()` rather than pinned values. - **Centralize the shape.** One JSON Schema per response type, asserted with `matchesJsonSchemaInClasspath`, catches dropped/renamed/retyped fields for the whole suite. When the API legitimately changes, one schema file changes instead of fifty test methods. Above that, put the endpoint calls behind a thin client layer — small methods like `createUser(request)` returning a typed response — so a URL or header change is one edit. This is the most valuable refactor and the one people skip. ## Diagnosability A failure that cannot be diagnosed from the CI output will be re-run rather than fixed. Minimum kit: `log().ifValidationFails()` on request and response suite-wide (with `Authorization` blacklisted in `LogConfig`), a correlation-id filter stamping a unique header on every request so a failure can be traced into server logs and traces, and the environment/build identity recorded in the report. ## Governance The technical fixes decay without ownership. Practices that hold: a named owner for the suite; a flake budget with a quarantine lane (a persistently flaky test is disabled with a ticket rather than left to erode trust); a rule that quarantined tests are fixed or deleted within a fixed window; and a review expectation that new API tests create their own data and use the shared client layer. ## Knowing what to delete Finally, some of the suite should not exist. Hundreds of HTTP tests usually include many that exercise validation permutations or branch logic reachable far more cheaply and reliably below the HTTP boundary. Migrating those down and deleting the HTTP versions shortens the run, removes flake surface, and leaves the API suite doing what only it can do: proving the deployed service's contract actually works end to end.
- Why does moving a REST Assured suite to parallel execution force you off RestAssured's static configuration fields?`RestAssured.baseURI`, `port`, `authentication`, `requestSpecification` and `filters` are static fields on a shared class with no thread scoping. Under concurrency, one worker's assignment is visible to every other worker, so tests intermittently hit the wrong host or send the wrong credentials, and the failures are effectively undebuggable. Passing an explicit `RequestSpecification` into `given().spec(...)` makes every request carry its own configuration, which is the only safe model once more than one test runs at a time.
- Some tests in the suite cannot create their own preconditions through the API. What do you do with them?First treat it as product feedback: a state that is unreachable through the API is also unreachable for real clients and support staff, and an endpoint or admin capability may be justified. Where that is not warranted, the honest answer is usually that the case does not belong in the HTTP suite — the branch it covers is reachable far more cheaply and deterministically below the API boundary. Keeping it as an HTTP test that depends on hand-seeded environment data is the worst option, because it will fail for reasons unrelated to the behaviour it claims to check.
- How do you decide which of the several hundred tests to delete rather than repair?Look for tests whose failure would never tell you something new: permutations of input validation, branch logic, and mapping rules that are fully determined below the HTTP layer and already covered there, or reachable there at a fraction of the cost. Keep the tests that only an over-the-wire call can prove — routing, serialization, authentication and authorization, status codes and headers, and the handful of end-to-end flows that represent real client journeys. Deleting the rest shortens the run and removes flake surface, and it is cheaper than repairing tests nobody will read.
It is the same discipline as taming a noisy alerting system: until you triage the pages into causes and cut the false ones, nobody responds to the real alert.
saying these in an interview costs you the question
- Proposing an immediate rewrite without measuring what actually fails and why.
- Adding retries or Thread.sleep as the primary fix, which hides real regressions and inflates runtime.
- Keeping shared seeded fixtures and per-test cleanup 'for speed', then wondering why tests fail only in the full suite.
- Turning on parallel execution while still setting RestAssured's static baseURI/authentication in setup.
- Fixing flakiness technically with no owner or flake budget, so the suite degrades again within months.