In a Gatling run where a third of the requests failed, why does index.html's Response Time Percentiles over Time chart still look flat and fast, and which panels of the same report show the failures?
answer
- The chart title ends in OK
- Failures leave the chart, not the table
- Ranges checks status before response time
- Responses per second places failures on the clock
basics
~20 sThat chart plots successful requests only, so timeouts and early failures never enter it. Read the % KO column, the Stats table's response-time columns, the failed bar of Response Time Ranges and the Errors table instead.
solid answer
~40 sGatling builds the **Response Time Percentiles over Time** chart from the records it marked `OK` — its title ends in `(OK)` precisely to say so — because failures can end prematurely or be timeouts, and would distort the curve. Every request that failed simply leaves that chart, so the survivors look fast. The failures are all still in the report, just elsewhere: the `KO` and `% KO` columns of the Stats table, the fourth bar of the **Response Time Ranges** panel, the failures series of **Responses per second over time**, and the **Errors** table with its message, count and percentage. Note too that the global Stats table's response-time columns are computed over **all** requests, so the table's 95th percentile and the chart's 95th line can legitimately disagree.
go deeper
Be ready to notice the (OK) in the chart's title and to say what it excludes, before quoting any number off that chart.
Be ready to explain that the global Stats table's percentile columns cover all requests while the chart covers successes only, so the two can disagree without either being wrong.
Be ready to walk a report in an order that cannot mislead: KO share first, errors second, failures on the clock third, response times last.
Be ready to argue for reporting conventions your team can hold — never quoting a response-time figure without the KO share beside it — and to say why a tool-side default caused the problem.
## Why the chart is clean The chart on `index.html` titled **Response Time Percentiles over Time (OK)** is fed one population and one only: the requests Gatling recorded as `OK`. The same is true of the matching chart on every request page and of both charts on a group page — all four carry `(OK)` in their titles. Gatling's own documentation gives the reason: failed requests can end prematurely or be caused by timeouts, and letting them into the computation would have a drastic effect on the figures. That design choice is defensible and it is also a trap. In a run where a third of the requests failed, a third of the sample has walked out of the chart — and it is exactly the abnormal third. Worse, a fast-failing target flatters the chart twice over: connections that are refused or reset in milliseconds never appear, while the successes that remain are the ones the target still had capacity to serve properly. The curve can therefore go **down** as the system gets **worse**. ## Where the failures actually are Every failure is still in the report. The panels differ in what they will tell you about them: | panel | what it carries about failures | |---|---| | Stats table, Executions half | `KO` count and `% KO` per row, and per group when the scenario uses `group(...)` | | Stats table, Response Time half | on the global page, figures over **all** requests — failures included | | Response Time Ranges | a fourth bar counting failures; status is tested **before** response time | | Responses per second over time | a failures series per second bucket, so you see *when* they happened | | Errors table | each distinct message with its count and its share of all errors | | Request detail page | `Min`, each percentile, `Max`, `Mean` and `Std Dev` split into `Total`, `OK` and `KO` columns | Two of those rows deserve spelling out. First, **the ranges panel classifies by status first**. A request that failed in 40 ms does not land in the fast bar — it lands in the failed bar, whatever its response time. The console summary Gatling prints at the end of the run uses the same four buckets with literal labels `OK: t < 800 ms`, `OK: 800 ms <= t < 1200 ms`, `OK: t >= 1200 ms` and `KO`, and the `OK:` prefixes are there for exactly this reason. Second, **the Stats table and the chart are computed over different populations**. The four percentile columns on the global page are computed over the `Total` sample — successes and failures together — while the chart takes `OK` only. Finding that the table says 4,200 ms at the 99th percentile while the chart's 99th line never leaves 300 ms is not a bug in either; it is the two populations showing through. The per-request detail pages remove the ambiguity by printing `Total`, `OK` and `KO` side by side for every statistic. ## Reading the report in an order that does not lie 1. Start at **`% KO`** on the Stats table's first row. If it is not effectively zero, no response-time figure in the report describes the run as a whole yet. 2. Open the **Errors** table and see whether the failures are one cause or many. 3. Look at **Responses per second over time** to place the failures on the clock — a band at the end reads very differently from a scatter across the run. 4. Only then read the **percentiles chart**, and read it as what it is: the shape of the requests that succeeded. 5. Cross-check against the **Stats table's response-time columns**, remembering that on the global page they include the failures. ## The border What a percentile *means*, whether per-interval percentiles can be combined into a single run figure, how to count timeouts honestly and what the difference between offered and achieved throughput implies are performance-testing questions, and they belong to performance-testing theory rather than to Gatling. What Gatling owns — and what an interviewer is testing here — is narrower and entirely checkable: **which records feed which panel**. Answer that, and the flattering chart stops being a surprise.
- Where can you read the response times of the failed requests themselves?On a request's detail page. Its stats table prints `Min`, each configured percentile, `Max`, `Mean` and `Std Dev` in three columns — `Total`, `OK` and `KO` — so the `KO` column is the failures' own distribution. The global page shows only the `Total` figure, which merges the two.
- Which panel tells you when in the run the failures happened?**Responses per second over time**, which plots successes and failures as separate series per time bucket. The Errors table gives you counts and messages but no time axis, and the percentiles chart cannot help because the failures are not in it.
- Does the same OK-only rule apply to group charts?Yes. A group's detail page draws both **Group Duration Percentiles over Time (OK)** and **Group Cumulated Response Time Percentiles over Time (OK)**, and both are built from successful group records only. The distribution charts beside them do split successes and failures.
It is the same distortion as publishing the average finishing time of a race after quietly dropping everyone who did not finish. The survivors look fast precisely because the slow and the broken left the data set.
saying these in an interview costs you the question
- Reads the percentiles chart as covering every request in the run
- Assumes a fast failure is counted in the sub-800 ms range bar
- Expects the Stats table 95th pct to match the chart's 95th line
- Treats a falling percentile curve as proof the target improved