Across a team's JMeter plans, how would you decide each Thread Group's sampler-error action?
answer
- Read the five choices as a scope ladder
- Fail fast in setUp, hold the load schedule
- Only setUp should be allowed to abort
- Check the tearDown-on-shutdown box first
- Prefer a visible element to a hidden button
basics
~20 sTreat it as a blast-radius decision. Let setUp groups fail fast with Stop Test, keep main load groups on Continue so the run holds its schedule, and express any narrower stop as a visible element rather than a group-wide policy.
solid answer
~50 sThe five choices form a ladder of blast radius: `Continue`, `Start Next Thread Loop`, `Stop Thread`, `Stop Test`, `Stop Test Now`. My default position is that **only a setUp Thread Group may end a run**. Give it `Stop Test`: if its sampler fails, the engine never starts the main Thread Groups at all. Main load groups stay on `Continue`, because anything else lets one thread's failure shorten the run and leave you with a results file covering less than the window you scheduled. Where a specific step genuinely must not be survived, express it locally — a Result Status Action Handler on that sampler, or an If Controller around a Flow Control Action — so the rule is visible next to the thing it guards. And check the Test Plan's tearDown-on-shutdown box before relying on any stop, or cleanup silently stops happening.
go deeper
Know that the choice is per Thread Group and that Continue is the default, so an unedited plan never stops itself on a failure.
Explain what each choice costs when it fires, and why a setUp group and a load group want different answers.
Show the operational reasoning: a truncated run, a skipped tearDown, an interrupt that reaches nothing, and how you would spot each after the fact.
Own the policy: who may abort a run, whether the rule is visible in the tree, and how you keep the setting from drifting across a repository of plans.
## Frame it as blast radius, not as strictness The five values of **Action to be taken after a Sampler error** are not a severity scale; they are a scope scale. Reading them that way makes the decision tractable: | Action | What it costs when it fires | |---|---| | `Continue` | nothing stops; you keep the sample and the schedule | | `Start Next Thread Loop` | one iteration of one thread | | `Stop Thread` | one virtual user, for the rest of the run | | `Stop Test` | everybody's run, gracefully | | `Stop Test Now` | everybody's run, with samplers interrupted mid-flight | The last two are the only ones where a single thread's bad luck changes what every other thread was doing. That asymmetry is the whole decision. ## A defensible default per group type 1. **setUp Thread Groups: `Stop Test`.** This is the one place fail-fast is unambiguously right. When a setUp sampler fails and raises a stop, the engine's `running` flag goes false, and the loop that starts the main Thread Groups is guarded by that flag, so none of them is ever started. A broken login, a missing seed data set or an unreachable environment costs you seconds instead of an hour. 2. **Main load Thread Groups: `Continue`.** A load run's job is to hold its schedule. Any other setting lets a transient failure truncate the run, and a truncated run leaves a shorter results file and a dashboard built from it. 3. **tearDown Thread Groups: `Continue`.** A cleanup step that gives up halfway is worse than one that grinds on. ## Where narrower rules belong When one step really is load-bearing — the token fetch that everything after it depends on — do not reach for the group-wide field. Two better options: - a **Result Status Action Handler** under that sampler, so only its failure reacts, with `Start Next Thread Loop` or `Stop Thread` selected; - an **If Controller** on `${__jexl3(!${JMeterThread.last_sample_ok})}` wrapping a **Flow Control Action** — the bare variable is true when the sample *succeeded*, so a failure guard has to negate it — when the condition is richer than “the sample failed” or when you want the rule to be obvious in the tree. The review question to ask of any plan is: *can a reader see, from the tree, what will stop this run?* A `Stop Test` radio button on a Thread Group is invisible unless you click the group; a named Flow Control Action is a line in the tree. ## Consequences worth checking before you commit - **tearDown Thread Groups run after a stop only if the Test Plan's 'Run tearDown Thread Groups after shutdown of main threads' box is ticked.** A policy of `Stop Test` on the main group plus an unticked box means your cleanup silently stops happening. - **`Stop Test Now` interrupts only samplers that support interruption**, so on a plan whose samplers do not, it buys nothing over `Stop Test` and risks half-measured samples. - **Third-party `jpgc` thread groups carry the same field.** The Concurrency, Arrivals, Stepping and Ultimate groups come from the plugins project rather than the Apache download, but they extend the same base class and store the same `ThreadGroup.on_sample_error` property, so a policy written for stock groups transfers to them. - **Whether the run passed is a separate question** from whether it continued. An error action decides how much of the run happens; it does not decide the verdict. ## Making it stick The realistic failure mode is drift: each plan's setting is whatever its author last clicked, and nobody notices until a nightly run ends after four minutes. Two cheap controls help. First, make the setting part of plan review — it is one grep over the repository, since the value is a plain string in the `.jmx`: ``` grep -l 'ThreadGroup.on_sample_error">stoptest' plans/*.jmx ``` Second, agree on the one sentence that justifies each non-default choice, and put it in the element's name. `Stop Test` on a group called `Load` is a mystery; on a group called `setUp - abort if the environment is down` it is self-documenting.
- What actually happens to the main Thread Groups when a setUp group raises Stop Test?The engine's running flag is cleared, and the loop that starts the main Thread Groups is guarded by that flag, so none of them is ever started. The setUp groups are allowed to finish, and then the run ends without any load having been applied.
- Why not simply put Stop Test on every Thread Group so nothing is ever missed?Because it hands every thread the power to end everyone's run. On a plan with hundreds of threads the run's length becomes a property of the flakiest endpoint, and you lose the rest of the window you meant to measure.
- How would you audit a repository of existing plans for this setting?Grep the .jmx files for ThreadGroup.on_sample_error; it is a plain string property with five possible values. Anything other than continue on a main load group is worth a sentence of justification in the element's name.
saying these in an interview costs you the question
- Treating the five choices as a strictness dial
- Putting Stop Test on the main load group by default
- Relying on a stop while tearDown-on-shutdown is unticked
- Assuming Stop Test Now is always faster than Stop Test
- Leaving the setting to whoever edited the plan last