An RSpec suite fails in CI only with --seed 4821; how does rspec --bisect narrow the failure down, and what makes it report nothing useful?
answer
- same seed, same options
- failing examples run alone first
- halve the non-failing examples
- minimal reproduction command
- bisect_runner :fork or :shell
basics
~20 srspec --seed 4821 --bisect reruns subsets in that order: it checks the failure needs other examples first, then repeatedly discards half of the passing ones, and prints a minimal reproduction command naming the victim and the example that pollutes it.
solid answer
~40 sRun `rspec --seed 4821 --bisect` with the same options as the failing run. rspec first runs the suite to collect failing and non-failing example ids, then runs the failures alone: if they still fail it reports `failure(s) do not require any non-failures to run first`, meaning the problem is not order. Otherwise it bisects in rounds, dropping halves of the non-failing examples that are not needed, and ends with `The minimal reproduction command is:` followed by something like `rspec ./spec/sms_spec.rb[1:2] ./spec/notifier_spec.rb[1:1] --seed 4821`. Ctrl-C prints the best command so far and `--bisect=verbose` shows every run. It is useless when the failure does not reproduce locally (`No failures found`), and the default `:fork` runner can misreport independence; set `config.bisect_runner = :shell` in a file loaded with `--require` and retry.
code
bash · 8 lines$ rspec --seed 4821 --bisect
Bisect started using options: "--seed 4821"
Running suite to find failures... (38.2 seconds)
Starting bisect with 1 failing example and 311 non-failing examples.
Checking that failure(s) are order-dependent... failure appears to be order-dependent
...
The minimal reproduction command is:
rspec ./spec/sms_spec.rb[1:2] ./spec/notifier_spec.rb[1:1] --seed 4821go deeper
Know that rspec --bisect exists for failures that appear only in some random orders, and that you must pass it the failing seed.
Describe the steps: original run, dependency check, halving rounds, and the minimal reproduction command with bracketed example ids.
Interpret each non-answer: no failures found, failures that need no other examples, the fork runner's caveat, and fix the leak in the polluting example.
Build the practice into the team's CI workflow: always log the seed and options so any engineer can bisect a failing order from the job output.
## The situation A notification-service suite runs in random order. CI fails with seed 4821, and the failing example, an SMS formatting spec, passes on its own. Some example that runs earlier leaves state behind: a stubbed constant that was not restored, a class-level cache, a mutated global. With hundreds of examples, finding which one by hand is slow. `rspec --bisect` automates that search. ## What bisect does, step by step Started as `rspec --seed 4821 --bisect` (plus any other options the failing run used), it prints `Bisect started using options: "--seed 4821"` and then: 1. **Original run.** Runs the suite once to collect the ids of failing examples and non-failing ones. If nothing fails, it stops with `No failures found. Bisect only works in the presence of one or more failing examples.` 2. **Dependency check.** Runs only the failing examples. If they still fail, it reports `failure(s) do not require any non-failures to run first`: the failure is not caused by order, and bisecting cannot help. 3. **Rounds.** Otherwise it splits the non-failing examples into halves and reruns the failures with each half, keeping only what is needed to reproduce: `Round 1: bisecting over non-failing examples 1-9 .. ignoring examples 6-9`. 4. **Result.** It prints the smallest set it found: ``` The minimal reproduction command is: rspec ./spec/sms_spec.rb[1:2] ./spec/notifier_spec.rb[1:1] --seed 4821 ``` The bracketed ids (`[1:2]` means the second example in the first top-level group of that file) identify examples exactly, and because the order depends only on the seed and each id, the two examples run in the same relative order as in CI. Two controls help on long runs: - **Ctrl-C** aborts and prints `The most minimal reproduction command discovered so far is:`; - **`--bisect=verbose`** lists every example id and every command it runs. ## When bisect reports nothing useful | Output | Meaning | Next step | |---|---|---| | `No failures found` | the failure did not reproduce with these options | match CI's seed, options, environment variables and data | | `failure(s) do not require any non-failures to run first` | the failing example fails alone | debug it directly; it is not order | | same, with a `:fork` NOTE | the fork runner may be wrong about independence | set `config.bisect_runner = :shell`, retry | | a long list of examples | several polluters, or a polluter that needs company | read the command and narrow by hand | ## The two runners `config.bisect_runner` chooses how subsets run: - **`:fork`**, the default where the platform supports `fork`, boots the application once in a parent process and forks a child for each subset. It is much faster, but one-time setup must live in a `before(:suite)` hook rather than at the top level of a file loaded with `--require`, and it can report a failure as independent when it is not. - **`:shell`**, the default elsewhere, shells out and boots rspec and the application for every subset. It is slower and the most compatible. The setting only takes effect when it is made in a file loaded via `--require`, such as `spec_helper.rb` through `.rspec`. ## What bisect costs Bisect is a search, so it reruns parts of the suite many times: - the first step is one full run, as long as the original; - each round reruns the failing examples with a subset of the non-failing ones, halving the candidates, so when one example is responsible the number of rounds grows roughly with the logarithm of the suite size; - application boot is paid per subset with the `:shell` runner and once with `:fork`, which is why `:fork` is the default where available. On a suite that takes many minutes, that means starting bisect and doing something else, or using Ctrl-C to accept a larger but still useful command. Narrowing the starting set first, for example to the directories that ran before the failure, shortens every round. ## After bisect The reproduction command names the **polluter** (the example that runs first) and the **victim**. Run the pair, read what the first one changes that it does not undo, and fix the leak there, not in the victim. Then rerun the full suite with the same seed to confirm.
- Bisect reports that the failures do not require any non-failures, but the example passes when run on its own with no seed. What could explain that?With the default `:fork` runner, the forked children inherit state from a parent that already booted the application, so bisect can wrongly conclude the failure is independent. rspec prints a NOTE saying so. Set `config.bisect_runner = :shell` in a file loaded with `--require` and rerun, which boots a clean process for every subset.
- Why must the bisect run use the same options as the failing CI run, not just the same seed?Bisect reproduces the failure by rerunning with the options it was started with. A different file list, tag filter, `--order` or environment changes which examples exist or run, so the first step may find no failures at all. Copy the seed and every relevant option, then compare the `Bisect started using options:` line with CI's command.
saying these in an interview costs you the question
- bisect picks its own random seeds to search for failing orders
- bisect can find the cause of a failure that also fails when run alone
- the minimal reproduction command runs the examples in a new random order
- the :shell runner is always faster than the :fork runner
- fix an order-dependent failure in the victim, not the polluter