How would you tune a GitHub merge queue's group size, concurrency, and wait time?
answer
- it is a queueing problem, so measure
- arrivals per CI cycle
- big batches lose big when red
- latency bought with runner minutes
- fix the flakes before the knobs
basics
~20 sSize the settings from two measurements: merge attempts per hour and CI duration, then adjust for failure rate. Larger groups cut runs but multiply rework when red; more concurrent builds cut latency but cost runner minutes; wait time trades a little latency for fuller batches.
solid answer
~50 sTreat it as a queueing problem with a measured arrival rate. Start from merge attempts per hour and end-to-end CI duration: if arrivals per CI cycle are near one, a small queue with low concurrency is enough; if they are five, you need roughly that many groups in flight to avoid a growing backlog. Then temper it with your **failure rate**, because that is what makes the settings interact — large groups minimise CI runs when everything is green and maximise discarded work when something is red, so a repository that evicts entries often should batch less, not more. Set the minimum-entries and wait-time knobs to batch arrivals during busy periods without adding latency when the queue is empty, and set the check timeout comfortably above your p99 CI duration so slow runs are not mistaken for dead ones. Above all, fix flaky tests first: in a queue every flake evicts an innocent pull request and throws away the speculative work behind it.
code
json · 12 lines{
"type": "merge_queue",
"parameters": {
"merge_method": "SQUASH",
"grouping_strategy": "ALLGREEN",
"max_entries_to_build": 5,
"min_entries_to_merge": 1,
"max_entries_to_merge": 2,
"min_entries_to_merge_wait_minutes": 2,
"check_response_timeout_minutes": 60
}
}go deeper
Recall that a merge queue has settings controlling how many pull requests are tested together and how many groups build at once, and that both affect CI cost.
Explain the direction of each knob: concurrency reduces waiting but adds runs, group size reduces runs but complicates blame, and the timeout must exceed real CI duration.
Derive settings from measured arrival rate, CI duration and failure rate, and argue why a flaky suite must be fixed before a queue is enabled at all.
Own the tradeoff end to end: what queue latency the organization will accept, what CI spend buys it, when to shrink the merge-group suite instead, and when the real answer is faster CI or a different repository boundary.
## Measure before you tune Four numbers determine every setting, and none of them are guessable: 1. **Arrival rate** — merge attempts per hour at peak, not on average. Queues fail at peak. 2. **CI duration** for the required checks on a merge group, at p50 and p99. The tail matters more than the median. 3. **Failure rate** — the share of entries that are genuinely red, split from the share that are flaky. These have different fixes and must be counted separately. 4. **Runner capacity and cost** — how many concurrent jobs you can actually run, and what a wasted group costs. The core ratio is *arrivals per CI cycle*: arrival rate multiplied by CI duration. Below one, a queue is barely doing anything and simple required checks with an up-to-date requirement would serve. Around three to five, the queue is genuinely earning its keep. Well above that, no queue configuration saves you and the real answer is to make CI faster or to split the repository. ## Concurrency: how many groups in flight Concurrency is the number of merge groups GitHub may build simultaneously. It converts queue depth into parallel CI runs, so it is the knob that fixes latency. If arrivals per CI cycle is `k`, you need at least `k` groups in flight to keep the backlog from growing; below that the queue lengthens without bound at peak. Above it, you are paying for speculation that rarely pays off, and every failure discards more in-flight work. Start at roughly the measured `k`, watch queue depth over a week, and adjust. The cost is direct: a queue that is `n` deep with full concurrency runs up to `n` CI builds to land `n` changes, versus one build per change without speculation. You are buying developer latency with runner minutes, and the exchange rate should be an explicit decision, not a default. ## Group size: how many entries per run Batching several entries into one group is the opposite lever: it reduces runs. Its cost appears only on failure, and it is sharp. A red batch does not identify its culprit, so the queue must re-form smaller groups to isolate it — extra runs, extra latency, and innocent authors waiting. The expected cost of a batch scales with the probability that *any* member is red, which grows with size. So: - **Low failure rate, expensive CI** → batch more. You rarely pay the isolation cost, and you save real money. - **Non-trivial failure rate** → batch less, ideally one entry per group, so failures are attributed instantly. - **Never batch to compensate for slow CI.** That trades a latency problem for an attribution problem and usually makes both worse. ## Wait time and minimum entries A short wait before forming a group lets several arrivals batch together during busy periods. Its cost is added latency for the first arrival in a quiet period. Keep it small — on the order of a minute or two — and pair it with a minimum-entries setting so the queue does not sit waiting when only one change is present. This knob is a rounding error compared with concurrency and group size; do not spend much on it. ## Check timeout The check response timeout must exceed your p99 CI duration with margin, or you will evict entries whose runs were merely slow, which looks exactly like flakiness and sends people hunting the wrong problem. But it must not be so long that a genuinely dead pipeline wedges the queue for hours. Set it from the measured tail, and revisit it whenever CI duration changes materially. ## The precondition nobody wants to hear Flaky tests dominate all of this. In a queue, a flake does not cost one re-run — it evicts a correct pull request, discards every speculative group behind it, and returns the author to the back of the line. The cost grows with queue depth, which means it is worst exactly when the queue matters most. Before tuning anything, quarantine flaky tests and measure the residual rate. A repository above a few percent flake rate will experience the queue as an unfair random-eviction machine no matter how the knobs are set. ## Consider what runs in the group at all A legitimate lever that is often forgotten: the merge group does not have to run the same suite as the pull request. A fast, high-signal subset that catches integration conflicts can gate the queue while slower suites run on the pull request or post-merge. That reduces both latency and cost more than any knob, at the price of a narrower guarantee — decide it deliberately. ## Interview framing Refuse to give fixed numbers, and say why: the settings are a function of arrival rate, CI duration, failure rate and runner capacity, and any figure without those is cargo-culting. Then give the shape of the answer — concurrency sized to arrivals per CI cycle, group size inversely to failure rate, timeout above p99, wait time small — and name flakiness as the precondition and CI duration as the lever with the largest effect.
- Queue depth grows through every peak no matter what you set. What is the real problem?Arrivals per CI cycle exceed what any concurrency setting can absorb. Once you are building as many groups as the runners allow and the backlog still grows, the levers left are outside the queue: make CI faster, shrink the suite that gates merge groups, or split the repository so merges do not all contend for one branch.
- Should the merge group run exactly the same test suite as the pull request?Not necessarily. A fast, high-signal subset that catches integration and semantic conflicts can gate the queue while slower suites run on the pull request or after merge. That cuts both latency and cost more than any setting, at the price of a narrower guarantee — so make it an explicit decision, not an accident.
- What goes wrong if the check response timeout is set below your p99 CI duration?Entries whose runs are merely slow get evicted as if they had failed. The pattern looks random and is easily mistaken for flakiness, so people hunt tests instead of configuration. Set the timeout from the measured tail with margin, and revisit it whenever CI duration changes materially.
saying these in an interview costs you the question
- Quotes fixed settings without measuring anything
- Increases batch size to compensate for a high failure rate
- Ignores that speculative work is discarded on failure
- Sets the check timeout below the p99 CI duration
- Treats the queue as a fix for slow or flaky CI