skip to content

In a GitHub merge queue, what happens when a speculative merge group's checks fail?

level: seniorimportance: should knowfreq 36%

answer

  1. nothing red is allowed to land
  2. the offender goes back to its author
  3. what was built on a false assumption
  4. batching makes blame ambiguous
  5. an entry that never reports must still leave

basics

~20 s

GitHub discards that group, removes the offending pull request from the queue and tells its author, then rebuilds the entries behind it without that change. Anything built speculatively on top of the failed group was based on a false assumption and is re-tested.

solid answer

~50 s

A red merge group never merges. GitHub drops the pull request responsible from the queue — it returns to the normal pull-request state for its author to fix and re-queue — and because every group built behind it assumed it would land, those speculative groups are invalidated and re-formed against the corrected order. Entries **ahead** of the failure are unaffected, since their groups did not contain the bad change. When a group batches several entries, GitHub cannot immediately tell which one broke it, so it re-forms smaller groups to isolate the culprit, which costs additional CI runs — the price of batching. The same removal happens when required checks never report within the configured check timeout, so an entry cannot wedge the queue indefinitely. The operational implication is that flaky tests are expensive here: each flake evicts an innocent pull request and forces a rebuild of everything behind it.

go deeper

for a junior

Recall that a failed queue run means nothing merges, the pull request leaves the queue, and its author fixes and re-queues it.

for a middle

Explain why entries behind the failure must be rebuilt while entries ahead are unaffected, and what happens when a group contains several pull requests at once.

for a senior

Diagnose repeated evictions in practice — genuine conflict versus flakiness versus checks that never report — and argue for group size based on your failure rate rather than merge volume.

for a principal

Own the economics: what rework a failed speculative group costs in runner minutes, what flake rate makes a queue counterproductive, and what you fix before turning one on.

## The invariant being protected The queue exists to guarantee one thing: nothing lands on the base branch unless the exact combination that will result was tested and green. Every failure behaviour follows from defending that invariant. A red group therefore cannot merge, and no group built on top of a red group can be trusted either, because its content included changes that are not going to land. ## What happens to the failing entry The pull request is **removed from the queue**. It goes back to being an ordinary open pull request; the author is notified, the queue view shows why it was dequeued, and the temporary `gh-readonly-queue` ref for the group is cleaned up. Nothing about the base branch changed — the whole point is that the bad combination never merged. The author fixes the problem on the pull request branch and queues it again, which places it at the back. ## What happens to everyone else - **Entries ahead of the failure** are untouched. Their groups did not include the failing change, so their results are still valid and they continue to merge in order. - **Entries behind the failure** were built speculatively on a base that included the failing change. Those groups are invalidated: GitHub discards them and re-forms the queue without the removed entry, which means new temporary refs and new CI runs. Work in flight is thrown away, which is precisely the cost of the speculation that bought you throughput in the first place. ## Batched groups and isolation If the queue is configured to merge more than one entry per group, a red result is ambiguous: any of the batched pull requests could be at fault, or the combination of two of them. GitHub responds by re-forming smaller groups so the culprit can be isolated, at the cost of extra runs. This is the fundamental batching tradeoff and it is worth stating explicitly in an interview: **large groups minimise CI runs when everything passes and maximise rework when something fails.** The right size therefore depends on your failure rate, not on your merge rate alone — a repository with frequent red entries should batch less, because it will be paying the isolation cost constantly. ## Timeouts and stalled entries A merge group whose required checks never report cannot be allowed to hold the queue forever. The queue has a **check response timeout**; when it expires, the entry is removed exactly as a failing one would be. This is what stops one misconfigured pipeline from freezing every merge in the repository, and it is also the symptom you see when CI does not run for merge groups at all: every entry marches into the queue, waits, and is evicted with no results. ## Why flakiness is disproportionately expensive On a pull request, a flaky test costs one re-run by an annoyed developer. In a queue, a flake in a merge group evicts a pull request that was entirely correct, discards every speculative group behind it, and puts the author at the back of the line. The cost is superlinear in queue depth, and the developer experience is worse than the raw numbers suggest because the eviction looks arbitrary. A repository with a meaningful flake rate should quarantine or fix those tests before enabling a queue; if it cannot, the queue will be blamed for problems it merely exposed. ## Operating a queue that keeps failing A few diagnostic habits: - Read the queue view for the *pattern*: one entry evicted repeatedly is a genuine bug; different entries evicted at random is flakiness or an unstable environment. - Distinguish evictions caused by check failures from evictions caused by the timeout — they have completely different fixes. - Watch whether groups are being discarded because of failures ahead rather than their own results; a single bad entry near the front can cause a burst of rebuilds that looks like widespread breakage. - If the same combination keeps failing while each entry is green alone, you have found a real semantic conflict, which is the queue doing its job. ## Interview framing State the invariant, then the three consequences: the offending entry is removed and returns to its author, entries ahead are unaffected, entries behind are invalidated and rebuilt. Add the batching-isolation cost and the timeout eviction, and finish with the flakiness point — that is the part that comes from having actually operated one.

  • Entries are being evicted at random rather than the same one repeatedly. What does that suggest?
    Flakiness or an unstable test environment rather than a genuine defect. A real bug evicts the same pull request every time it is queued; a flake takes out whichever entry happened to be in the group when the test misbehaved. Fix or quarantine the unstable tests, because in a queue each flake also discards the speculative work behind it.
  • Why does the queue evict an entry whose checks never report at all?
    Because otherwise one misconfigured pipeline would freeze every merge in the repository. The check response timeout bounds the wait: when it expires the entry is removed exactly as a failing one would be, the temporary ref is cleaned up, and the rest of the queue continues.
  • How do you decide how many pull requests to put in one merge group?
    From your failure rate rather than your merge rate. Large groups save CI runs when everything passes but force isolation runs whenever anything is red, and the chance that some member is red grows with size. Frequent failures argue for one entry per group so blame is instant; a very stable suite with expensive CI can batch more.

saying these in an interview costs you the question

  • Says the failing change merges and is reverted afterwards
  • Thinks entries ahead of the failure are rebuilt too
  • Believes a batched failure pinpoints the culprit for free
  • Assumes a stalled entry blocks the queue forever
  • Treats flaky tests as a minor annoyance inside a queue

context