skip to content

In GitLab Runner's config.toml, what is the difference between the global `concurrent` setting and a runner entry's `limit`?

level: seniorimportance: should knowfreq 38%

answer

  1. one process, several registrations
  2. top-level key versus per-entry key
  3. the smaller of the two wins
  4. polling throughput is a third, separate knob

basics

~20 s

concurrent is a process-wide ceiling on jobs running across every runner entry in that config.toml, while limit caps how many jobs one specific [[runners]] entry may run at once. The effective parallelism is the smaller of the two.

solid answer

~50 s

One `gitlab-runner` process can serve several runner entries, and the two settings sit at different levels. `concurrent`, at the top of `config.toml`, is the total number of jobs the whole process will execute simultaneously — it is the machine's capacity guard. `limit`, inside a `[[runners]]` block, caps that one registration, so you can say "this runner may take at most four jobs even though the host allows sixteen". The two multiply badly if you only set one: leaving `concurrent = 1` (the install default) means a fleet of runner entries still runs one job at a time, and raising `limit` alone changes nothing. There is a third knob, `request_concurrency`, which controls how many job requests the runner fetches from GitLab in parallel — a queue-polling throughput setting, not an execution one. Size `concurrent` against real CPU, memory and disk on the host, not against how many jobs are waiting.

code

toml · 19 lines
toml
concurrent = 10
check_interval = 3

[[runners]]
  name = "docker-general"
  url = "https://gitlab.example.com/"
  token = "glrt-EXAMPLE"
  executor = "docker"
  limit = 8
  request_concurrency = 2
  [runners.docker]
    image = "alpine:3.20"

[[runners]]
  name = "deploy-protected"
  url = "https://gitlab.example.com/"
  token = "glrt-EXAMPLE2"
  executor = "shell"
  limit = 1

go deeper

for a junior

Know that a runner can execute more than one job at a time and that this is configured in config.toml, with a global setting and a per-runner setting.

for a middle

Explain that concurrent is process-wide and limit is per registration, that the effective value is the smaller, and that the install default of 1 is why untuned runners feel slow.

for a senior

Demonstrate sizing from host CPU, memory, disk and executor type, and recognise the failure signatures — OOM-killed builds, disk exhaustion, uniformly slow jobs — that come from over-commitment.

for a principal

Own capacity strategy across the fleet: where to spend on autoscaling versus fixed hosts, how per-registration limits act as spend controls, and how deploy serialisation is guaranteed independently of any one runner's configuration.

## The shape of config.toml ```toml concurrent = 10 check_interval = 3 [[runners]] name = "docker-general" url = "https://gitlab.example.com/" token = "glrt-EXAMPLE" executor = "docker" limit = 8 request_concurrency = 2 [runners.docker] image = "alpine:3.20" [[runners]] name = "deploy-protected" url = "https://gitlab.example.com/" token = "glrt-EXAMPLE2" executor = "shell" limit = 1 ``` One process, one file, two registrations. Understanding which setting lives at which indentation level is most of the answer. ## concurrent — the process ceiling `concurrent` is a top-level key. It is the maximum number of jobs this runner *process* will have in flight at any moment, summed across every `[[runners]]` entry. If it is 10 and both entries above are busy, the process stops asking GitLab for work once ten jobs are running, no matter how much either entry's own `limit` would allow. The default in a fresh installation is `1`. That single line explains a large fraction of "our runner is so slow" reports: an eight-core machine dutifully processing one job at a time while a queue builds up. ## limit — the per-registration ceiling `limit` lives inside a `[[runners]]` block and caps that registration alone. It exists because different registrations have different costs and different risks: - A deployment runner set to `limit = 1` serialises production deploys, so two pipelines cannot race each other onto the same environment. - A memory-hungry integration-test runner might be held at 2 while a lint runner on the same host is allowed 8. - For autoscaling executors, `limit` bounds how many machines or pods can exist for that registration — it becomes a spend control, not just a parallelism one. The effective parallelism for one entry is `min(limit, concurrent - jobs running elsewhere)`. Raising `limit` without raising `concurrent` accomplishes nothing. ## request_concurrency — a different axis entirely `request_concurrency` (default 1) is how many job *requests* the runner will make to GitLab in parallel. It affects how quickly a runner can fill empty slots when a burst of jobs appears, not how many it can execute. On a runner with a high `concurrent` value and a bursty queue, leaving it at 1 makes slot-filling serial and the fleet looks underused; raising it too far adds API load for no benefit. Related is `check_interval`, the seconds between polls for new work — lowering it reduces pickup latency at the cost of more requests. ## Sizing it honestly The number to set is a function of the host, not the backlog: - **CPU.** Most CI jobs are compile- or test-bound and will happily use every core. Over-committing turns a two-minute job into a six-minute one for everybody, which looks like a broken runner rather than a saturated one. - **Memory.** The binding constraint more often than CPU: a JVM or bundler job with a large heap times four concurrent jobs exceeds the box, and the kernel's OOM killer starts terminating builds — which surfaces as random, unreproducible job failures. - **Disk and I/O.** Each concurrent job has its own build directory, cache extraction and image layers. Parallel jobs contend for I/O and for disk space; a full disk fails every job on the host at once. - **Executor.** With the Kubernetes executor, `concurrent` is a limit on simultaneous pods and the real capacity question moves to the cluster's resource requests and autoscaler. ## Symptoms to recognise - Jobs pending while the runner sits at low CPU: `concurrent` too low, or `request_concurrency` throttling slot-filling. - Every job slow and timing out together: `concurrent` too high for the host, everything contending. - Random OOM kills or "no space left on device" appearing only under load: concurrency exceeding memory or disk, not a code problem. - Two deploys colliding on one environment: a deployment runner without `limit = 1`, and no resource-group serialisation in the pipeline.

  • What does request_concurrency change, and why is it not a parallelism setting?
    It sets how many job requests the runner sends to GitLab at once, defaulting to 1. It affects how fast empty execution slots are filled during a burst, not how many jobs may run. A runner with `concurrent = 20` and `request_concurrency = 1` fills slots serially and can look underused while a queue exists.
  • How would you stop two pipelines deploying to the same environment simultaneously?
    Set `limit = 1` on the deployment runner so it executes one job at a time, and pair it with a `resource_group` on the deploy job so GitLab itself serialises those jobs even if capacity grows later. The runner limit is the machine-level guard; the resource group is the pipeline-level guarantee that survives a fleet change.
  • What symptoms suggest concurrent is set too high for the host?
    Jobs that all slow down together rather than queueing, timeouts appearing only at peak, random OOM-killed processes with no code change, and "no space left on device" under load. The tell is that failures correlate with load rather than with a particular job, and CPU or memory on the host is pinned while throughput drops.

saying these in an interview costs you the question

  • Thinking limit alone raises a runner's parallelism
  • Leaving concurrent at the default 1 and blaming GitLab for slow CI
  • Confusing request_concurrency with how many jobs run at once
  • Sizing concurrency from the queue length instead of host resources
  • Assuming concurrent applies per registered runner rather than per process

context