GitLab CI jobs on an autoscaled runner fleet almost never get a cache hit, even though `cache:` is configured correctly. Why, and what fixes it?
answer
- the YAML names the key, the runner picks the place
- default storage lives on the machine that ran the job
- disposable compute plus local cache equals永never a hit
- object storage, and shared across the fleet
basics
~20 sBy default GitLab Runner stores caches locally on the machine that ran the job, so a fleet that creates a fresh instance or pod per job never finds the previous cache. Configuring distributed cache storage in [runners.cache] with Shared enabled puts caches in object storage instead.
solid answer
~40 s`cache:` in `.gitlab-ci.yml` only declares *what* to save and under which key; *where* it lands is the runner's job. With the default local configuration, the docker executor keeps caches in a volume on the runner host, so hits only happen when the next job lands on the same machine. On an autoscaled fleet — a VM or Kubernetes pod created per job — that machine is gone, so every job cold-starts and the cache silently becomes upload-only cost. The fix is distributed cache: a `[runners.cache]` block with `Type = "s3"` (or another supported object store), the bucket details under `[runners.cache.s3]`, and `Shared = true` so runners in the fleet read each other's entries. Watch out that latency to the bucket can exceed the time to just reinstall dependencies for small caches.
code
toml · 13 lines[[runners]]
name = "autoscaled-fleet"
url = "https://gitlab.example.com/"
token = "glrt-EXAMPLE"
executor = "kubernetes"
[runners.cache]
Type = "s3"
Shared = true
[runners.cache.s3]
ServerAddress = "s3.amazonaws.com"
BucketName = "gitlab-runner-cache"
BucketLocation = "eu-central-1"
AuthenticationType = "iam"go deeper
Know that cache: declares what to save, and that the runner decides where it is stored — so a cache is not automatically available on another machine.
Explain the default local storage per runner host, why disposable compute never hits it, and what a distributed cache configuration with object storage changes.
Diagnose from the job log rather than the YAML, weigh transfer cost against reinstall cost, and treat a shared cache as a write surface that untrusted pipelines must not be able to poison.
Own the economics and the trust model of the cache tier: lifecycle policies and storage spend, which pipelines may write shared entries, and when promotion of a built artifact must replace caching entirely.
## Two halves again: the key and the location GitLab's caching splits cleanly. `.gitlab-ci.yml` says what to cache and how to name it; the runner says where that archive physically lives. Job authors only ever see the first half, which is why "the cache config is correct" and "there are no cache hits" are perfectly compatible statements. After the script finishes, the runner's helper archives the cached paths and stores the archive under a path derived from the key. On the next job with the same key, it looks the archive up and extracts it before the script runs. Everything interesting happens in the lookup. ## Where the archive goes by default With no cache configuration on the runner, storage is **local to the runner host**. For the docker executor that means a volume on the machine running the daemon; for the shell executor a directory under the runner's cache path. The consequence is simple: a cache hit requires the next job to run on the *same host*. That works fine for a couple of long-lived Docker runners — in fact it works very well, because there is no network transfer at all. It fails completely for: - **Autoscaled fleets** where a VM is created for a job and destroyed after it. - **The Kubernetes executor**, where each job is a fresh pod, possibly on a fresh node. - **Any multi-host fleet**, where the job simply lands somewhere else half the time and hit rate becomes a coin toss. The failure is quiet. The job log shows the cache being created and uploaded every time, the pipeline is merely slower than it should be, and nobody notices because nothing is red. ## Distributed cache The fix is to point the runner at object storage: ```toml [[runners]] name = "k8s-runner" executor = "kubernetes" [runners.cache] Type = "s3" Shared = true [runners.cache.s3] ServerAddress = "s3.amazonaws.com" BucketName = "gitlab-runner-cache" BucketLocation = "eu-central-1" AuthenticationType = "iam" ``` Two details carry the weight. `Type` selects the backend — S3-compatible storage (including MinIO) is the common one, with GCS and Azure also supported. `Shared = true` removes the runner-token component from the cache object path, so *every* runner in the fleet reads and writes the same entries; leave it false and each runner effectively keeps its own private namespace in the same bucket, which reproduces the original problem with an extra network hop. Credentials should come from an instance role or workload identity (`AuthenticationType = "iam"`) rather than static keys pasted into `config.toml`, since that file sits on every runner host. ## The tradeoffs nobody mentions Distributed cache is not free: - **Latency and egress.** Every job downloads and re-uploads an archive. For a 2 GB `node_modules`, that can be slower than a fresh install from a nearby package mirror. Measure before assuming caching helps. - **Unbounded growth.** Object storage keeps every key forever unless you add a lifecycle policy to expire old objects. Cache buckets quietly become the largest line on a storage bill. - **Trust.** Anything a job can write, a later job will execute — cache poisoning is real. Scope caches so an untrusted merge-request pipeline cannot write an entry that a protected-branch job will then restore, and never cache build outputs you intend to ship: promote a built artifact instead of trusting a cache to reproduce it. - **`Shared = true` means shared.** Two projects using the same key now collide. Keys must carry enough identity to keep projects and branches apart. ## Diagnosing it Read the job log rather than guessing. It states whether a cache was found for the key and where it was fetched from, and whether the archive was created and uploaded at the end. Two consecutive jobs on the same branch that both report creating the cache and never restoring it are telling you the storage is local while the compute is disposable. Cross-check with the executor: if the fleet creates a pod or VM per job, local cache was never going to work.
- What does Shared = true change in a runner's cache configuration?It removes the per-runner token from the object path, so every runner using that bucket reads and writes the same cache entries. Without it, each runner keeps a separate namespace inside the shared bucket, so an autoscaled fleet still misses despite paying for the network round trip. It also means keys must distinguish projects and branches themselves.
- When is caching dependencies in CI actually slower than not caching?When the archive is large and the object store is far away. Downloading and re-uploading a multi-gigabyte archive over the network can exceed the cost of installing from a nearby package mirror, especially with an incremental installer. Measure both paths; the honest answer is sometimes to cache a smaller, denser directory or nothing at all.
- Why should a cache never be trusted to carry a build artifact between stages?Because a cache is a speed optimisation with no integrity guarantee: it can be missing, stale, or written by a less-trusted pipeline. Use `artifacts:` to pass build output between jobs and promote one built artifact through environments. Restoring a shipped binary from a cache means what you deploy is whatever the cache happened to hold.
saying these in an interview costs you the question
- Assuming cache: alone makes caches shared across runners
- Confusing cache with artifacts for passing build output
- Setting up a bucket but leaving Shared disabled
- Never expiring cache objects and blaming storage cost on logs
- Ignoring that an untrusted pipeline can poison a shared cache