skip to content

What are ephemeral and just-in-time GitHub Actions runners, and when would you use them?

level: seniorimportance: nice to knowfreq 42%

answer

  1. One job, then it deregisters
  2. The machine must go too, not just the agent
  3. An API call replaces the registration token
  4. Single-use configuration suits autoscalers
  5. What you lose is warmth

basics

~20 s

An ephemeral GitHub Actions runner accepts exactly one job and then deregisters, so nothing carries over. A just-in-time runner goes further: GitHub's API issues a single-use configuration for a runner that has not existed before, so autoscalers can create one per queued job.

solid answer

~50 s

Registering with `./config.sh --ephemeral` makes the runner take one job and then unregister itself; you pair it with a disposable VM or container so the machine dies too. **Just-in-time (JIT) runners** remove the registration-token dance: you call the REST API endpoint that generates a JIT configuration (`POST /repos/{owner}/{repo}/actions/runners/generate-jitconfig`, with an organization equivalent), receive an encoded config, and start the agent with `./run.sh --jitconfig <value>`. That configuration is single-use and inherently ephemeral, which is exactly what an autoscaler needs: it can react to a `workflow_job` webhook, launch a container with a fresh JIT config, and let it evaporate after the job. This is the model behind Actions Runner Controller on Kubernetes. The tradeoff is cold start on every job — no warm caches, no pre-pulled images — so you either pre-bake the image or accept the latency.

code

bash · 8 lines
bash
# ask GitHub for a single-use runner configuration
gh api --method POST /repos/acme/api/actions/runners/generate-jitconfig \
  -f name=jit-runner-1 \
  -F runner_group_id=1 \
  -f 'labels[]=self-hosted' -f 'labels[]=linux'

# start the agent with it; it takes one job and disappears
./run.sh --jitconfig "$ENCODED_JIT_CONFIG"

go deeper

for a junior

Know that a runner can be configured to accept a single job and then remove itself, which keeps jobs from affecting each other.

for a middle

Explain the difference between the --ephemeral flag and an API-issued single-use JIT configuration, and why the machine must be discarded too.

for a senior

Argue the tradeoff in production terms: isolation and elasticity against cold-start latency and lost cache locality, plus the workflow_job-driven autoscaling loop.

for a principal

Own the fleet architecture — per-job pods versus a warm pool, image baking strategy, API rate limits at scale, and which workloads justify keeping any persistent runners at all.

## Two related ideas **Ephemeral runner.** A runner configured with `--ephemeral` picks up one job, completes it, and then removes its own registration. The agent process exits; it will never take a second job. The important discipline is that ephemerality must extend to the *machine*: if you restart the same VM with the same disk, you have merely re-registered a dirty host and thrown away the benefit. **Just-in-time runner.** JIT configuration is the API-native way to create such a runner. Instead of minting a short-lived registration token and running `config.sh`, you ask GitHub for a complete, single-use runner configuration: ``` gh api --method POST /repos/acme/api/actions/runners/generate-jitconfig \ -f name=jit-1 -F runner_group_id=1 -f 'labels[]=self-hosted' -f 'labels[]=linux' ``` The response carries an encoded configuration blob you hand straight to the agent via `--jitconfig`. No token file is written, no `config.sh` step, no removal step — and the configuration cannot be reused to bring up a second runner. ## Why this matters **Security.** This is the primary driver. A persistent runner lets job N influence job N+1 through the workspace, `PATH`, dependency caches, background processes or leftover credentials. A one-job-and-gone runner removes that channel entirely, which is the standard answer to "how do I self-host safely". **Correctness and flakiness.** Persistent runners drift. Disk fills with old workspaces and container images; a half-finished job leaves a lock file; two runners on one host contend for the Docker daemon or a fixed port. A large share of "works on my machine, flakes in CI" on self-hosted fleets is state, and ephemerality is the cure. **Elasticity.** Ephemeral and JIT runners are the substrate for autoscaling. GitHub emits a `workflow_job` webhook when a job is queued; a controller creates one runner for that job and lets it disappear afterwards. Actions Runner Controller implements this on Kubernetes, giving a pod per job. Persistent runners cannot scale to zero, because a scaled-down runner might be mid-job. ## The costs, honestly - **Cold start.** Every job pays for provisioning plus any tool installation. Mitigations: bake tooling into the runner image, keep a small warm pool of pre-created runners, or lean harder on `actions/cache` — which now has to fetch over the network instead of hitting a warm local directory. - **Cache locality lost.** The big attraction of a persistent self-hosted runner is a hot dependency directory and pre-pulled images. You trade that for isolation; for many teams a shared cache service or a well-designed `actions/cache` key recovers most of it. - **Operational surface.** Something must now create runners: an autoscaler, a controller, or a cloud API loop. That is a system to run and monitor. Naive designs also hit API rate limits when they mint configuration per job at high volume. ## Choosing Use ephemeral/JIT when the workload is untrusted (public repositories, many contributors), when jobs need clean isolation, or when demand is spiky enough that scale-to-zero matters. Keep persistent runners for a small, trusted, homogeneous internal fleet where the warm state genuinely pays — and even then, wipe the workspace and monitor drift.

  • Is an ephemeral runner useful if you reuse the same virtual machine for the next runner?
    Barely. Deregistering the agent while keeping the disk means the next job still inherits the workspace, caches, installed tools and any planted binary. Ephemerality is a property of the whole execution environment: destroy the VM or container, or at minimum recreate from a known image.
  • What does an autoscaler use to know a runner is needed?
    The workflow_job webhook, which GitHub delivers when a job is queued, including its requested labels. A controller matches those labels to a runner pool and creates one runner for the job. Actions Runner Controller does exactly this on Kubernetes, scheduling a pod per queued job and scaling to zero when idle.
  • What is the main performance cost you must design around?
    Cold start. Every job provisions a fresh environment, so tool installation and image pulls repeat. Bake the toolchain into the runner image, keep a small warm pool to absorb bursts, and make actions/cache keys good enough that the network fetch is cheaper than rebuilding.

saying these in an interview costs you the question

  • Thinking ephemeral means the workspace is merely cleaned
  • Reusing the same disk under a new registration
  • Expecting warm caches on a per-job container
  • Assuming persistent runners can scale to zero safely
  • Confusing a JIT config with a reusable registration token

context