skip to content

Using the Jenkins Kubernetes plugin, how is a build executed, and in a pod template that defines a maven container alongside the default agent container, what decides which container an `sh` step runs in?

level: seniorimportance: should knowfreq 46%

answer

  1. one pod per build, then deleted
  2. one container is the agent itself
  3. tool containers must not exit
  4. shared emptyDir is the workspace
  5. wrap the step or set a default

basics

~20 s

The plugin creates one pod per build from a pod template; its inbound agent container connects back to the controller and the workspace lives on a volume shared by all containers. Steps run in the agent container unless wrapped in container('maven') or covered by defaultContainer.

solid answer

~60 s

You configure a Kubernetes cloud on the controller and a pod template — inline YAML in `agent { kubernetes { yaml … } }`, or a template defined in the cloud configuration and selected by label. When a build needs that agent, the plugin creates a pod in the target namespace. One container in it is the Jenkins inbound agent, named `jnlp` by default, which connects back to the controller over WebSocket or the JNLP port; the other containers hold your toolchains. The workspace is an `emptyDir` volume mounted into every container, which is what lets them see each other's files. By default `sh` steps execute in the agent container, which contains no build tools. You direct them with `container('maven') { sh 'mvn verify' }`, or set `defaultContainer 'maven'` so the whole block targets it. Tool containers need a long-running command — a `sleep` or a TTY — or they exit immediately and the step fails. When the build ends the pod is deleted, so nothing outside the workspace volume survives.

code

groovy · 28 lines
groovy
pipeline {
  agent {
    kubernetes {
      yaml '''
        apiVersion: v1
        kind: Pod
        spec:
          containers:
          - name: maven
            image: maven:3.9-eclipse-temurin-17
            command: ["sleep"]
            args: ["infinity"]
            resources:
              requests: {cpu: "500m", memory: "1Gi"}
              limits:   {memory: "2Gi"}
      '''
    }
  }
  stages {
    stage('Build') {
      steps {
        container('maven') {
          sh 'mvn -B verify'
        }
      }
    }
  }
}

go deeper

for a junior

Know that this model creates a fresh pod for each build from a template, that one container in the pod is the Jenkins agent itself, and that the pod is deleted when the build ends.

for a middle

Explain container selection: steps default to the agent container, container('name') { } or defaultContainer targets a tool container, tool containers need a long-running command, and the workspace volume is mounted into all of them.

for a senior

Diagnose it live: unschedulable pods from resource requests or taints, OOMKilled containers taking the workspace with them, image-pull secrets, and the resource requests you set per container rather than per pod.

for a principal

Weigh elasticity against cold start and cluster coupling. Decide caching strategy without persistent workspaces, whether idleMinutes reuse is acceptable, namespace and quota isolation between teams, and who owns the pod templates every pipeline now depends on.

## The execution model The Kubernetes plugin makes Jenkins agents disposable. Instead of a fleet of long-lived machines, the controller holds a **cloud** configuration pointing at a cluster, and each build that requests a matching agent causes a **pod** to be created, used once, and deleted. The sequence: 1. A build requests an agent — either by label matching a pod template defined in the cloud configuration, or by declaring the pod inline in the Jenkinsfile. 2. The plugin creates the pod in the configured namespace, with a generated name. 3. One container is the **inbound agent**, named `jnlp` by default. It runs the Jenkins agent JAR, dials the controller (WebSocket over HTTPS in modern setups, or the older TCP JNLP port), and registers as a node. 4. The controller schedules the build onto that node. Steps execute inside the pod. 5. When the build finishes, the node is disconnected and the pod deleted. A declarative pipeline expresses it like this: ```groovy pipeline { agent { kubernetes { defaultContainer 'maven' yaml ''' apiVersion: v1 kind: Pod spec: containers: - name: maven image: maven:3.9-eclipse-temurin-17 command: ['sleep'] args: ['infinity'] resources: requests: {cpu: '500m', memory: '1Gi'} ''' } } stages { stage('Build') { steps { sh 'mvn -B verify' } } } } ``` ## Which container runs the step This is the part candidates get wrong. A pod has several containers, but `sh` has to pick one. The rule: - **By default, steps run in the agent container** (`jnlp`). That image contains a JRE and the agent JAR — no Maven, no Node, no kubectl. A `sh 'mvn -v'` there fails with "command not found", which is the single most common first-day confusion with this plugin. - **`container('maven') { … }`** switches the enclosed steps into that named container. The step is executed there via the Kubernetes exec API. - **`defaultContainer 'maven'`** in the declarative `kubernetes` block makes that container the target for the whole block, so you stop writing `container(...)` around everything. Two constraints follow from how containers are used here. First, every tool container must **stay running** — a container whose image entrypoint exits immediately (most language images do) leaves nothing to exec into. Give it `command: ['sleep']` with `args: ['infinity']`, or equivalent, or enable a TTY. Second, all containers see the same **workspace volume**: the plugin mounts an `emptyDir` (the `workspace-volume`) at the agent's working directory, typically `/home/jenkins/agent`, into every container. That is what makes it legal for a `maven` container to compile and a later `container('kubectl')` block to read the resulting files. ## Failure modes worth knowing - **Pod never schedules.** Requests exceed what any node can offer, or a node selector, taint or quota blocks it. Symptom: the build sits waiting for an executor, sometimes for the plugin's provisioning timeout, then errors. Diagnose with the pod's events, not the Jenkins log alone. - **Container OOMKilled or pod evicted mid-build.** The build dies abruptly, typically with a channel-closed or agent-disconnected error, and because the workspace was an `emptyDir` inside the pod, it is gone. Nothing is recoverable except what you stashed or archived along the way. Right-size memory requests and limits on the tool containers, not just the agent container. - **Image pull failures** show up as an unschedulable-looking hang; a private registry needs an `imagePullSecret` in the pod spec. - **Every build pays a cold start.** Pod creation plus image pull plus agent connection is tens of seconds. `idleMinutes` on a pod template keeps a pod alive for reuse by subsequent builds, and `podRetention` controls whether pods are kept after failure for debugging — both trade away the clean-per-build property in exchange for latency or forensics. ## Why teams choose it The pull is that agent capacity becomes elastic and every build starts from a known image. The dirty-workspace class of bug disappears, because there is no workspace to inherit. Toolchains are declared in the Jenkinsfile rather than installed by hand on VMs, so "works on agent 3, fails on agent 7" stops happening. The costs are equally real: a cold start per build, a dependency on cluster capacity and quota, YAML that is now part of every pipeline's blast radius, and the fact that caches must be deliberately engineered — a persistent volume or a remote cache — because nothing survives the pod. ## Interview framing The question is a good discriminator because the naive answer ("it runs builds in Kubernetes") says nothing. What demonstrates real use is the container-selection rule, the sleeping-tool-container requirement, the shared workspace volume, and knowing that an evicted pod takes the workspace with it.

  • Why must a tool container in a Jenkins pod template be given something like sleep infinity?
    Steps are executed by exec'ing into an already-running container. Most language images have an entrypoint that runs and exits — `maven` prints help, `node` starts a REPL that ends with no TTY — so the container terminates before any step can use it. Overriding `command` with a long sleep, or enabling a TTY, keeps it alive for the life of the pod.
  • A build dies with an agent-disconnected error and the workspace is gone. What likely happened?
    The pod stopped existing. A container was OOMKilled, the pod was evicted under node pressure, or the node was drained. Because the workspace is an `emptyDir` inside that pod, it disappears with it — only what was already stashed or archived survives. Check the pod's events and the container's last state, then raise the memory request and limit on the container that died.
  • What do idleMinutes and podRetention change, and what do they cost?
    `idleMinutes` keeps a provisioned pod alive for a while so later builds reuse it, trading the clean-per-build guarantee for a faster start. `podRetention` (never, onFailure, always) decides whether finished pods are left in the cluster for inspection, which helps debugging but consumes quota and leaves workspaces around. Both weaken the disposability that made the model attractive, so scope them narrowly.

saying these in an interview costs you the question

  • Expecting sh steps to find tools in the default agent container
  • Omitting a long-running command from tool containers
  • Assuming the workspace survives after the pod is deleted
  • Sizing only the agent container's resources
  • Treating pod startup time as negligible per build

context