skip to content

Should a pipeline run ZAP as a fresh container per job or as one long-lived daemon?

level: principalimportance: should knowfreq 35%

answer

  1. verdict, state, surface, cost
  2. what survives between two targets
  3. which lifetime can report an outcome
  4. pin the tag, own the listener

basics

~20 s

Default to a fresh container per job running the one-shot lifetime: it is the shape whose exit code is a verdict and whose state cannot leak between targets. Share a daemon only when several steps need the control API.

solid answer

~50 s

The two topologies differ on axes that are mechanical rather than a matter of taste. **Verdict**: only the one-shot bootstrap hands the control singleton's exit status back as the process result, and the automation framework refuses to record a status at all unless the process type is the one-shot one. **State**: a container that dies with its job cannot carry a session, a site tree or a configuration edit into the next target; a standing daemon carries all three. **Surface**: a long-lived listener exists between runs, which somebody has to own. The shared daemon's real advantage is amortised start-up and the ability to drive one instance from several steps in an order you control. If you need that, the honest compromise is still a per-job container that runs the daemon lifetime and is shut down through its own API so a status is returned.

go deeper

for a junior

You will usually be handed the topology rather than choosing it. Know which one your pipeline uses and what it means for how your job finds out whether the run passed.

for a middle

Be able to argue both sides on mechanism rather than preference: which lifetime can return a status, what state a long-lived instance carries between targets, and what start-up actually costs.

for a senior

Show you would migrate an existing pipeline safely — verdict first, then isolation, then the listener's ownership — rather than rewriting the topology and hoping the gate still means something.

for a principal

This is a standard you own. Pick a default, name the condition for the exception, decide where verdicts come from and who owns the listener, and make the image tag a reviewed dependency rather than a moving name nobody refreshes.

## The question behind the question Both topologies work, and teams adopt one by accident far more often than by decision — usually whichever the first person to wire it up happened to copy. It is worth deciding, because the two differ on things a pipeline actually depends on. ## The two shapes | | container per job, one-shot lifetime | one long-lived daemon | |---|---|---| | how the job learns the outcome | the process exit code, produced after the work finished | something you read yourself; the start-up return value is meaningless | | can the automation framework record a status | yes — its status helper is gated on exactly this lifetime | no, the helper returns early | | state between targets | none survives the container | session, history, site tree and configuration edits all persist | | start-up cost | paid per job: image pull, add-on load, model init | paid once | | concurrency | naturally one run per container | several steps share one instance, and can interfere | | the listener | exists only while the job runs | exists continuously, and has an owner | | what a tag change does | changes what that job can do | changes what everything can do | ## What the per-job container buys Isolation is the headline, and it is stronger than it looks. The tool accumulates a session as it works — recorded traffic, a site tree, alerts, whatever a plan set along the way. In a standing instance, the second target inherits all of it, and "we saw an alert on service B that was actually raised against service A" is a real and very hard-to-diagnose failure. A container that dies with its job cannot produce it. The second thing it buys is a verdict that is an actual verdict, for the mechanical reason above rather than by convention. That matters because the alternative — reading a status out of the run's own output — is a thing someone has to build and keep working. ## What the standing daemon buys Amortised start-up is real: loading the program and its add-ons is not free, and a pipeline with many small jobs pays it many times. Beyond that, the genuine reason to want a service is **control**: several steps that drive the run through the API in an order you decide, with your own polling and your own decisions between them. That is a shape a one-shot invocation cannot express, and if you need it you need it. What you take on with it is: no meaningful exit code, shared state, and a listener that outlives any single job. ## What to standardise, and in what order 1. **Pick a default topology and write it down.** Teams should have to justify the other one, not invent one from scratch each time. 2. **Decide where the verdict comes from** before you decide anything else, and make it something a job cannot satisfy by accident. "The process exited zero" is not a verdict on the service topology, and a gate that reports success on every run is worse than no gate. 3. **Pin the image tag** in the pipeline definition, and own the refresh. A moving tag changes what every job can do with no change of yours. 4. **Decide who may reach the listener and with what credential.** A background instance exposes a control surface for as long as it is up; the control-API surface is where that question is answered, and it should be answered on purpose rather than inherited from an example command. 5. **Decide how a run is stopped.** On the service topology, shutting down through the API is what lets any status that was set be returned at all; killing the process discards it. 6. **Decide the isolation boundary between targets** explicitly if you keep a standing instance — a fresh session per target, or an instance per target, or an accepted and documented risk. ## The honest middle The compromise most teams land on is a **container per job that runs the service lifetime inside it**: the steps that need the API get the API, the container still dies with the job so nothing leaks, and the run is shut down through the API at the end so a status can come back. It keeps the isolation and the ownership story of the per-job shape while giving a multi-step job the control surface it wanted. The cost is the start-up you were trying to amortise, which is usually the cheaper thing to give up. ## How to answer this in an interview Do not pick a side and defend it. Name the axes — verdict, state, surface, cost — say which mechanism forces each one, state your default and the condition under which you would switch, and say what you would standardise so that the choice is not remade badly in every repository. The mechanical detail that shows you have used the tool rather than read about it is that the exit status is not merely inconvenient on the service topology: the automation framework does not record one there at all.

  • A team already runs one shared daemon and does not want to change. What do you insist on?
    Three things: a verdict the pipeline reads deliberately rather than inferring from an exit code, an explicit isolation step between targets so one run's session cannot colour the next, and a named owner for the listener including who may reach it. If those three are in place the topology is defensible.
  • Where does the container per job topology actually hurt?
    On start-up cost in pipelines with many small jobs, and on any workflow that genuinely needs several steps to drive one instance in an order the caller controls. The usual fix is one container per job that runs the service lifetime inside it, rather than one instance shared across jobs.
  • What makes a moving image tag a governance problem rather than a convenience?
    Because it silently changes what every job can do. A tag that moves can bring a different base, a different add-on set and a different architecture story, none of which appears in your diff. Pin it, and make refreshing it a change someone reviews.

saying these in an interview costs you the question

  • A shared daemon is fine, the exit code still tells us if it passed
  • State does not accumulate because each job sends its own target
  • Whichever is faster is the right answer
  • Pin nothing, always pull the newest image so it stays current
  • The listener is internal, so nobody needs to own who reaches it