In ZAP, why does a `-cmd` run return a failure exit code where a `-daemon` run returns zero?
answer
- zero can mean not yet started
- who reads the status on the way out
- one bootstrap returns it, one does not
- the recording is gated, not just the reading
basics
~20 sOnly the one-shot bootstrap returns an exit status; the daemon bootstrap returns zero as soon as its background thread is alive. The automation add-on also refuses to record a status unless the process type is cmdline.
solid answer
~40 sTwo mechanisms stack. First, `CommandLineBootstrap.start()` finishes with `rc = control.getExitStatus()` and `ZAP.main` turns a non-zero value into a process exit; `DaemonBootstrap.start()` returns zero the moment the `ZAP-daemon` thread is alive, because the process has not finished anything yet. Second, the automation add-on's own `setExitStatus` helper is wrapped in a check that the process type is `cmdline`, so on a daemon the plan's failure is never even written to the control singleton. A daemon can still exit non-zero — the out-of-space path and the `insights` add-on's `exitAutoOnHigh` path both set a status and exit when no view is attached — but never because the plan you ran said so.
go deeper
Learn the shape first: a run that finishes can report how it went, and a service that has just started up cannot. Reach for the one-shot lifetime whenever a job needs a pass or fail.
Explain both halves of the mechanism — the bootstrap that returns the control singleton's status versus the one that returns zero immediately, and the automation helper that refuses to record a status unless the process type is the one-shot one.
This is the question behind "why is our security job always green". Show that you would check what actually set the exit status, notice that a service lifetime returns before any work happened, and move the verdict to something you read deliberately.
Decide organisation-wide where a security gate's verdict comes from, and make it something a job cannot accidentally satisfy. An exit code that means "the process started" is worse than no gate, because it reports success on every run.
## Where a process exit code comes from here The control singleton holds a single integer, set through `setExitStatus(status, logMessage)` and read through `getExitStatus()`. Nothing about that field is mode-specific. What *is* mode-specific is whether anybody reads it on the way out, and whether the thing you ran was allowed to set it in the first place. Both halves break on the service path, for different reasons. ## The one-shot path reads it `CommandLineBootstrap.start()` runs the registered work, shuts the program down in a `finally`, and then does the equivalent of "if I have no error of my own, adopt the control singleton's status". `ZAP.main` takes that return value and calls `System.exit` on anything non-zero. So on this path the exit code is a genuine end-of-run verdict, produced after the work has actually finished. The field's own documentation says as much: it notes that setting an exit status works however the program is run, but that it makes more sense in command-line mode. That sentence is the whole leaf in miniature. ## The service path returns before there is anything to report `DaemonBootstrap.start()` builds a thread, starts it, and returns zero. At that instant the plan has not run, the scan has not started, and no verdict exists. The process then stays alive because that thread loops instead of returning. It ends only when something calls `System.exit` explicitly — the control API's shutdown action does, and it passes `getExitStatus()` through, so a daemon *can* deliver a status if you shut it down through the API after something set one. ## And off that path the automation framework will not even record one This is the part that surprises people. The automation add-on has a private helper that every one of its status-setting paths goes through, and that helper begins by asking whether the process type is `cmdline`. If it is not, the helper returns without touching the control singleton and without logging the error line. So under `-daemon`, a plan with errors does not merely fail to *report* a status — it never *records* one. That includes the explicit case. The plan job whose entire purpose is to choose an exit value does not set one directly either: it stores an override on the add-on, and the same gated helper is what would later apply it. The gate is upstream of everything. | what fails | sets an exit status? | gated on | |---|---|---| | an automation plan's errors or warnings | only under the one-shot lifetime | the process type being `cmdline` | | the database or disk filling up | yes, then exits | no view being attached | | a high-level insight with the auto-exit option on | yes, then exits | no view being attached | Notice that the second and third rows are gated on the *absence of a view*, not on the one-shot lifetime. That is the tree-wide asymmetry running the other way: those two are self-termination paths that fire in headless runs and not on the desktop. ## So what does a daemon's exit code mean? It means the process stopped, and usually it means somebody stopped it. It does not mean the scan passed. A pipeline that starts the program with `-daemon`, waits, and then reads the shell's status variable is reading "the service came up", which is true of a run that scanned nothing at all. ## What to do instead - **Run the work under the one-shot lifetime** when the pipeline's decision is "did this pass". That is the only lifetime whose exit code is produced after the work finished and is allowed to carry the automation framework's verdict. - **If you need the service lifetime** — because several steps drive the control API in an order you control — then decide explicitly where the verdict comes from, read it yourself, and shut the process down through its API rather than killing it, so that whatever status was set is actually returned. - **Do not treat a green job as evidence that anything was scanned.** That is true of both lifetimes and doubly true of the service one, where zero is the value you get before the work has even started. - **Diagnose on the service path from the log, not the console.** The reporting helpers that command-line listeners use only write to standard output and standard error when the process type is `cmdline`; everywhere else the same message goes to the log and nowhere visible. ## The trap to name in an interview Say the two halves separately. "The daemon's start returns zero before the work runs" is the obvious half. "And the automation framework's status-setting is gated on the process type, so nothing would have been recorded even if something read it" is the half that shows you have looked.
- Can a ZAP daemon ever exit with a non-zero code?Yes, through paths that are gated on there being no view rather than on the lifetime: the out-of-space handler and the `insights` add-on's auto-exit-on-high option both set a status and then exit. The control API's shutdown action also passes the current status through. What cannot reach that field on a daemon is the automation framework's own verdict.
- Why does the error message from a command-line listener not appear on the console under `-daemon`?Because the reporting helpers those listeners use switch on the process type: they print to standard output or standard error only when it is `cmdline`, and otherwise write to the log alone. On the service path the message exists, but you have to go and read the log for it.
A batch job is asked a question and answers when it is done. A server is asked to open for business, and "I opened" is the only thing it can tell you at start-up — whatever happens to a customer later is not in that answer.
saying these in an interview costs you the question
- A daemon exit code of zero proves the scan passed
- Exit codes are produced by the wrapper, so the lifetime is irrelevant
- A daemon can never exit with a non-zero code
- The automation framework records a verdict the same way in every mode
- Killing the daemon process returns whatever status was set