A CI runner is killed mid-suite. What does the results store hold afterwards if the run was pushed as a directory at the end of the job, and what does it hold if the run was being reported to a ReportPortal server as it went?
answer
- two delivery shapes, opposite failures
- all at once, or a conversation
- the launch nobody ever closed
- stopped on purpose, or reaped
- IN_PROGRESS until a finish call
basics
~20 sA directory pushed at the end leaves the store with nothing: the push never ran and the workspace died with the runner. A live-reported run leaves a launch stuck at IN_PROGRESS, because its finish call never arrived.
solid answer
~50 sThe two delivery shapes fail in opposite directions. A results directory pushed once at the end is all-or-nothing — the adaptor was writing files to local disk the whole time, but the upload step never executed, so the store holds nothing at all and the partially written directory is destroyed with the workspace. A live-reporting client is the reverse: the launch was opened on the server before the first test, items streamed in as they finished, and then the finish call never arrived. The server is left holding a launch stuck at `IN_PROGRESS` with items still open beneath it. Closing it explicitly records `STOPPED`; a server-side sweep that reaps launches which have run longer than a configured duration records `INTERRUPTED` instead. Those two are worth distinguishing: one means somebody ended it, the other means nobody did.
go deeper
Know that an end-of-job upload is all-or-nothing: if the runner dies first, nothing about that run ever reaches the store and the local files go with the workspace.
Explain why live reporting leaves the opposite trace — a launch opened on the server that never received its finish call, so it sits at IN_PROGRESS holding whatever arrived.
Show that you would resolve stuck launches deliberately and keep the STOPPED versus INTERRUPTED distinction, and that you would not let a truncated directory read as a small passing run.
Own the choice between the two delivery shapes across an estate, weighing total loss on a kill against the ongoing cost of reaping open runs and the operational habits each demands.
## Two delivery shapes, two opposite failure modes Almost every results store is fed one of two ways, and a killed runner exposes the difference immediately. **Push a directory at the end.** The adaptor writes result files to local disk as tests finish; a single step at the end of the job archives that directory and sends it. Delivery is one event. **Report live as the run goes.** A client opens a run on the server before the first test, reports each item as it completes, and closes the run at the end. Delivery is a conversation. ## The directory push: all or nothing When the runner dies mid-suite, the upload step is simply never reached. The store holds **nothing** for that run — not a partial run, not an empty one, not a marker. From the store's point of view the run does not exist and never did. The partially written directory does exist, briefly, on the runner. Every test that had finished has its file; the ones still running or not yet started do not. That directory is then reclaimed with the workspace, so it is not recoverable. The subtle hazard here is what would happen if that directory *were* pushed — by a retry, a salvage step, or a job that dies after the upload but before its final steps. A truncated results directory is **structurally indistinguishable from a complete one**. There is no field saying "this run was cut short"; the results simply describe fewer tests. The store records a smaller run that looks entirely healthy, and any comparison against previous runs quietly reads a shrunken suite as normal. ## Live reporting: the launch that never closes The streaming shape has the opposite problem: the server already knows about the run. In ReportPortal terms, a launch was created and its `StatusEnum` value is `IN_PROGRESS`, which is the state a launch occupies from the moment it is opened until a finish call arrives. When the runner is killed, that call never arrives. The launch stays `IN_PROGRESS`, holding whatever items were reported, and test items that were open when the process died stay open too. A stuck launch does not resolve itself by timing out at the client — the client is gone. It is closed from the server side, in one of two ways, and which one happened is recorded: - **`STOPPED`** — a person or an API call explicitly ended the launch. The launch is given an end time, and the items still in progress underneath it are interrupted. `STOPPED` therefore means *somebody decided this run was over*. - **`INTERRUPTED`** — a server-side job that looks for launches which have been running longer than a configured duration finished this one on its own. `INTERRUPTED` means *nobody closed it and the server reaped it*, and it is counted on the failed side of the run's statistics rather than the passing side. ## The two shapes side by side | | Directory pushed at the end | Reported live | |---|---|---| | Store state after a kill | nothing exists | a launch stuck at `IN_PROGRESS` | | Partial data reaches the store | no | yes, up to the moment of death | | Is the incompleteness visible? | no — there is no run to look at | yes — the state itself says so | | Who resolves it | nobody; the run is simply absent | an explicit stop (`STOPPED`) or a server sweep (`INTERRUPTED`) | | Cost of the mode | you lose everything | you accumulate open runs that need reaping | ## What to build, given both 1. **Make incompleteness representable.** If you push a directory, the store cannot tell a truncated run from a small one. Give it something to compare against — have the job declare what it expected to report, so a run arriving with far fewer results can be flagged rather than averaged in. 2. **Decide who closes stuck runs, and when.** With live reporting, the question is never *whether* stuck launches happen but who ends them. Leaving them to a server sweep is a legitimate choice; leaving them to nobody is not, because open runs distort every view built over the project. 3. **Preserve the distinction the states give you.** Collapsing `STOPPED` and `INTERRUPTED` into one "didn't finish" bucket throws away the difference between a run somebody cancelled and a run whose machine vanished, which are different incidents with different fixes. 4. **Do not treat a kill as a test failure.** A run cut short by a lost runner says nothing about the software under test. Whoever reads the store needs to see that distinction, or the next person will chase a product bug that was really a reclaimed spot instance.
- Why is it worth keeping `STOPPED` and `INTERRUPTED` as different states rather than collapsing both into "did not finish"?`STOPPED` records a deliberate end — a person or an API call closed the launch. `INTERRUPTED` records that nobody closed it and the server reaped it after it outran its allowed duration. Those are different incidents: one is a decision, the other is a lost runner or a wedged client. Merging them hides which of the two you have.
- With a directory pushed at the end, how would you make a truncated run visibly incomplete rather than merely small?Give the store something to check the run against — have the job declare what it expected to deliver, or push a small marker as the last thing in the same upload. Then a run whose contents fall short of what it announced can be flagged, instead of being read as a healthy suite that happens to have shrunk.
A results directory pushed at the end is like a shop that banks the day's takings at closing time: if the shop burns down at three in the afternoon, the bank has nothing at all. A live-reporting client is like a till that wires each sale as it rings up — the bank has the whole morning, and a till that is still sitting open until someone goes and shuts it.
saying these in an interview costs you the question
- Assumes a killed run leaves a partial upload in the store
- Thinks a stuck launch times out and closes itself
- Reads a truncated results directory as a small healthy run
- Treats an interrupted run as evidence the product failed
- Believes the runner's workspace can be recovered after the job dies