On a rented browser fleet, why must a failing run's evidence be arranged before the run rather than after it?
answer
- what is left afterwards?
- no shell, no host, no rotation
- the run ended and so did the machine
- sent, recorded, kept - and nothing else
- evidence before the failure, not after
basics
~20 sA rented session leaves no machine to log into: it ran on a host you hold no account on, and it ended with the run. Only what your side sent, what your side recorded, and what the provider kept survive.
solid answer
~50 sA failure on a fleet you operate leaves an object behind: a host with a name, a container that may still be up, a log on disk, a machine you can hold out of rotation and re-run against. A rented session leaves none of that. It ran on a host you have no account on, it ended when your run ended, and there is no shell to open afterwards. What survives is exactly what your side sent, what your side recorded, and whatever the far side kept unasked — and you cannot add to any of it later. So the discipline inverts: on a rented fleet evidence is a precondition of the run rather than a response to a failure, and the identifiers that make a run nameable have to be captured while it is still alive.
go deeper
Understand that a remote session is gone once the run ends. If you did not capture something while the test was running, you cannot go back for it, so screenshots and logs have to be wired into the failing path in advance.
Be able to list what actually survives a remote run: what you sent, what you recorded, what the far side kept. Being able to recite that list is what stops you promising an investigation you cannot perform.
Expect to design the evidence path, not just use it. Say where identifiers are captured, what teardown records, and how you would raise a specific run with a provider rather than a vague complaint about last night.
You will be asked what a hosted fleet costs in diagnosability across many teams. Set the standard once — what every suite records at session creation and at failure — because unrecorded runs are unrecoverable and no central tooling can fix them later.
## Investigation on a machine you administer When a suite fails on a fleet you run, the failure leaves an object behind. There is a host with a name. There is a container that may still be up. There is a service log on disk, a browser profile, perhaps a core file. You can log in. You can hold that machine out of rotation and re-run the case against exactly it, with a debugger attached if you want. Every one of those moves rests on a single fact: **the machine is an asset you administer.** Reachability is not something you thought about, because it was never in question. ## Investigation on a machine you rent A rented fleet removes that asset, and it removes it completely rather than partially. The session ran on a host you have no account on. It ended when your run ended. There is no shell to open, no log directory to read, no container to keep alive "just for a bit", and no way to go back afterwards and ask for something nobody captured while it was happening. What survives a rented run is a short list, and it is worth being able to recite: - **What your side sent.** The request that created the session, the commands that followed, and their timings. - **What your side recorded.** Assertion output, whatever your harness captured at the moment of failure, and the identifiers that name this particular run. - **What the far side kept.** A provider commonly records something about a session without being asked, and that record can usually be retrieved afterwards. What a given provider records, and how long it keeps it, is a property of that provider: establish it, do not assume it. - **Nothing else.** Anything outside those categories did not survive the run, and no amount of asking afterwards creates it. ## Why that inverts the discipline On your own fleet, evidence gathering is a **response** to a failure. You see red, you go and look. The looking happens after, on a thing that is still there. On a rented fleet, evidence gathering is a **precondition** of the run. The failure and the disappearance of the machine are the same event, so anything you did not arrange to keep is simply gone. This is not a difference of emphasis; it is a different order of operations, and it is the commonest way a team gets less out of a hosted fleet than it paid for. ## What "arranging it in advance" concretely means 1. **Capture the identifiers that make a run nameable.** A conversation with a provider about "last night's failure" goes nowhere; a conversation about a specific run has somewhere to start. Those identifiers exist only while the session is alive unless your harness writes them down. 2. **Record what you asked for and what you were given.** The requested target and the environment the created session reported back are the only description of that machine you yourself will hold. 3. **Make the harness capture at the moment of failure, not after it.** Whatever your framework collects on a failing case is collected while the session still exists. A collection step that runs after teardown collects nothing. 4. **Make teardown record the outcome rather than exiting silently.** Teardown is the last code that runs with the session still open; it is the last chance to note anything. 5. **Know, before you need it, what the provider keeps and for how long** — by asking and checking once, rather than assuming a retention window exists. ## The hedge that keeps this honest "You get nothing back from a rented fleet" is as wrong as "you can just go and look". A hosted provider typically retains a record of a session it ran, and retrieving that record is a real and useful capability — it is one of the things renting adds rather than takes away. Two things about it are nonetheless true at once, and a strong answer holds both: - It is **the provider's record, on the provider's terms.** You did not choose what it contains. - It is **closed to additions after the fact.** Whatever it holds is what it holds; nothing you realise the next morning can be added to it retroactively. So the pre-arranged half of the evidence is the half you actually govern, and that is why it deserves the design attention. ## Working it through on a real suite Take a regional railway's seat-reservation site. A booking case fails at the seat-map step, at night, on a rented fleet. Ask yourself, before the run rather than after it, what you would want in your hand at nine the next morning: - The application build under test and the fixture data the case used — yours, if you recorded them. - The target requested and the environment given — yours, if the harness wrote them down. - Where in the case it failed and what the page looked like then — yours, if capture is wired into the failing path. - The run's identifiers, so you can ask the provider about that run — yours only if captured while it lived. - Anything about the host itself — not yours, beyond whatever the provider chose to keep. The first four are design decisions you make once. The last is the one you accept. That asymmetry, and not any particular tool, is what changes when the machines stop being yours.
- Is it fair to say you get no diagnostic information back from a hosted fleet?No, and that overstates it badly. A provider commonly records something about a session without being asked, and retrieving that is genuinely useful — it is one of the things renting adds. But it is their record on their terms, and it is closed to additions afterwards. What a given provider keeps, and for how long, is something to establish rather than assume.
- Why is a re-run a poor substitute for inspecting the original machine?Because a re-run is a different session on a different host. It cannot tell you what the failed run's machine looked like; at best it tells you whether the condition reproduces. On shared infrastructure a green re-run is weak evidence, since the host, its load and its neighbours have all changed underneath you.
- Which single artefact most often turns out to be missing the morning after a remote failure?The identifiers that make the run findable — the ones that let a conversation be about a specific session rather than about last night in general. They exist only while the session is alive, so a harness that does not write them down at creation time loses them permanently, along with any chance of asking the provider a precise question.
Investigating a failure on your own fleet is like inspecting a car in your own garage; investigating one on a rented fleet is like tracing a parcel after handing it to a carrier — the only account of the journey is whatever somebody arranged to record before it left.
saying these in an interview costs you the question
- Expects to open a shell on the host after a remote failure
- Assumes the provider keeps everything about every session forever
- Plans to start capturing evidence once a failure has appeared
- Treats a green re-run as equivalent to inspecting the original machine
- Cannot name which identifiers make a past run findable