skip to content

Vibe Coding

Prompt-driven, agent-in-the-loop development: state the intent, let the agent write, review, iterate. Interviewers probe this to see whether you drive the loop deliberately or simply accept whatever comes back.

part ofAI-assisted developmentoverview, primer and where to startread it →
on this pageshow

questions

29

Why does one wrong assumption early in an agent session make the turns that follow more confident rather than less?

level: juniorimportance: must knowfreq 62%

answer

  1. The agent reads its own output back
  2. Nothing marks which lines were checked
  3. Qualifiers disappear when a claim is restated
  4. Its edits remove the evidence it would find
  5. Confidence tracks repetition, not verification

basics

~20 s

An agent's own earlier conclusions sit in the session record, and later turns read them as given rather than as guesses. Each restatement drops the hedge the first one carried, so certainty grows while evidence does not.

solid answer

~40 s

A session is a record, and the agent wrote most of it. A turn that says `I searched and did not find other readers of this field` becomes, several turns later, `this field has no other readers` — the hedge rarely survives restatement, and nothing in the record marks which lines were checked and which were guessed. The session has also been editing the code the whole time, so a later search for the old name comes back empty because this run removed every occurrence it could see. Agreement accumulates inside the session while verification does not, which is what rising confidence is made of. The premises worth checking yourself are the ones the run decided rather than the ones you supplied.

code

text · 13 lines
text
turn 3   "I searched the files I could read and did not find
          other readers of the status field."

turn 6   "Since status has no other readers, I renamed it in
          the model and updated the callers."

turn 11  "Searched for the old name: no occurrences left.
          The rename is complete."

turn 15  "No migration is needed here — the field was internal."

# the claim never got stronger; the wording did
# turn 11's search was empty because turn 6 emptied it

go deeper

for a junior

Recall the shape: the agent reads its own earlier statements back as facts, so an early guess quietly becomes a premise. Name one thing you would verify yourself rather than take from the session.

for a middle

Explain the mechanism: no provenance in the record, hedges lost in restatement, and a run whose own edits erase the traces a later search would find. Say why that makes confidence rise.

for a senior

Show the judgment call — correct in place while the premise is cheap, or abandon and restart with the fact stated up front once it is in the code. Say what each choice costs you.

for a principal

Own the framing: a long run's certainty is an artefact of its record, not a measurement of the codebase, so the premises a change rests on belong in the open where somebody can check them.

## A session is a record, and the agent wrote most of it When an agent works across many turns, what it reasons from next is not your codebase directly. It is the running record of the session: your request, the files it has read, the edits it has made, and **everything it has previously said**. That last category is the trouble. A conclusion the run reached on turn three is, by turn twenty, just a line in the record, sitting beside the lines you wrote and the file contents it read, and **nothing in it is stamped with where it came from.** So `I searched the files I could read and did not find other readers of this field` and `this field has no other readers` end up doing much the same work on a later turn, even though the first is a report about a search and the second is a claim about the world. **Context poisoning** is the usual name for the wrong item entering the session's working record; **context drift** is the usual name for what the session then does with it, as later turns inherit it and build on it. It is unlike an ordinary mistake in a single edit, because the wrong item is not in the output where you would review it — it is in the material that produces all the later output. ## The hedge decays with each restatement Long sessions summarise themselves constantly: a turn refers back to what was decided, a plan is restated, a change is described before it is made. Qualifiers tend to be the first words a summary drops. The claim survives; the uncertainty attached to it does not. | what the early turn actually established | what later turns treat as established | |---|---| | a search over the files it could read found no other readers | the field has no other readers | | the tests that already existed still passed after the rename | nothing outside the code depends on the old name | | no caller inside this repository used the old name | no stored data, configuration or outside consumer uses it | Each right-hand side is a larger claim than its left, and the step between them was taken by restatement rather than evidence. ## The run edits the evidence it would check against Suppose the run renames the field on turn five. On turn eleven, being careful, it searches for the old name to confirm nothing is left behind. The search comes back empty — largely **because the run itself removed every occurrence it could see.** What it still cannot see, such as a column in stored data or a key in a configuration file it never opened, is exactly what it could not see on turn three. The confirmation therefore adds close to no information, and it reads like confirmation. That is the shape of the whole failure: agreement accumulates inside the session while verification does not. ## Why a late correction lands badly Telling the run on turn twenty that the premise was wrong helps less than it should, for reasons worth separating: - The earlier line is still in the record. Your correction sits after it; both are read. - The premise is no longer only in the record. It is in the code the session has written, and the code agrees with it. - A run that has produced a lot of work on a premise has a lot of material agreeing with it, and agreement is what it reasons from. - The correction has to be applied backwards through work built on the wrong footing, which is a bigger request than the one that caused the problem. So a correction lands better the earlier it is made, and once the premise has produced many edits, a fresh session with the corrected fact stated at the top is often cheaper than arguing inside the poisoned one. That trade costs the session's accumulated useful work, which is why it is worth deciding deliberately rather than drifting into. ## What to check, and when 1. **Separate what the run was told from what the run decided.** You supplied some facts; it inferred others. Both kinds can be wrong, and the decided ones are the ones nothing outside the session ever checked. 2. **Check the decided ones yourself, outside the session,** against the thing they are claims about — the stored data, the configuration, the consumers you know of. 3. **Do it early, while the premise is cheap.** A premise checked on turn four costs one question; the same premise checked on turn thirty costs whatever was built on it. 4. **Read a confident restatement as a reason to check, not as evidence.** A confident restatement tells you the claim has been repeated; it does not tell you it has been supported. ## What this is not This is not the context window filling up — how much a session can hold is a separate subject, and a session can be poisoned while well inside its capacity. It is also not an argument that assumptions are avoidable: a run that assumed nothing would ask about everything and get nothing done. The practical question is narrower: which assumptions is this change resting on, and which of those have you confirmed with something other than the session's own say-so.

  • How would you tell a poisoned premise from the agent simply being wrong on one turn?
    A one-turn mistake is local: the edit is wrong and the turns around it are unaffected. A poisoned premise shows up as several unrelated-looking edits that are all consistent with one thing the run believes, including edits that look like tidy-up. If correcting one edit implies correcting the others, you are looking at a premise.
  • You spot the wrong premise on turn four rather than turn thirty. Does that change what you do?
    Yes, and it is the whole reason for looking early. On turn four almost nothing rests on it, so stating the correct fact and continuing is usually enough. By turn thirty the premise is in the code as well as the record, and a fresh session with the fact stated up front is often the cheaper route.
  • Can you ask the agent which of its statements were verified?
    You can ask, and you will get an answer, but the answer is produced from the same record that contains the claim. It is a restatement, not a provenance trail. Treat it as a place to start looking rather than as the check itself, and confirm the premise against the code, the data or the configuration it is a claim about.

A rumour grows more certain with each retelling, not because anyone learned anything new, but because each teller drops the hedge the last one used. A long session does the same thing to its own early guesses.

saying these in an interview costs you the question

  • Telling the agent it was wrong removes the wrong assumption from the session
  • How confidently the agent states something tracks how well supported it is
  • A bad assumption will surface as a failing build before it spreads
  • Asking the agent to double-check its own earlier conclusion settles it
  • A bigger context window would prevent this
  • Drift is just the model being unreliable, so nothing you do about the session matters
open as a page

Why is “add a retry policy to the webhook sender” under-specified even though it reads complete?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The request names an outcome and no behaviour. Retrying a payload the receiver rejected, sending the same one twice, and giving up silently all satisfy it — so the agent picks one, and you usually find out from the diff.

open as a page

A check failed after an agent's change — why hand back the run's own output rather than your summary of it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The run's output carries the failing case, the expected and actual values, and the line that produced them. A summary replaces those facts with your diagnosis, and if the diagnosis is wrong the next turn inherits it.

open as a page

An agent changed nine files unattended and no author can explain them — how does that change your review?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A generated change carries no author intent, so an unexplained line is not evidence of a reason you lack — it is an open question. Reconstruct intent from the request, verify the risky parts, and answer for what you merge.

open as a page

Why does building a whole feature in one agent run fail differently from building it as five?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A single run has no interior — nothing inside it meets a check you already trusted. An early wrong decision surfaces only after later work is built on it; five units catch it while it is still one unit wide.

open as a page

An agent's change adds a dependency that does not exist — what makes this a security problem?

level: middleimportance: must knowfreq 54%

basics

~20 s

The broken build is not the risk; the reflex is. Fetching a name to make the build pass hands the decision to whoever claimed that name, and a plausible invented name is worth claiming in advance.

open as a page

Your agent retried deliveries the receiver had permanently refused — what do you add to the request?

level: middleimportance: must knowfreq 56%

basics

~20 s

Add the fact that tells the two cases apart — a refusal the receiver would repeat is not retried, a silence is — and name the wrong behaviour as an explicit non-goal. Emphasis carries no information; a fact does.

open as a page

An agent's loop ended with every check green. Why might the fault it was chasing still be there?

level: middleimportance: must knowfreq 62%

basics

~20 s

A loop optimises the verdict it was given, and the verdict is editable. Green can be reached by weakening an assertion, skipping a case or widening a catch — all cheaper than fixing the cause.

open as a page

How do you check that an agent's change did what you asked, and not just something reasonable?

level: middleimportance: must knowfreq 58%

basics

~20 s

Write the request out as observable outcomes before reading the change, then walk it both ways: every outcome present, and every behaviour change traceable to an outcome. The backwards walk is the half reviewers skip.

open as a page

Which properties of the work, not of the tool, decide which AI coding tool shape you should use?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The task decides it: whether a fast, trustworthy check exists; whether the change is wide and mechanical or narrow and subtle; whether you know the code well enough to review at speed; and whether it must be explainable later.

open as a page

How do an inline completer, an in-editor agent, a terminal agent and a hosted agent differ in what they can see and reach?

level: middleimportance: should knowfreq 52%

basics

~20 s

Four shapes, two axes: what each can read, and what each can act on. A completer writes at one point; an in-editor agent edits an open project; a terminal agent also runs commands; a hosted agent changes only its own copy.

open as a page

You approved every turn of an agent session, so why is the finished change one you would not approve?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each turn is judged against the state the turn before it left, so the baseline moves with the change. A turn that repairs the previous turn's side effect is locally right and carries the session further from what you asked.

open as a page

An agent asked to rename one field is editing its fortieth file — what should have told you sooner?

level: middleimportance: should knowfreq 46%

basics

~20 s

File count is the weakest signal: a rename legitimately reaches every caller. The warning is new territory the request never named — a stored-data change, an externally visible shape, an edited test expectation — and it appears well before the count does.

open as a page

Before an agent starts a multi-turn run on your webhook sender, what must the request state beyond the goal?

level: middleimportance: should knowfreq 50%

basics

~20 s

Three things the goal itself does not carry: the fence, naming artefacts the run must not change; what done means, in terms you could check against the diff; and the decisions it must bring back rather than settle alone.

open as a page

Before an agent starts iterating on a change, what makes a check usable as the loop's steering signal?

level: middleimportance: should knowfreq 54%

basics

~20 s

A usable verdict runs without you, returns the same answer on the same input, is cheap enough to run every turn, and names what failed rather than that something did — and it has to be failing now, for the right reason.

open as a page

In a generated change, why is a plausible call to a real API harder to catch than an invented one?

level: middleimportance: should knowfreq 52%

basics

~20 s

A name that does not exist fails the moment anything resolves it. A real name used on a wrong assumption resolves and fits the code around it, since one run wrote both sides; the disagreement is with the real contract.

open as a page

Your five agent units can only be tested once the fourth lands — how do you re-cut the plan?

level: middleimportance: should knowfreq 48%

basics

~10 s

Re-order so evidence arrives first: open with a behaviour-preserving unit the existing suite already checks, then the smallest unit that settles the riskiest assumption. Merge any unit that cannot fail on its own.

open as a page

What tells you a unit of work is too big for one agent run before you start it?

level: middleimportance: should knowfreq 56%

basics

~10 s

Three tells, all visible before you type: the unit has no end state where the project still builds, no check that existed before the unit did, and no area you can name.

open as a page

In a long agent coding session, what fills the context window, and why does the run degrade rather than stop?

level: middleimportance: should knowfreq 52%

basics

~20 s

A long session is filled mostly by its own by-products: files read entire, full command output, failed attempts, its own restatements. The room is finite, so older material goes and the work carries on with less.

open as a page

Beyond the change itself, what does a long autonomous agent run cost you that a small assisted edit does not?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A long autonomous run concentrates a large change into one review decision, widens what a mistake can touch before anyone notices, leaves less of the session reconstructable, and produces a result a re-run will not reproduce.

open as a page

What has to be true before an agent session starts for abandoning it later to be cheap?

level: seniorimportance: should knowfreq 50%

basics

~20 s

A point you verified — built, tested, recorded in version control — and a run whose reach stays inside that tree. Recovery cost is set by what the session could touch outside it, not by how many files changed.

open as a page

How much surrounding code do you hand an agent at the start of a session, and what does too much cost?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Hand over what changes a decision the run would otherwise get wrong — the contract, the record shape, one example of the code you want. Surplus material gets imitated, and it makes a wrong result impossible to attribute.

open as a page

Four turns in, each fix breaks a neighbouring case — how do you tell a stalled loop from a converging one?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Watch the failing set, not the last message. A converging loop shrinks it and keeps it shrunk; a stalled one moves it sideways — same size, different members. Fix the stopping condition before turn one, while the sunk cost is small.

open as a page

An agent's change adds an interface with a single implementation — how do you decide whether to keep it?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Name the second implementation, caller or value the structure exists for. If it exists today or is genuinely planned, keep it; if it is only imaginable, the code is paying now for a prediction, and it comes out.

open as a page

After a two-week team trial of an agentic coding tool, what can you legitimately conclude from it?

level: principalimportance: should knowfreq 38%

basics

~20 s

A short trial mostly tells you about yourselves: how well your checks hold, how much of a run you can reconstruct, where your codebase resists being read. It generalises poorly to untried areas, other teams, and later versions.

open as a page

Which engineering goals can you state well enough to hand to an agent run, and which can you not?

level: principalimportance: should knowfreq 33%

basics

~20 s

A goal is statable when its acceptance exists apart from the implementation. Where acceptance is what the work discovers, or two teams disagree what done means, no wording fixes it — settle it first, or run to learn and discard.

open as a page

In a loop steered by a failing check, what should the agent be barred from editing, and what does barring it cost?

level: principalimportance: should knowfreq 36%

basics

~20 s

Whatever decides the verdict should not move silently: assertions, expected values, skip lists, rule configuration, the threshold that fails the build. The cost is real — checks are sometimes wrong, and a locked loop burns turns failing against a stale expectation.

open as a page

When is re-running an agent with a corrected request cheaper than repairing the change it produced?

level: seniorimportance: nice to knowfreq 34%

basics

~20 s

Re-run when the premise is wrong — a misread request or an unstated constraint — because that mistake is spread through every file. Repair when the premise holds and the defects are few. A re-run discards the review already done.

open as a page

When a feature is split across several agent runs, what stops the fourth reinventing what the first built?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Only what is legible in the code carries across a boundary. A later run reads the repository, not the earlier run's reasoning — so a decision survives as a helper with no route around it, or as a failing test.

open as a page