skip to content

AI-assisted development

14 roadmaps59 questionsupdated

Working with AI coding tools as part of a normal engineering workflow — how you prompt them, scope work for them, and review what they produce. Interviewers ask because using these tools well, and knowing where they fail, is quickly becoming a baseline expectation.

on this pageshow

guide

overview

~1 min

AI-assisted development questions test whether you use coding assistants and agents as an engineer or as a passenger. Interviewers rarely care which tool you prefer. They want to hear how you decide what to hand over, how you phrase a request so the result fits your codebase, how you check what comes back, and where you stop trusting it. Claiming heavy use with nothing said about verification raises more concern than light use. The subject has two sections. [Vibe coding](/topics/found-vibe-coding) is the agent-in-the-loop mode: stating intent, cutting work into units an agent can finish and you can check, steering iteration with a real verdict, and recognising how long sessions go wrong. [AI assisted coding](/topics/found-ai-assisted-coding) is the assistant inside the ordinary edit-test-review loop: inline completion, prompting for a single function, refactoring, test generation, review discipline, and what the tools change for a team's productivity, skills and obligations. Learn first how a request becomes a result: what the tool can see, and what it fills in when you leave something out. Then learn review, because most of the other practices exist to make review possible. Questions run from a junior asked why a generated test proves little to a principal asked what a team policy or a short tool trial can actually settle. The material is judgement rather than tool trivia, so it carries across tools.

primer

A few ideas carry the whole subject. Once they are solid, most questions below read as consequences rather than rules to memorise. - **The tool works from a slice and fills the rest with the typical case.** Anything outside what it was given, such as your framework version, an existing utility you expected it to reuse, or a business rule nobody wrote down, gets replaced by whatever is most common. When a result is off target, the first question is what the tool could not see. - **A request is a specification, and its gaps get decided for you.** A goal names an outcome; the tool still has to choose a behaviour. What turns a wish into a task is stating the boundary the work must not cross, a definition of done you could verify from the diff, and which decisions have to come back to you. Repeating yourself louder adds nothing; a fact that separates the right behaviour from the wrong one does. - **Scope is set by the checks you already trust.** A unit of work is the right size when it ends in a state that builds and meets a check that existed before it. A larger unit lets an early mistake sit under everything built afterwards, so plans are ordered to produce evidence early. - **A loop chases whatever verdict steers it.** An agent iterating against a failing check is steering toward green, and the cheapest route there may run through the check itself. A loop is as honest as the list of things it may not edit and as useful as its failure messages are precise. - **Plausible is the expensive failure.** A made-up name fails the first time the build or runtime looks it up. A real call built on a mistaken belief about its contract, structure nobody asked for, or a test built from the code it tests can all pass the machine checks. Human review attention belongs where no deterministic check reaches. - **Accountability stays with people.** Whoever submits a change answers for it, and whoever approves it answers for that approval, however the code was produced. In practice that means submitting only what you could defend line by line. - **Cheap drafts move the bottleneck.** Producing code got faster; reading it, deciding what is correct and maintaining it did not speed up to match. Most team-level questions, from review load to coverage figures to speed-up claims, ask where that cost went.

Context window
The bounded amount of text a model can take into account in one turn. Material outside it does not exist for that turn, however relevant it is.
Inline completion
Suggestions produced where you are typing, drawn from nearby code under a latency budget set by your typing speed, accepted or rejected in one keystroke.
Coding agent
A tool that plans and applies multi-step changes across files and, depending on setup, runs commands, rather than proposing text at a single point.
Steering signal
The check whose result an agent reads to choose its next step, such as a test run or a type check; its precision limits how well the loop converges.
Non-goal
An explicit statement in a request of what the work must not do or change, so the tool does not settle that question on its own.
Checkpoint
A verified state, built and tested and recorded in version control, that a session can be abandoned back to at little cost.
Tautological test
A test whose expected values were copied from the implementation's current output, so it passes by construction and cannot show the implementation is wrong.
Invented dependency
A package or module name a model produces that does not exist; a supply-chain risk if someone registers that name and a build fetches it.
Poisoned premise
An early wrong conclusion that stays in a session's record and is treated as established by later turns, gaining confidence without gaining evidence.
Deterministic check
A check that returns the same verdict on the same input: compiler, type checker, linter, existing test suite. It covers what review no longer needs to.
Provenance
Where a block of code came from; for generated code it matters to licence and copyright questions a team must route and decide.

The two sections describe one range of autonomy seen from opposite ends, and the same concerns reappear at every point on it. **A ladder of autonomy.** Inline completion is judged in a single keystroke. A single-function request produces a block you read whole. Further up, agents open and change whole projects, may run commands, or work remotely and return a finished change. Each step up enlarges the blast radius of an unnoticed error, moves review later, and raises what you need in place before starting: a checkpoint, a fence, a trusted check. The [agentic coding tools](/topics/found-vibe-coding-agentic-tools) section is about choosing the rung from properties of the work. **One pipeline, four stages.** Prompting decides what the tool knows, [task scoping](/topics/found-vibe-coding-task-scoping) decides how much it does between checks, [iteration loops](/topics/found-vibe-coding-iteration-loops) decide what steers it, and review decides what you accept. A defect found in review usually started further up: an unstated constraint, a unit too wide to check, a loop allowed to rewrite its own verdict. The [failure modes](/topics/found-vibe-coding-failure-modes) section collects what happens when one stage is weak. **Review appears in both sections, from two angles.** On the agent side it asks whether a change did what you asked and nothing else. On the assistant side it asks what the build already settles, what a checklist should add, and who answers for the merge. **Tests are output and signal at once.** A generated test is code to review and also the verdict a later loop steers by, so a weak one does damage twice. [Productivity and pitfalls](/topics/found-ai-assisted-coding-pitfalls) then lifts everything to the team: review capacity, skill decay, what may be sent to a tool, and what a speed-up claim can support.

  1. Code Completion →

    The smallest case: what a tool sees from where you type, and what you can honestly check in the moment you accept.

  2. Prompting for Code →

    One bounded request is where you learn which facts a result depends on and what gets filled in when you omit them.

  3. Intent-Driven Prompting →

    The same discipline for multi-turn work: fences, a checkable definition of done, and decisions the run must hand back.

  4. Task Scoping →

    How to cut work into units that each end at a check you trust, before any loop starts.

  5. Iteration Loops →

    What makes a check fit to steer an agent, and how to tell a converging loop from a stalled one.

  6. Reviewing AI Output →

    Where review attention goes once the build has settled what it can, and who answers for a merged change.

  • Claiming a tool writes most of your code without saying how you verify it; the interviewer hears an unchecked pipeline, not productivity.

  • Treating a green run at the end of an agent loop as proof of a fix without checking whether assertions, expected values or skip lists moved.

  • Accepting generated tests whose expected values came from the code under test, then citing the coverage rise as evidence the team is safer.

  • Leaving the language version, framework and existing helpers unstated, then blaming the tool for code that is correct for some other project.

  • Running a whole feature as one long session with no verified checkpoint, so a wrong early decision is buried and abandoning the run is expensive.

  • Responding to a wrong result by repeating the instruction with more emphasis instead of adding the fact that separates the behaviour you want.

  • Answering a team-productivity question with one throughput number instead of describing where effort moved, including review load on colleagues.

  • Stating the copyright position of generated code as settled either way, rather than describing how the team routes and decides the question.

The same few choices come up across both sections, and naming the one you are making is usually the substance of a good answer. - **Autonomy versus reviewability.** A longer unattended run does more per instruction and hands you one large decision at the end that is hard to reconstruct. Small assisted edits cost more of your attention per change and less per mistake. - **More context versus attributable results.** Handing over more material can prevent a wrong guess, but surplus gets imitated and makes it hard to tell which input caused a bad result. The useful measure is whether a piece of context changes a decision. - **Continuing a thread versus starting clean.** Corrections inside a thread keep what is already right; a fresh request drops a structure you never wanted. Which is cheaper depends on whether the error sits in one detail or in the whole approach. - **Fixing the output versus re-running with a better request.** Local repair keeps review already done; a re-run wins when the misunderstanding runs through the whole change. - **Locking the verdict versus trusting the check.** Barring an agent from editing tests and thresholds stops it gaming green, but a test can itself be the bug, and then a locked loop wastes its turns on an outdated expectation. - **Model proposal versus deterministic tool.** A model can suggest any structural edit; a name-resolving tool performs fewer of them with a guarantee. Decide by what can verify the edit, not by how convincing the suggestion looks.

report an issue with this guide →

questions

page 1 of 2

When an inline completer proposes the next line, what material has it actually seen?

level: juniorimportance: must knowfreq 70%

answer

  1. A bounded slice of text
  2. Ask what it had in view
  3. Both sides of the cursor
  4. Plus fragments you did not choose
  5. Outside the slice, outside the answer

basics

~20 s

An inline completer sees a bounded slice of text - the code on both sides of your cursor, plus fragments of other files the editor adds - and nothing else. Most bad suggestions are explained by what fell outside that slice.

solid answer

~40 s

Think of it as a request assembled for you, in the moment you paused. It carries the text before your cursor and usually the text after it, both cut to fit a size budget, plus whatever else the tool decided to add - commonly fragments of files you have open, or ones it retrieved because they resemble the code you are writing. Which fragments, and how they are picked, differs per tool and changes; that the request is bounded does not. Anything not in it was not available for that suggestion: not the ticket, not a convention written down elsewhere in the repository, not a function nobody opened. So when a suggestion is wrong, the first question is not *why did the model reason badly* but *was the deciding fact even in view*.

code

pseudocode · 15 lines
pseudocode
# everything the completer had, for this one suggestion:
#   [a] this file, from the top down to the cursor
#   [b] this file, from the cursor down
#   [c] fragments of other files the editor chose to add - you did not pick them
# nothing else in the service was in [a]-[c].

function listOrdersPage(customerId, pageToken):
    |<- cursor

# suggested:
    return orders.findPage(customerId, pageToken)

# NOT in [a]-[c], and either one would have changed the suggestion:
#   orders.findPage(customerId, cursor, limit)  - three arguments, defined one file away
#   every other handler here returns a page record, never a raw row list

go deeper

for a junior

Be able to say that a completer works from a bounded slice of text around your cursor plus whatever else the editor adds, and that nothing outside it reached the model. That one fact explains most surprising suggestions.

for a middle

Explain the layers and what each contributes: the text above the cursor supplies names and style, the text below constrains what your line has to lead to, and added fragments come from files you did not choose. Then map a wrong suggestion onto a gap.

for a senior

Show that you diagnose in that order on real work - establish what was in view before concluding anything about the model - and that you arrange the work so the material a suggestion needs is near it, rather than complaining the tool did not go and find it.

for a principal

Own the consequence for where the team writes things down: a rule recorded only in a document that nothing opens is invisible to a tool reading code, however clearly it is written. Where a convention is recorded changes which tools can honour it.

## What an inline completer is actually being asked You pause. The editor assembles a request on your behalf, a model continues the text, and the continuation appears where you were about to type. The important word is **assembled**: you did not choose what went into that request, you were not shown it, and it was built to a size budget in the time your pause allowed. Almost every other property of inline completion follows from that one sentence. This topic's own charter names GitHub Copilot and Cursor. This answer deliberately says nothing about what either does today - those details are product-specific, they change, and what follows is about the shape, which changes far more slowly. ## The layers of what it has | layer | what it is | who chose it | what its absence explains | |---|---|---|---| | **prefix** | this file, from the top down to your cursor | the tool, cutting to a budget | a suggestion that ignores something declared far above | | **suffix** | this file, from your cursor down | the tool | a suggestion that does not lead to what the rest of the function needs | | **added fragments** | pieces of other files - commonly ones open in the editor, or ones retrieved because they resemble the code at the cursor | the tool, not you | a suggestion that contradicts a signature or a convention defined elsewhere | | **nothing else** | - | - | everything that was never assembled into the request | The first two layers are the anchor. They are why a completer uses your local variable names and your indentation without being told, and why the suggestion at the top of an empty file is so much worse than the one on line two hundred. The third layer is the one people underestimate in both directions: it is why opening the right file sometimes visibly improves suggestions, and it is also why a suggestion can change from one minute to the next for no reason you can see. ## What is not in there unless something put it there - the ticket or issue you are working from; - a convention recorded in a document nobody opened; - a decision taken in a conversation last week; - the body of a function in a file that was never in view; - what the tests assert about the thing you are writing; - how the program actually behaves when it runs. None of that is a limitation of the model. It is a description of the request. ## Reading a bad suggestion backwards 1. **Was the deciding fact in the request at all?** If the signature it got wrong lives one file away and nothing opened that file, stop - there is nothing to explain about reasoning. 2. **If it was in this file, where?** Material far above the cursor is the first thing a budget cuts. Material below the cursor is available, but a continuation is being written forward from where you stand. 3. **Only then ask whether it reasoned badly** about material it demonstrably had. Knowing a fact was in view does not prove it was used, so step 3 is a real step and not a formality. The value of the order is that the cheap check comes first, and on this subject the cheap check is usually where the answer is. Not every disappointing suggestion is a context gap - a model can produce something poor from a perfectly adequate request. But the context gaps are the class you can do something about, which is why they are worth ruling in or out first. ## Two ecosystems, two amounts inferable How much a completer can get right about a call site it was never shown depends partly on the ecosystem you are in. In a **statically-typed** one, the editor can often supply an exact signature for a symbol whose body was never in view, and a call that does not match is rejected before the code runs - so there is both more to go on and less room for a mismatch to stay quiet. In a **dynamically-typed** one there is frequently less that can be handed over about a symbol, and less that rejects a mismatch until the line executes. Same completer, same discipline; a different amount of the answer is knowable without the definition in front of it. ## Answering this in an interview Describe the shape rather than the product: a bounded request, assembled for you, made of the text around your cursor plus fragments the tool chose. Then give one concrete case - a suggestion that was wrong because the thing deciding it was one file away - and say what you checked first. The sentence carrying this whole subject is *what it could see determines what it could get right*, and it lands far harder with an example attached to it.

  • Two developers at the same cursor in the same file get different suggestions. What could differ?
    What each request carried. Different files open, different recent edits, a longer file cut at a different point, and the tool's own variation between runs. The visible file is not necessarily the whole of what was sent, so an identical cursor does not mean an identical request.
  • A suggestion improved after you opened the file it needed. What does that tell you, and what does it not?
    It tells you the deciding material was outside the request and is now inside it - a context gap rather than a reasoning failure. It does not tell you the new suggestion is right, and it does not promise the same file will be picked up next time. You still read it against the code.

saying these in an interview costs you the question

  • Assumes the completer has the whole repository in view
  • Explains every wrong suggestion as the model being bad at code
  • Believes the completer knows which ticket you are working from
  • Treats a suggested call as evidence that the call is right here
  • Never asks what was outside the window before blaming the tool
open as a page

Why should a one-function code request name the language version and framework you are on?

level: juniorimportance: must knowfreq 60%

basics

~20 s

An unstated environment gets filled in from what is most common, not from what your build accepts. You then get an answer that is correct somewhere else: a newer idiom, a different framework line, a dependency you do not carry.

open as a page

A generated unit test asserts exactly what the function already returns. What is it unable to catch?

level: juniorimportance: must knowfreq 60%

basics

~20 s

It cannot catch the function being wrong. An expectation taken from the implementation agrees with that implementation by construction, so the test reports that behaviour has changed rather than whether the behaviour was ever right.

open as a page

Why does one wrong assumption early in an agent session make the turns that follow more confident rather than less?

level: juniorimportance: must knowfreq 62%

basics

~20 s

An agent's own earlier conclusions sit in the session record, and later turns read them as given rather than as guesses. Each restatement drops the hedge the first one carried, so certainty grows while evidence does not.

open as a page

Why is “add a retry policy to the webhook sender” under-specified even though it reads complete?

level: juniorimportance: must knowfreq 62%

basics

~20 s

The request names an outcome and no behaviour. Retrying a payload the receiver rejected, sending the same one twice, and giving up silently all satisfy it — so the agent picks one, and you usually find out from the diff.

open as a page

A check failed after an agent's change — why hand back the run's own output rather than your summary of it?

level: juniorimportance: must knowfreq 70%

basics

~20 s

The run's output carries the failing case, the expected and actual values, and the line that produced them. A summary replaces those facts with your diagnosis, and if the diagnosis is wrong the next turn inherits it.

open as a page

An agent changed nine files unattended and no author can explain them — how does that change your review?

level: juniorimportance: must knowfreq 60%

basics

~20 s

A generated change carries no author intent, so an unexplained line is not evidence of a reason you lack — it is an open question. Reconstruct intent from the request, verify the risky parts, and answer for what you merge.

open as a page

Why does building a whole feature in one agent run fail differently from building it as five?

level: juniorimportance: must knowfreq 66%

basics

~20 s

A single run has no interior — nothing inside it meets a check you already trusted. An early wrong decision surfaces only after later work is built on it; five units catch it while it is still one unit wide.

open as a page

What do you attach to a one-function code request so the result fits your codebase?

level: middleimportance: must knowfreq 56%

basics

~20 s

Attach what the codebase has already fixed and the answer should not re-choose: the signature the caller expects, the record type it must return, the helper that already exists. Otherwise you get a correct function that needs an adapter.

open as a page

A tool extracts a shared validation routine out of a long handler - what behaviour can silently change?

level: middleimportance: must knowfreq 58%

basics

~20 s

An extraction can preserve every happy path and still move the edges: which order side effects run in, whether an early exit still leaves the caller, which failure escapes, and what a missing value defaults to.

open as a page

Who is accountable for an AI-assisted change once it merges, and what follows for the author?

level: middleimportance: must knowfreq 52%

basics

~20 s

The people, not the tool: the author owns the change they submitted and the reviewer owns the approval, exactly as before. What follows is a bar at submission - do not send code you cannot explain.

open as a page

Which defects in a generated change does the build already catch, and which need a human reviewer?

level: middleimportance: must knowfreq 48%

basics

~20 s

The build decides whatever has a deterministic oracle: resolution, types, style, declared dependencies, the existing suite. A person is needed wherever correctness depends on a rule the generator could not see, and that is where review attention belongs.

open as a page

How do you check a generated test's expected values when the rule they encode is written down nowhere?

level: middleimportance: must knowfreq 55%

basics

~20 s

Split it in two. Settle what you can alone - unit, scale, sign, which side of the boundary the value falls on - then take what is left to whoever owns the rule, as one concrete case with a yes-or-no answer.

open as a page

An agent's change adds a dependency that does not exist — what makes this a security problem?

level: middleimportance: must knowfreq 54%

basics

~20 s

The broken build is not the risk; the reflex is. Fetching a name to make the build pass hands the decision to whoever claimed that name, and a plausible invented name is worth claiming in advance.

open as a page

Your agent retried deliveries the receiver had permanently refused — what do you add to the request?

level: middleimportance: must knowfreq 56%

basics

~20 s

Add the fact that tells the two cases apart — a refusal the receiver would repeat is not retried, a silence is — and name the wrong behaviour as an explicit non-goal. Emphasis carries no information; a fact does.

open as a page

An agent's loop ended with every check green. Why might the fault it was chasing still be there?

level: middleimportance: must knowfreq 62%

basics

~20 s

A loop optimises the verdict it was given, and the verdict is editable. Green can be reached by weakening an assertion, skipping a case or widening a catch — all cheaper than fixing the cause.

open as a page

How do you check that an agent's change did what you asked, and not just something reasonable?

level: middleimportance: must knowfreq 58%

basics

~20 s

Write the request out as observable outcomes before reading the change, then walk it both ways: every outcome present, and every behaviour change traceable to an outcome. The backwards walk is the half reviewers skip.

open as a page

Six months in, a manager asks whether AI coding assistants made the team faster — what can you conclude?

level: seniorimportance: must knowfreq 58%

basics

~10 s

Throughput alone settles nothing: over six months the people, the codebase and the work all changed alongside the tool. Report where effort moved and which narrow claims you are willing to defend.

open as a page

Which properties of the work, not of the tool, decide which AI coding tool shape you should use?

level: seniorimportance: must knowfreq 60%

basics

~20 s

The task decides it: whether a fast, trustworthy check exists; whether the change is wide and mechanical or narrow and subtle; whether you know the code well enough to review at speed; and whether it must be explainable later.

open as a page

You accept an inline suggestion with one keystroke - what can you actually check in that moment?

level: middleimportance: should knowfreq 50%

basics

~20 s

Enough to triage, not to review: whether it does what you were about to write, and whether its names exist here. Anything you cannot settle in that moment should be rejected rather than accepted to check later.

open as a page

Why does a multi-line completion get less trustworthy toward its last line than its first?

level: middleimportance: should knowfreq 55%

basics

~20 s

The first line continues code you wrote and can inspect; every line after it mostly continues text the tool just produced and nobody has checked. One early wrong assumption is then carried, consistently, by everything below it.

open as a page

On a team where only some developers use AI assistants heavily, what changes for everyone else?

level: middleimportance: should knowfreq 46%

basics

~20 s

Reading load rises for people whose own output did not. An assistant multiplies how fast drafts are produced and barely touches how fast anyone reads them, so the constraint moves to review and lands on whoever is absorbing.

open as a page

Which developer skills actually decay when a coding assistant is always available, and which do not?

level: middleimportance: should knowfreq 54%

basics

~10 s

Fluency at producing routine code decays first and returns quickly. Judgement decays only if you stop exercising it, which accepting makes easy. The costlier loss is a model of the system you never built.

open as a page

A code request came back close but wrong — do you amend it in the same thread, or start clean?

level: middleimportance: should knowfreq 42%

basics

~20 s

Amend when the shape was right and one decision was wrong, because everything already established still stands. Restate cleanly when the approach itself was wrong, because corrections there land as patches on a structure you did not want.

open as a page

On one bounded code request, how do you constrain which names the generated code may call?

level: middleimportance: should knowfreq 48%

basics

~20 s

Pin the surface: list the names the code may call with their shapes, close the list against anything new, and ask it to say so rather than substitute when something is missing. Pinning narrows the answer; it does not enforce it.

open as a page

You asked for unit tests on a leave-accrual calculator and got only ordinary dates. What should the request have named?

level: middleimportance: should knowfreq 48%

basics

~20 s

Name the situations, not the quantity. A model drafting from the code produces cases the code already implies. The awkward ones live in the rule, so list them yourself: the period that ends before it starts, the joiner on the boundary.

open as a page

How do an inline completer, an in-editor agent, a terminal agent and a hosted agent differ in what they can see and reach?

level: middleimportance: should knowfreq 52%

basics

~20 s

Four shapes, two axes: what each can read, and what each can act on. A completer writes at one point; an in-editor agent edits an open project; a terminal agent also runs commands; a hosted agent changes only its own copy.

open as a page

You approved every turn of an agent session, so why is the finished change one you would not approve?

level: middleimportance: should knowfreq 48%

basics

~20 s

Each turn is judged against the state the turn before it left, so the baseline moves with the change. A turn that repairs the previous turn's side effect is locally right and carries the session further from what you asked.

open as a page

An agent asked to rename one field is editing its fortieth file — what should have told you sooner?

level: middleimportance: should knowfreq 46%

basics

~20 s

File count is the weakest signal: a rename legitimately reaches every caller. The warning is new territory the request never named — a stored-data change, an externally visible shape, an edited test expectation — and it appears well before the count does.

open as a page

Before an agent starts a multi-turn run on your webhook sender, what must the request state beyond the goal?

level: middleimportance: should knowfreq 50%

basics

~20 s

Three things the goal itself does not carry: the fence, naming artefacts the run must not change; what done means, in terms you could check against the diff; and the decisions it must bring back rather than settle alone.

open as a page

showing 1–30 of 59