AI-assisted development
Working with AI coding tools as part of a normal engineering workflow — how you prompt them, scope work for them, and review what they produce. Interviewers ask because using these tools well, and knowing where they fail, is quickly becoming a baseline expectation.
on this pageshowhide
guide
overview
~1 minAI-assisted development questions test whether you use coding assistants and agents as an engineer or as a passenger. Interviewers rarely care which tool you prefer. They want to hear how you decide what to hand over, how you phrase a request so the result fits your codebase, how you check what comes back, and where you stop trusting it. Claiming heavy use with nothing said about verification raises more concern than light use. The subject has two sections. [Vibe coding](/topics/found-vibe-coding) is the agent-in-the-loop mode: stating intent, cutting work into units an agent can finish and you can check, steering iteration with a real verdict, and recognising how long sessions go wrong. [AI assisted coding](/topics/found-ai-assisted-coding) is the assistant inside the ordinary edit-test-review loop: inline completion, prompting for a single function, refactoring, test generation, review discipline, and what the tools change for a team's productivity, skills and obligations. Learn first how a request becomes a result: what the tool can see, and what it fills in when you leave something out. Then learn review, because most of the other practices exist to make review possible. Questions run from a junior asked why a generated test proves little to a principal asked what a team policy or a short tool trial can actually settle. The material is judgement rather than tool trivia, so it carries across tools.
primer
A few ideas carry the whole subject. Once they are solid, most questions below read as consequences rather than rules to memorise. - **The tool works from a slice and fills the rest with the typical case.** Anything outside what it was given, such as your framework version, an existing utility you expected it to reuse, or a business rule nobody wrote down, gets replaced by whatever is most common. When a result is off target, the first question is what the tool could not see. - **A request is a specification, and its gaps get decided for you.** A goal names an outcome; the tool still has to choose a behaviour. What turns a wish into a task is stating the boundary the work must not cross, a definition of done you could verify from the diff, and which decisions have to come back to you. Repeating yourself louder adds nothing; a fact that separates the right behaviour from the wrong one does. - **Scope is set by the checks you already trust.** A unit of work is the right size when it ends in a state that builds and meets a check that existed before it. A larger unit lets an early mistake sit under everything built afterwards, so plans are ordered to produce evidence early. - **A loop chases whatever verdict steers it.** An agent iterating against a failing check is steering toward green, and the cheapest route there may run through the check itself. A loop is as honest as the list of things it may not edit and as useful as its failure messages are precise. - **Plausible is the expensive failure.** A made-up name fails the first time the build or runtime looks it up. A real call built on a mistaken belief about its contract, structure nobody asked for, or a test built from the code it tests can all pass the machine checks. Human review attention belongs where no deterministic check reaches. - **Accountability stays with people.** Whoever submits a change answers for it, and whoever approves it answers for that approval, however the code was produced. In practice that means submitting only what you could defend line by line. - **Cheap drafts move the bottleneck.** Producing code got faster; reading it, deciding what is correct and maintaining it did not speed up to match. Most team-level questions, from review load to coverage figures to speed-up claims, ask where that cost went.
- Context window
- The bounded amount of text a model can take into account in one turn. Material outside it does not exist for that turn, however relevant it is.
- Inline completion
- Suggestions produced where you are typing, drawn from nearby code under a latency budget set by your typing speed, accepted or rejected in one keystroke.
- Coding agent
- A tool that plans and applies multi-step changes across files and, depending on setup, runs commands, rather than proposing text at a single point.
- Steering signal
- The check whose result an agent reads to choose its next step, such as a test run or a type check; its precision limits how well the loop converges.
- Non-goal
- An explicit statement in a request of what the work must not do or change, so the tool does not settle that question on its own.
- Checkpoint
- A verified state, built and tested and recorded in version control, that a session can be abandoned back to at little cost.
- Tautological test
- A test whose expected values were copied from the implementation's current output, so it passes by construction and cannot show the implementation is wrong.
- Invented dependency
- A package or module name a model produces that does not exist; a supply-chain risk if someone registers that name and a build fetches it.
- Poisoned premise
- An early wrong conclusion that stays in a session's record and is treated as established by later turns, gaining confidence without gaining evidence.
- Deterministic check
- A check that returns the same verdict on the same input: compiler, type checker, linter, existing test suite. It covers what review no longer needs to.
- Provenance
- Where a block of code came from; for generated code it matters to licence and copyright questions a team must route and decide.
The two sections describe one range of autonomy seen from opposite ends, and the same concerns reappear at every point on it. **A ladder of autonomy.** Inline completion is judged in a single keystroke. A single-function request produces a block you read whole. Further up, agents open and change whole projects, may run commands, or work remotely and return a finished change. Each step up enlarges the blast radius of an unnoticed error, moves review later, and raises what you need in place before starting: a checkpoint, a fence, a trusted check. The [agentic coding tools](/topics/found-vibe-coding-agentic-tools) section is about choosing the rung from properties of the work. **One pipeline, four stages.** Prompting decides what the tool knows, [task scoping](/topics/found-vibe-coding-task-scoping) decides how much it does between checks, [iteration loops](/topics/found-vibe-coding-iteration-loops) decide what steers it, and review decides what you accept. A defect found in review usually started further up: an unstated constraint, a unit too wide to check, a loop allowed to rewrite its own verdict. The [failure modes](/topics/found-vibe-coding-failure-modes) section collects what happens when one stage is weak. **Review appears in both sections, from two angles.** On the agent side it asks whether a change did what you asked and nothing else. On the assistant side it asks what the build already settles, what a checklist should add, and who answers for the merge. **Tests are output and signal at once.** A generated test is code to review and also the verdict a later loop steers by, so a weak one does damage twice. [Productivity and pitfalls](/topics/found-ai-assisted-coding-pitfalls) then lifts everything to the team: review capacity, skill decay, what may be sent to a tool, and what a speed-up claim can support.
- Code Completion →
The smallest case: what a tool sees from where you type, and what you can honestly check in the moment you accept.
- Prompting for Code →
One bounded request is where you learn which facts a result depends on and what gets filled in when you omit them.
- Intent-Driven Prompting →
The same discipline for multi-turn work: fences, a checkable definition of done, and decisions the run must hand back.
- Task Scoping →
How to cut work into units that each end at a check you trust, before any loop starts.
- Iteration Loops →
What makes a check fit to steer an agent, and how to tell a converging loop from a stalled one.
- Reviewing AI Output →
Where review attention goes once the build has settled what it can, and who answers for a merged change.
Claiming a tool writes most of your code without saying how you verify it; the interviewer hears an unchecked pipeline, not productivity.
Treating a green run at the end of an agent loop as proof of a fix without checking whether assertions, expected values or skip lists moved.
Accepting generated tests whose expected values came from the code under test, then citing the coverage rise as evidence the team is safer.
Leaving the language version, framework and existing helpers unstated, then blaming the tool for code that is correct for some other project.
Running a whole feature as one long session with no verified checkpoint, so a wrong early decision is buried and abandoning the run is expensive.
Responding to a wrong result by repeating the instruction with more emphasis instead of adding the fact that separates the behaviour you want.
Answering a team-productivity question with one throughput number instead of describing where effort moved, including review load on colleagues.
Stating the copyright position of generated code as settled either way, rather than describing how the team routes and decides the question.
The same few choices come up across both sections, and naming the one you are making is usually the substance of a good answer. - **Autonomy versus reviewability.** A longer unattended run does more per instruction and hands you one large decision at the end that is hard to reconstruct. Small assisted edits cost more of your attention per change and less per mistake. - **More context versus attributable results.** Handing over more material can prevent a wrong guess, but surplus gets imitated and makes it hard to tell which input caused a bad result. The useful measure is whether a piece of context changes a decision. - **Continuing a thread versus starting clean.** Corrections inside a thread keep what is already right; a fresh request drops a structure you never wanted. Which is cheaper depends on whether the error sits in one detail or in the whole approach. - **Fixing the output versus re-running with a better request.** Local repair keeps review already done; a re-run wins when the misunderstanding runs through the whole change. - **Locking the verdict versus trusting the check.** Barring an agent from editing tests and thresholds stops it gaming green, but a test can itself be the bug, and then a locked loop wastes its turns on an outdated expectation. - **Model proposal versus deterministic tool.** A model can suggest any structural edit; a name-resolving tool performs fewer of them with a guarantee. Decide by what can verify the edit, not by how convincing the suggestion looks.
explore
- Vibe Coding29 questions
- Intent-Driven Prompting5 questions
- Task Scoping5 questions
- Reviewing AI Output5 questions
- Iteration Loops5 questions
- Agentic Coding Tools4 questions
- Failure Modes5 questions
- AI Assisted Coding30 questions
- Code Completion5 questions
- Prompting for Code5 questions
- AI-Assisted Refactoring4 questions
- Test Generation5 questions
- Reviewing AI Output6 questions
- Productivity & Pitfalls5 questions
- AI & Data Scientistroleanchors this topic
- AI Engineerroleanchors this topic
- Backend Developerroleanchors this topic
- Blockchain Developerroleanchors this topic
- Data Engineerroleanchors this topic
- DevOps / SRE Engineerroleanchors this topic
- Frontend Developerroleanchors this topic
- Full Stack Developerroleanchors this topic
- Java Backend Developerroleanchors this topic
- Java SDETroleanchors this topic
- Kotlin Backend Developerroleanchors this topic
- QA Engineerroleanchors this topic
- Vibe Codingskillanchors this topic
- BI Analystrole
questions
page 2 of 2Before an agent starts iterating on a change, what makes a check usable as the loop's steering signal?
basics
~20 sA usable verdict runs without you, returns the same answer on the same input, is cheap enough to run every turn, and names what failed rather than that something did — and it has to be failing now, for the right reason.
In a generated change, why is a plausible call to a real API harder to catch than an invented one?
basics
~20 sA name that does not exist fails the moment anything resolves it. A real name used on a wrong assumption resolves and fits the code around it, since one run wrote both sides; the disagreement is with the real contract.
Your five agent units can only be tested once the fourth lands — how do you re-cut the plan?
basics
~10 sRe-order so evidence arrives first: open with a behaviour-preserving unit the existing suite already checks, then the smallest unit that settles the riskiest assumption. Merge any unit that cannot fail on its own.
What tells you a unit of work is too big for one agent run before you start it?
basics
~10 sThree tells, all visible before you type: the unit has no end state where the project still builds, no check that existed before the unit did, and no area you can name.
In a long agent coding session, what fills the context window, and why does the run degrade rather than stop?
basics
~20 sA long session is filled mostly by its own by-products: files read entire, full command output, failed attempts, its own restatements. The room is finite, so older material goes and the work carries on with less.
In a service you joined last week, where should you check an inline completer's suggestions hardest?
basics
~20 sWherever this service departs from the widespread way of doing something. A continuation follows shapes that recur, so on your unusual code the recurring shape is the wrong one - and it reads as entirely reasonable to anyone new here.
How do you request a structural edit on a file the assistant cannot read in one piece?
basics
~20 sGive the tool a complete view of something small, not a partial view of something large: the block being moved, the uses you found yourself, the tests that pin it. What it cannot see, it fills in.
How do you review a rename an assistant applied across forty files without reading all forty?
basics
~20 sReview by shape, not by file: decide what every hunk should look like, read the ones that break the pattern, and search separately for the names a rename cannot reach - in configuration, stored data or message text.
What belongs on a review checklist for AI-assisted changes that the ordinary checklist does not?
basics
~20 sOnly what the build and the ordinary checklist cannot already cover. In practice a handful of lines: invariants the generator could not see, the caller's authority, provenance for a distinctive block, and whether the author can explain every line.
How should a review handle the licence and provenance of generated code while the law is unsettled?
basics
~20 sSurface the question, do not rule on it. Reviewers flag blocks distinctive enough to need a provenance answer and route them to whoever owns that risk; the team decides once, and sets a higher bar for code it redistributes.
Which security weaknesses deserve a named line on a review checklist for generated code?
basics
~10 sThe ones whose correctness depends on something outside the file the generator was working in: the caller's authority, the trust boundary the input crossed, how credentials are obtained, and what the failure path discloses.
Your coverage number rose after a batch of generated unit tests. Why might the team be no safer?
basics
~20 sThe number counts code that ran. Generation made running code cheap and left the hard half - deciding what the right answer is - exactly as expensive, so a figure that stood in for effort no longer does.
Beyond the change itself, what does a long autonomous agent run cost you that a small assisted edit does not?
basics
~20 sA long autonomous run concentrates a large change into one review decision, widens what a mistake can touch before anyone notices, leaves less of the session reconstructable, and produces a result a re-run will not reproduce.
What has to be true before an agent session starts for abandoning it later to be cheap?
basics
~20 sA point you verified — built, tested, recorded in version control — and a run whose reach stays inside that tree. Recovery cost is set by what the session could touch outside it, not by how many files changed.
How much surrounding code do you hand an agent at the start of a session, and what does too much cost?
basics
~20 sHand over what changes a decision the run would otherwise get wrong — the contract, the record shape, one example of the code you want. Surplus material gets imitated, and it makes a wrong result impossible to attribute.
Four turns in, each fix breaks a neighbouring case — how do you tell a stalled loop from a converging one?
basics
~20 sWatch the failing set, not the last message. A converging loop shrinks it and keeps it shrunk; a stalled one moves it sideways — same size, different members. Fix the stopping condition before turn one, while the sunk cost is small.
An agent's change adds an interface with a single implementation — how do you decide whether to keep it?
basics
~20 sName the second implementation, caller or value the structure exists for. If it exists today or is genuinely planned, keep it; if it is only imaginable, the code is paying now for a prediction, and it comes out.
What must a team's policy on what may be sent to an AI coding tool actually decide?
basics
~20 sWhich repositories a tool may be attached to, what it may see once attached, what never goes anywhere, the default for the unlisted case and who answers it quickly. A rule phrased as per-paste judgement decides nothing.
Which structural edits do you let a model propose, and which do you insist a deterministic tool perform?
basics
~20 sSort the edit by what can check it, not by how good the suggestion looks. Name resolution covers mechanical edits inside the scope it indexes, a build check catches shape, and order, defaults and failure need a test.
After a two-week team trial of an agentic coding tool, what can you legitimately conclude from it?
basics
~20 sA short trial mostly tells you about yourselves: how well your checks hold, how much of a run you can reconstruct, where your codebase resists being read. It generalises poorly to untried areas, other teams, and later versions.
Which engineering goals can you state well enough to hand to an agent run, and which can you not?
basics
~20 sA goal is statable when its acceptance exists apart from the implementation. Where acceptance is what the work discovers, or two teams disagree what done means, no wording fixes it — settle it first, or run to learn and discard.
In a loop steered by a failing check, what should the agent be barred from editing, and what does barring it cost?
basics
~20 sWhatever decides the verdict should not move silently: assertions, expected values, skip lists, rule configuration, the threshold that fails the build. The cost is real — checks are sometimes wrong, and a locked loop burns turns failing against a stale expectation.
An inline completer must answer in the pause between keystrokes - what does that deadline cost?
basics
~20 sYour typing sets the deadline, not the tool: a suggestion that arrives after you wrote the line is worthless however good it is. Everything done to meet that deadline narrows what the suggestion could account for.
Copyright in model-generated code is unsettled — what can a team still decide for itself?
basics
~20 sEverything about its own exposure: what it accepts into which products, what its existing contracts and licence obligations already require, and who it asks. The law is not the team's to settle, and stating it as settled is the error.
On a code request where the approach is the risk, why ask for that before any code?
basics
~20 sAsking for the approach first moves your review to the cheapest artefact: four lines read and rejected in seconds, against a working function that is expensive to read and hard to reject once it runs. Use it where the decisions carry the risk.
When is re-running an agent with a corrected request cheaper than repairing the change it produced?
basics
~20 sRe-run when the premise is wrong — a misread request or an unstated constraint — because that mistake is spread through every file. Repair when the premise holds and the defects are few. A re-run discards the review already done.
When a feature is split across several agent runs, what stops the fourth reinventing what the first built?
basics
~20 sOnly what is legible in the code carries across a boundary. A later run reads the repository, not the earlier run's reasoning — so a decision survives as a helper with no route around it, or as a failing test.
Should a change record that an AI tool helped write it, and what should that record change?
basics
~20 sRecord it as a fact, for provenance later, and do not turn it into a scrutiny tier. Review attention belongs on what a change touches and can break, which is verifiable, rather than on a self-reported label.
showing 31–59 of 59