Prompt engineering
Writing the instruction the model actually follows: how you structure the task, when examples help, how you pin down output format, and how you iterate with measurement instead of vibes. Interviewers treat this as the baseline skill for anyone building on LLMs.
on this pageshowhide
guide
overview
~1 minPrompt engineering is the working skill behind any product that puts a language model in front of a task. Interviewers probe it less as wordcraft than as engineering discipline: can you tell a vague instruction from a precise one, explain why a prompt that worked on five inputs fails on the sixth, and show that a change actually helped rather than merely looked better on the example you had open. At senior levels the questions shift toward where a prompt stops being the right tool, and which guarantees have to live in code instead. The subject splits into six sections. [Instruction design](/topics/found-prompt-engineering-instruction-design) covers how the task is phrased, structured and broken into stages. [Few-shot prompting](/topics/found-prompt-engineering-few-shot-prompting), the largest, is teaching by demonstration, and how the choice, order and format of examples steer the result. [System and role prompts](/topics/found-prompt-engineering-system-role-prompts) deals with the layers of a chat request and what a standing instruction can and cannot hold. [Chain-of-thought](/topics/found-prompt-engineering-chain-of-thought) asks when making the model reason first is worth its tokens. [Output formatting](/topics/found-prompt-engineering-output-control) is about output a program can consume, and [evaluation and iteration](/topics/found-prompt-engineering-evaluation-iteration) is about measuring all of the above. Start with instruction design, since every other technique modifies a base instruction, then few-shot prompting. Take evaluation third rather than last: without it, the advice in every other section is untestable. The same idea often returns at a deeper level in senior questions, so revisit a section as your level rises.
primer
A few ideas sit under every section. Hold these and most of the questions below read as consequences rather than facts to memorise. - **The model continues text; it does not execute a contract.** The instruction, the examples, the conversation so far and any pasted document arrive as one sequence the model extends. Wording shifts probabilities rather than fixing outcomes, which is why a format request occasionally fails, why rules fade in long chats and why pasted content can behave like an order. - **What you show tends to outweigh what you say.** When demonstrations and the written instruction disagree, the demonstrations usually win. Examples define the answer shape, the set of labels in play and even the implied base rates, so a careless example set can quietly rewrite a careful instruction. - **Precise, structured and deliberately placed.** A good instruction names the task, the audience, the constraints and the output in concrete terms, marks off material from instructions, and sits where the model will weigh it. Many "the model is bad at this" complaints dissolve once the task is stated precisely or split into smaller steps. - **A prompt lowers a failure rate; code bounds it.** Anything where one violation counts as an incident, such as a spending limit, a permission or a format a machine must read, needs a check outside the model: a tool that refuses, a validator, constrained decoding. - **Every technique has a price.** More examples, reasoning steps, extra calls in a chain and repeated sampling each add tokens and latency to every request. The right amount is the smallest that measurably fixes a failure you have actually observed. - **Measure instead of eyeballing.** A prompt tuned against the handful of inputs you keep rereading is overfit to them. A fixed evaluation set, pinned conditions and a record of which prompt version produced which output turn prompt editing from opinion into engineering.
- Zero-shot prompting
- Asking for a task with an instruction alone and no worked examples; the baseline every added technique should be measured against.
- Few-shot prompting
- Including a small number of input-output demonstrations in the prompt so the model infers the task, the answer shape and the label set from them.
- Exemplar
- One demonstration in a few-shot prompt: an example input paired with the output the model should imitate.
- Delimiter
- A marker, such as a named tag or a fence, that separates parts of a prompt, typically the instructions from the material they operate on.
- System prompt
- The standing instruction an application places ahead of the conversation to set role, rules and tone; weighted heavily by the model but enforced by nothing.
- Instruction hierarchy
- The trained ordering a model uses to weigh conflicting instructions from different layers of a request, with retrieved and tool content meant to carry no authority.
- Prompt injection
- Text inside material the model processes, such as a document, web page or tool result, written to be read as instructions and redirect its behaviour.
- Chain-of-thought
- Prompting the model to produce intermediate reasoning before its final answer, which tends to help multi-step problems and adds cost everywhere else.
- Self-consistency
- Sampling several independent reasoning chains for one question and keeping the most common final answer, trading multiplied cost for accuracy on discrete answers.
- Constrained decoding
- Restricting which tokens the model may generate at each step, usually from a schema or grammar, so the output cannot break the required structure.
- Label bias
- A skew toward particular answers caused by the prompt itself, such as the mix, order or recency of labels in the examples, rather than by the input.
- Held-out test set
- Evaluation inputs set aside and not read during tuning, so their score estimates behaviour on unseen traffic rather than on the examples you tuned against.
- Ablation
- Removing one part of a prompt at a time and rescoring an unchanged evaluation set to learn which parts actually affect results.
- LLM judge
- A model that scores another model's outputs against a rubric; quick to scale, but itself a component to validate and pin.
The sections are layers of one prompt and one working method, not separate tricks. **Instruction design is the base layer.** Every other technique modifies an instruction that should already be clear. Delimiters and a stable section order are what later let examples, retrieved documents and conversation history be inserted by code without blurring into the instructions. Splitting a task into a chain is also where [output formatting](/topics/found-prompt-engineering-output-control) starts to matter, because every seam between stages is a place to parse and check. **Examples and instructions compete.** Demonstrations are a second channel of instruction, and the [demonstration format](/topics/found-prompt-engineering-few-shot-prompting-demonstration-format) section shows how that channel can override the first. Chain-of-thought often arrives through the same channel, as worked examples that show their reasoning, so the two sections share questions about format and cost. Choosing examples per request ties few-shot prompting to retrieval and to prompt caching. **Roles decide whose words count.** The [system and role prompts](/topics/found-prompt-engineering-system-role-prompts) section sorts everything above into layers: application rules, user turns, and pasted or retrieved content that should be treated as material. That is where prompt engineering meets security, and where the recurring answer is to move hard limits into tools and validators. **Evaluation closes the loop.** Every choice in the other sections, from how many examples and in which order to whether to reason first or whether to merge a chain, is an empirical question. [Evaluation and iteration](/topics/found-prompt-engineering-evaluation-iteration) supplies the method for answering it, which is why its ideas appear in senior questions across all six sections.
- Instruction Design →
Every other technique edits a base instruction; learn to state a task precisely and keep material apart from instructions first.
- Few-Shot Prompting →
How demonstrations teach the answer shape and label set, and how they can quietly contradict the instruction they accompany.
- Evaluation and Iteration →
Learn to measure a change before collecting more techniques; later choices about examples, reasoning and chains depend on it.
- System and Role Prompts →
The layers of a chat request, and why a standing instruction is guidance rather than a guarantee.
- Output Formatting →
Needed as soon as code reads the answer: schemas, validation, and what happens when output fails to parse.
- Chain-of-Thought →
Last, because deciding when reasoning pays relies on the cost awareness and measurement habits the earlier sections build.
Treating a system prompt as a security boundary; a limit written in prose holds only as often as the model complies, so enforce it in the tool or service.
Pasting untrusted text into a prompt unmarked, so the model cannot tell material to process from instructions to follow; marking it helps but does not make it safe.
Writing a careful instruction, then supplying examples that show a different answer shape; the examples usually win.
Leaving a label out of the few-shot set, or giving one label most of the examples, and then blaming the model for a skewed classifier.
Trusting a format instruction to produce parseable output and then scraping or coercing whatever came back, instead of validating it and rejecting failures.
Adding a step-by-step reasoning trigger to every request; on lookup and formatting work it adds latency and cost and can make answers worse.
Declaring a prompt better after rereading the same few inputs, with no held-out set and with the model version or sampling settings left unpinned.
Most senior questions in this hub are one of these choices in disguise, and naming the one you are making is often most of the answer. - **Describing versus demonstrating.** A written rule is cheap and transparent; examples convey nuance a description misses, but they cost tokens on every request and can teach rules you never intended. - **One call versus a chain.** A single call is faster and simpler to operate; a chain gives each step a narrow job and a place to check it, at the price of latency, more moving parts and errors that travel between stages. - **Fixed versus retrieved examples.** A fixed block is predictable and friendly to prompt caching; picking examples per request fits the input better but makes behaviour depend on the pool's contents and freshness. - **Reasoning versus a direct answer.** Asking for reasoning helps multi-step problems and costs tokens and latency everywhere else, so the decision belongs to a segment of traffic, not to the whole system. - **Prompt versus code.** Wording lowers a failure rate; a validator, a tool check or constrained decoding bounds it. The more a failure costs, the further its guard belongs from the prompt. - **Strict versus expressive output.** A tight schema makes output easy to consume; too tight, and it can force the model to invent a value or choose a wrong option rather than say it does not know.
Several shapes recur across the sections under different names; spotting one is often the quickest way into an unfamiliar question. - **Separate material from instructions.** Delimiters around documents, a fixed prompt skeleton, and tool or retrieved content treated as data rather than orders are the same idea at the text, template and role level. - **Check every seam.** Between chain stages, after structured output, and at the entrance to an example pool: validate what crosses the boundary and decide in advance what happens when it fails. - **Restate near the point of generation.** Rules that fade over a long conversation or a long output are repeated close to where the model is writing, not only at the top. - **Escalate only on measured failure.** Instruction alone, then examples, then reasoning, then sampling and voting: each rung is climbed when the cheaper one is shown to fall short. - **Change one variable, rescore the same set.** Ablations, shot-count sweeps, order searches and chain-versus-single-call comparisons are one experimental design pointed at different variables.
explore
- Instruction Design9 questions
- Task Decomposition & Prompt Chaining5 questions
- Delimiters & Prompt Structure4 questions
- Few-Shot Prompting17 questions
- Exemplar Selection5 questions
- Dynamic Exemplar Retrieval4 questions
- Ordering Effects & Label Bias4 questions
- Demonstration Format4 questions
- System and Role Prompts9 questions
- Instruction Hierarchy & Conflicts5 questions
- Multi-Turn Instruction Drift4 questions
- Chain-of-Thought4 questions
- Output Formatting5 questions
- Evaluation and Iteration5 questions
questions
page 2 of 2How does failure-mode analysis of wrong outputs drive the next prompt revision?
basics
~20 sRead the wrong outputs, group them by hand into a few named failure modes, and count each. The biggest group sets the next revision's target, and each mode implies a different edit — or tells you the fix is not a prompt edit at all.
How do few-shot demonstrations signal the label space, and what happens to unseen classes?
basics
~20 sThe label strings appearing in the demonstrations act as the effective output vocabulary. A class named only in the instruction but never demonstrated is emitted rarely or never, so every class you want back must appear at least once, spelled exactly as your parser expects.
A retrieved-exemplar store keeps teaching a product name retired six months ago — how do you fix it?
basics
~20 sTreat the exemplar pool as a versioned production dataset, not a folder of examples. Give every entry a timestamp and a validity window, filter expired entries out at retrieval, verify anything written back before it becomes teachable, and audit the pool against the current taxonomy on a schedule.
Why does a three-approve, one-deny shot set push a few-shot classifier toward approve?
basics
~20 sMajority-label bias: the model reads the label mix in the demonstrations as evidence about how often each outcome occurs, so a 3:1 approve-heavy shot set inflates the predicted approve rate on borderline building-permit applications, independent of what each application actually says.
Why add negative exemplars showing what a classifier must not flag?
basics
~20 sPositive-only demonstrations teach where the rule applies but never where it stops, so the model over-generalizes and flags look-alikes. Near-miss negatives — cases that resemble a hit but are not one — draw the boundary the positives leave undefined.
How do you handle source text that contains the same delimiter your prompt uses?
basics
~20 sHandle it at interpolation time, before the prompt is built: pick a delimiter the content cannot contain, escape or normalise the offending marker in the span, or extend the fence beyond any run inside. Never concatenate raw content and hope.
Why does an LLM's long structured output drift, like row 60 of a 100-row table losing a field?
basics
~20 sAdherence decays over a long generation: the format spec recedes into distant context, each new row is copied from nearby rows rather than the schema, and one omission propagates. Fix by batching, per-row validation with row-count postconditions, and constrained decoding.
Why is an LLM's instruction hierarchy a trained preference rather than an enforced rule?
basics
~20 sNothing in the serving stack checks compliance. Every layer arrives as tokens in one context window, separated only by role markers the model was post-trained to weight. That produces a strong statistical preference that can be argued down, not a rule the runtime enforces.
How would you build a regression eval that catches multi-turn instruction drift?
basics
~20 sReplay fixed conversation scripts and assert the same constraint at several depths — say turn 5, 20 and 50 — instead of testing one-shot. Run each script repeatedly because sampling is nondeterministic, and report an adherence curve per depth so a regression shows up as decay, not a single failure.
How much explicit chain-of-thought prompting is redundant on reasoning models?
basics
~20 sMost of the trigger phrasing is. Models that deliberate internally already produce a chain, and long hand-written reasoning scaffolds can cut across it. What still pays is task-specific direction — what to check, in what order, and what the answer must contain.
How would you manage production prompts as versioned artifacts across a team?
basics
~20 sTreat each prompt as code: a reviewed file under version control, an identifier recorded with every production response, an eval run attached to each change, a named owner, and a rollback that is one deploy away.
After searching twenty exemplar orders on a held-out dev set, is the winning order worth shipping?
basics
~20 sOnly after a second, untouched split confirms the margin. Picking the best of twenty candidates inflates the winner's dev score by selection noise, and orders tuned on one model version frequently stop being best on the next — so the ongoing revalidation cost is part of the decision.
How do you standardize section order across a library of thirty production prompts?
basics
~20 sFix one skeleton — role, rules, tools, data, task — and make every prompt fill its slots rather than invent an order. The gain is reviewability, clean diffs, eval attribution and a stable shared prefix; the cost is ceremony, so keep a documented escape hatch.
How would you decide whether to collapse a five-call prompt chain into one call?
basics
~20 sTreat it as an experiment, not a preference. Run both variants over a labelled set, compare quality, p95 latency, and cost per request including retries, then price the observability and routing you give up. Collapse adjacent stages whose boundary carries no check.
When does forcing a strict JSON schema on an LLM hurt output quality?
basics
~20 sWhen the schema makes truth unrepresentable or removes reasoning room. Required non-nullable fields push the model to fabricate, a closed enum forces a wrong nearest label, and constraining from the first token denies it space to think before committing.
How should a multi-tenant LLM platform resolve tenant rules that conflict with its own policy?
basics
~20 sPlatform policy is a floor tenants may narrow but never widen. Assemble the prompt server-side so tenant text sits in a lower layer it cannot edit or precede, define the default resolution as failing toward the platform rule, and enforce anything consequential in tools rather than prose.
How do you triage persona slip versus hard-rule violations in long chats?
basics
~20 sSplit standing constraints by blast radius. Persona and tone slip is a quality signal with a tolerable decay budget, managed by measurement and light re-anchoring. Rules whose violation is an incident — prohibitions, disclosures, machine-parsed formats — must not depend on prompt adherence at all; enforce them outside the model.
What is skeleton-of-thought prompting, and when does parallel expansion hurt quality?
basics
~20 sSkeleton-of-thought first asks for a terse outline, then expands each outline point in its own call, running the expansions concurrently and stitching the results. It cuts wall-clock time for long outputs but breaks down when the sections depend on each other.
Who owns a few-shot exemplar pool, and when must it be refreshed?
basics
~20 sAn exemplar pool is policy encoded as data, so it needs a named owner in the domain team, provenance and PII review on every entry, versioning alongside the prompt, and a refresh triggered by policy changes, label-set changes and drift — not by the calendar alone.
showing 31–49 of 49