skip to content

Prompt engineering

7 roadmaps49 questionsupdated

Writing the instruction the model actually follows: how you structure the task, when examples help, how you pin down output format, and how you iterate with measurement instead of vibes. Interviewers treat this as the baseline skill for anyone building on LLMs.

on this pageshow

guide

overview

~1 min

Prompt engineering is the working skill behind any product that puts a language model in front of a task. Interviewers probe it less as wordcraft than as engineering discipline: can you tell a vague instruction from a precise one, explain why a prompt that worked on five inputs fails on the sixth, and show that a change actually helped rather than merely looked better on the example you had open. At senior levels the questions shift toward where a prompt stops being the right tool, and which guarantees have to live in code instead. The subject splits into six sections. [Instruction design](/topics/found-prompt-engineering-instruction-design) covers how the task is phrased, structured and broken into stages. [Few-shot prompting](/topics/found-prompt-engineering-few-shot-prompting), the largest, is teaching by demonstration, and how the choice, order and format of examples steer the result. [System and role prompts](/topics/found-prompt-engineering-system-role-prompts) deals with the layers of a chat request and what a standing instruction can and cannot hold. [Chain-of-thought](/topics/found-prompt-engineering-chain-of-thought) asks when making the model reason first is worth its tokens. [Output formatting](/topics/found-prompt-engineering-output-control) is about output a program can consume, and [evaluation and iteration](/topics/found-prompt-engineering-evaluation-iteration) is about measuring all of the above. Start with instruction design, since every other technique modifies a base instruction, then few-shot prompting. Take evaluation third rather than last: without it, the advice in every other section is untestable. The same idea often returns at a deeper level in senior questions, so revisit a section as your level rises.

primer

A few ideas sit under every section. Hold these and most of the questions below read as consequences rather than facts to memorise. - **The model continues text; it does not execute a contract.** The instruction, the examples, the conversation so far and any pasted document arrive as one sequence the model extends. Wording shifts probabilities rather than fixing outcomes, which is why a format request occasionally fails, why rules fade in long chats and why pasted content can behave like an order. - **What you show tends to outweigh what you say.** When demonstrations and the written instruction disagree, the demonstrations usually win. Examples define the answer shape, the set of labels in play and even the implied base rates, so a careless example set can quietly rewrite a careful instruction. - **Precise, structured and deliberately placed.** A good instruction names the task, the audience, the constraints and the output in concrete terms, marks off material from instructions, and sits where the model will weigh it. Many "the model is bad at this" complaints dissolve once the task is stated precisely or split into smaller steps. - **A prompt lowers a failure rate; code bounds it.** Anything where one violation counts as an incident, such as a spending limit, a permission or a format a machine must read, needs a check outside the model: a tool that refuses, a validator, constrained decoding. - **Every technique has a price.** More examples, reasoning steps, extra calls in a chain and repeated sampling each add tokens and latency to every request. The right amount is the smallest that measurably fixes a failure you have actually observed. - **Measure instead of eyeballing.** A prompt tuned against the handful of inputs you keep rereading is overfit to them. A fixed evaluation set, pinned conditions and a record of which prompt version produced which output turn prompt editing from opinion into engineering.

Zero-shot prompting
Asking for a task with an instruction alone and no worked examples; the baseline every added technique should be measured against.
Few-shot prompting
Including a small number of input-output demonstrations in the prompt so the model infers the task, the answer shape and the label set from them.
Exemplar
One demonstration in a few-shot prompt: an example input paired with the output the model should imitate.
Delimiter
A marker, such as a named tag or a fence, that separates parts of a prompt, typically the instructions from the material they operate on.
System prompt
The standing instruction an application places ahead of the conversation to set role, rules and tone; weighted heavily by the model but enforced by nothing.
Instruction hierarchy
The trained ordering a model uses to weigh conflicting instructions from different layers of a request, with retrieved and tool content meant to carry no authority.
Prompt injection
Text inside material the model processes, such as a document, web page or tool result, written to be read as instructions and redirect its behaviour.
Chain-of-thought
Prompting the model to produce intermediate reasoning before its final answer, which tends to help multi-step problems and adds cost everywhere else.
Self-consistency
Sampling several independent reasoning chains for one question and keeping the most common final answer, trading multiplied cost for accuracy on discrete answers.
Constrained decoding
Restricting which tokens the model may generate at each step, usually from a schema or grammar, so the output cannot break the required structure.
Label bias
A skew toward particular answers caused by the prompt itself, such as the mix, order or recency of labels in the examples, rather than by the input.
Held-out test set
Evaluation inputs set aside and not read during tuning, so their score estimates behaviour on unseen traffic rather than on the examples you tuned against.
Ablation
Removing one part of a prompt at a time and rescoring an unchanged evaluation set to learn which parts actually affect results.
LLM judge
A model that scores another model's outputs against a rubric; quick to scale, but itself a component to validate and pin.

The sections are layers of one prompt and one working method, not separate tricks. **Instruction design is the base layer.** Every other technique modifies an instruction that should already be clear. Delimiters and a stable section order are what later let examples, retrieved documents and conversation history be inserted by code without blurring into the instructions. Splitting a task into a chain is also where [output formatting](/topics/found-prompt-engineering-output-control) starts to matter, because every seam between stages is a place to parse and check. **Examples and instructions compete.** Demonstrations are a second channel of instruction, and the [demonstration format](/topics/found-prompt-engineering-few-shot-prompting-demonstration-format) section shows how that channel can override the first. Chain-of-thought often arrives through the same channel, as worked examples that show their reasoning, so the two sections share questions about format and cost. Choosing examples per request ties few-shot prompting to retrieval and to prompt caching. **Roles decide whose words count.** The [system and role prompts](/topics/found-prompt-engineering-system-role-prompts) section sorts everything above into layers: application rules, user turns, and pasted or retrieved content that should be treated as material. That is where prompt engineering meets security, and where the recurring answer is to move hard limits into tools and validators. **Evaluation closes the loop.** Every choice in the other sections, from how many examples and in which order to whether to reason first or whether to merge a chain, is an empirical question. [Evaluation and iteration](/topics/found-prompt-engineering-evaluation-iteration) supplies the method for answering it, which is why its ideas appear in senior questions across all six sections.

  1. Instruction Design →

    Every other technique edits a base instruction; learn to state a task precisely and keep material apart from instructions first.

  2. Few-Shot Prompting →

    How demonstrations teach the answer shape and label set, and how they can quietly contradict the instruction they accompany.

  3. Evaluation and Iteration →

    Learn to measure a change before collecting more techniques; later choices about examples, reasoning and chains depend on it.

  4. System and Role Prompts →

    The layers of a chat request, and why a standing instruction is guidance rather than a guarantee.

  5. Output Formatting →

    Needed as soon as code reads the answer: schemas, validation, and what happens when output fails to parse.

  6. Chain-of-Thought →

    Last, because deciding when reasoning pays relies on the cost awareness and measurement habits the earlier sections build.

  • Treating a system prompt as a security boundary; a limit written in prose holds only as often as the model complies, so enforce it in the tool or service.

  • Pasting untrusted text into a prompt unmarked, so the model cannot tell material to process from instructions to follow; marking it helps but does not make it safe.

  • Writing a careful instruction, then supplying examples that show a different answer shape; the examples usually win.

  • Leaving a label out of the few-shot set, or giving one label most of the examples, and then blaming the model for a skewed classifier.

  • Trusting a format instruction to produce parseable output and then scraping or coercing whatever came back, instead of validating it and rejecting failures.

  • Adding a step-by-step reasoning trigger to every request; on lookup and formatting work it adds latency and cost and can make answers worse.

  • Declaring a prompt better after rereading the same few inputs, with no held-out set and with the model version or sampling settings left unpinned.

Most senior questions in this hub are one of these choices in disguise, and naming the one you are making is often most of the answer. - **Describing versus demonstrating.** A written rule is cheap and transparent; examples convey nuance a description misses, but they cost tokens on every request and can teach rules you never intended. - **One call versus a chain.** A single call is faster and simpler to operate; a chain gives each step a narrow job and a place to check it, at the price of latency, more moving parts and errors that travel between stages. - **Fixed versus retrieved examples.** A fixed block is predictable and friendly to prompt caching; picking examples per request fits the input better but makes behaviour depend on the pool's contents and freshness. - **Reasoning versus a direct answer.** Asking for reasoning helps multi-step problems and costs tokens and latency everywhere else, so the decision belongs to a segment of traffic, not to the whole system. - **Prompt versus code.** Wording lowers a failure rate; a validator, a tool check or constrained decoding bounds it. The more a failure costs, the further its guard belongs from the prompt. - **Strict versus expressive output.** A tight schema makes output easy to consume; too tight, and it can force the model to invent a value or choose a wrong option rather than say it does not know.

Several shapes recur across the sections under different names; spotting one is often the quickest way into an unfamiliar question. - **Separate material from instructions.** Delimiters around documents, a fixed prompt skeleton, and tool or retrieved content treated as data rather than orders are the same idea at the text, template and role level. - **Check every seam.** Between chain stages, after structured output, and at the entrance to an example pool: validate what crosses the boundary and decide in advance what happens when it fails. - **Restate near the point of generation.** Rules that fade over a long conversation or a long output are repeated close to where the model is writing, not only at the top. - **Escalate only on measured failure.** Instruction alone, then examples, then reasoning, then sampling and voting: each rung is climbed when the cheaper one is shown to fall short. - **Change one variable, rescore the same set.** Ablations, shot-count sweeps, order searches and chain-versus-single-call comparisons are one experimental design pointed at different variables.

explore

report an issue with this guide →

questions

page 1 of 2

Why must few-shot examples use the same field names and delimiters as the live query?

level: juniorimportance: must knowfreq 65%

answer

  1. the prompt is one continuous sequence
  2. the pattern is the instruction
  3. separator drift between examples and query
  4. stop the prompt mid-record
  5. render both sides from one template

basics

~20 s

Demonstrations teach by pattern continuation. If the examples use one separator and field set but the live query arrives in a different shape, the query stops reading as the next item in the series, and the model's output format drifts or it keeps writing examples.

solid answer

~50 s

A few-shot prompt is one long sequence, and the model's job is to continue it. The examples establish a template — field names, their order, the separator between items, the casing and spacing — and the live query is supposed to be an incomplete instance of that same template, stopped right where the answer should begin. Consistency is what makes that work. If your wildlife-sighting exemplars are separated by `###` with `Report:` / `Species:` / `Confidence:` fields, but the live report is wrapped in `---` and labelled `Observation:`, the model sees a *different* pattern starting and may invent its own field names, answer in prose, or continue generating more demonstrations instead of answering. The fix is mechanical: build the prompt from one template function so exemplars and the live query cannot drift apart, and end the query mid-pattern at the field you want filled.

code

markdown · 14 lines
markdown
Report: Two grey wolves crossing the ridge at dusk, clear view
Species: grey wolf
Confidence: high
###
Report: Something large moved in the brush, never got a look
Species: unknown
Confidence: low
###
Report: Small deer-like animal at the salt lick, seen briefly
Species: roe deer
Confidence: medium
###
Report: {{live report text}}
Species:

go deeper

for a junior

Be able to say that the model continues a pattern, so examples and the live query must look like items in the same series: same field names, same separator, same spacing.

for a middle

Explain the specific failure modes — invented field names, dropped fields, the model continuing the example series — and the trick of ending the prompt mid-record at the field you want filled.

for a senior

Show how you keep the format from drifting in a real codebase: exemplars stored as structured records, one renderer for exemplars and live query, and a check that catches divergence before it ships.

for a principal

Be clear that demonstration formatting is a probabilistic nudge, not a guarantee, and own the decision about where the real parsing contract lives and what happens on the outputs that still come back malformed.

## The prompt is one sequence There is no structural boundary in a prompt between "your examples" and "the real question". The model receives a single token stream and predicts what comes next. Everything few-shot prompting does rests on that: the examples set up a visible, repeating pattern, and the live query is arranged so that the single most natural continuation is the answer you want. This is why formatting consistency is not cosmetic. The pattern *is* the instruction. Break it and you have quietly asked a different question. ## What has to stay identical Four things drift in practice, and each has a characteristic failure: - **Field names.** `Report:` in the examples and `Observation:` in the query. The model no longer recognises the query as an instance of the demonstrated task, and often re-labels the output with field names it invented. - **Field order and set.** Examples with three fields, query with two, or the same fields in a different order. The model tends to fill in what it saw, producing fields you did not ask for or omitting the one you did. - **The separator between items.** Exemplars split by `###` while the live query is wrapped in `---`. This is the classic case: `---` reads as the start of a *new* section rather than the next item in the list, and the model frequently continues the example series — inventing another fake report and labelling it — instead of answering yours. - **Whitespace and casing.** `Species: grey wolf` in one exemplar, `species:grey wolf` in another. Inconsistency here weakens the pattern and shows up as unstable output casing, which then breaks whatever parses the response. ## Stop mid-pattern The second half of the technique is where the prompt ends. If each exemplar is `Report: … / Species: … / Confidence: …`, the live query should end with `Report: <text>` and then a bare `Species:` — with nothing after it. The model is now completing a partially written record, which is the easiest possible continuation and leaves almost no room to add a preamble like "Sure, here's the classification". Ending the prompt after the input text without the trailing field name invites exactly that preamble, because there is no half-finished line to complete. ## Why it drifts in real systems Almost nobody writes an inconsistent prompt on purpose. It happens because the exemplars live in one place — a fixture file, a spreadsheet, a constant at the top of a module — and the live query is assembled somewhere else, in application code, by a different person, months later. Someone adds a field to the query builder for a new use case and does not update the fixtures. Someone reformats the examples file and a trailing newline changes. A migration rewrites the separator. The durable fix is structural rather than a matter of care: render both the exemplars and the live query through **one template function**, so a change to the shape is applied to every instance by construction. If exemplars are stored as structured records rather than as pre-rendered strings, they cannot fall out of sync with the query at all — the renderer is the single source of the format. ## The distinction worth naming in an interview A sharp answer separates two concerns that often get conflated. Formatting the demonstrations is about **teaching a shape by example** — making the model's most likely continuation be the answer you want, in the layout you want. It is a soft, probabilistic mechanism. It is not the same as *enforcing* a shape at generation time, which is a decode-side concern with different tools and different guarantees. Good demonstration format raises the odds that a parse succeeds and makes the output stable enough to read; it does not make malformed output impossible. ## Prose versus fields One more question follows naturally: should the examples be strictly fielded at all, or is natural prose fine? Fielded exemplars — explicit `Field:` labels on separate lines — give the model an unambiguous slot to fill and give your parser an unambiguous thing to read. Prose exemplars ("the report describes two grey wolves, so this is a confident grey wolf sighting") are more natural and can suit open-ended tasks, but the boundary between input and answer becomes fuzzy, the model's stopping point becomes unpredictable, and downstream extraction turns into regex archaeology. For anything whose output is consumed by code rather than a human, fielded and consistent wins.

  • How should the prompt end so the model does not add a preamble before the answer?
    End mid-record, at the field name you want filled. If exemplars are Report / Species / Confidence, the live query should end with the report text and a bare `Species:` on its own line. The model is then completing a half-written record, which is the easiest continuation and leaves no natural place for "Sure, here's the classification".
  • Is it better to write demonstrations as labelled fields or as natural prose?
    Fielded, whenever code consumes the output. Explicit field names give the model an unambiguous slot to fill and give your parser an unambiguous thing to read, and they make the stopping point predictable. Prose exemplars suit open-ended tasks read by humans, but they blur the input-answer boundary and turn extraction into fragile pattern matching.
  • Why does this consistency decay in production prompts, and how do you prevent it?
    Because exemplars usually live in a fixture or constant while the live query is assembled in application code, by a different person, later. Someone adds a field or reformats one side. Prevent it structurally: store exemplars as structured records and render both them and the live query through a single template function, so the shape cannot diverge by construction.

saying these in an interview costs you the question

  • Treats separators and field names as cosmetic polish
  • Believes the model knows where the examples end and the query starts
  • Ends the prompt after the input without the trailing field name
  • Hand-writes the query format separately from the exemplar format
  • Confuses teaching a shape by example with enforcing it at generation time

context

open as a page

What is dynamic exemplar retrieval in few-shot prompting, and how does it differ from a fixed example block?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Dynamic exemplar retrieval picks the few-shot demonstrations per request, pulling the most similar labelled examples from a pool by vector similarity. A fixed block hardcodes the same demonstrations into the prompt for every request, whatever the input looks like.

open as a page

When picking few-shot exemplars, why must every label class appear?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Demonstrations define the label space the model treats as live. A class with no example is predicted far less often than it should be, even when the instruction names it, so coverage of every label comes before adding more examples.

open as a page

What is prompt chaining, and why split a task across several model calls?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Prompt chaining runs one task as an ordered series of model calls, each stage's output feeding the next stage's prompt. Every stage gets a narrow instruction and a checkable result, so a single step can be validated, retried, or replaced.

open as a page

When do you escalate from a direct answer to chain-of-thought or self-consistency?

level: middleimportance: must knowfreq 72%

basics

~20 s

Escalate only when the cheaper rung measurably fails. Direct answers suit lookup and formatting; chain-of-thought pays on multi-step reasoning; sampling several chains and voting pays only when the answer is a short discrete value worth the multiplied cost.

open as a page

Why hold out a separate test set when iterating on a prompt by hand?

level: middleimportance: must knowfreq 62%

basics

~20 s

Repeated hand-tuning fits a prompt to the examples you keep looking at. A held-out test set, scored rarely and never read during tuning, is the only honest estimate of how the prompt behaves on inputs you have not seen.

open as a page

In few-shot prompting, do wrong labels in the demonstrations hurt accuracy?

level: middleimportance: must knowfreq 55%

basics

~20 s

Wrong labels hurt far less than most people expect. On classification tasks, swapping gold labels for random ones drawn from the same label set often costs only a few points, while destroying the input-label format costs much more.

open as a page

Why can permuting the same four few-shot examples swing accuracy by double digits?

level: middleimportance: must knowfreq 62%

basics

~20 s

Order is not cosmetic: the model reads the demonstration sequence as evidence about which label is likely, and examples nearest the end pull hardest. The same four exemplars in different orders can move a sentiment task from the fifties into the high eighties.

open as a page

How many few-shot exemplars are enough, and how do you find that point?

level: middleimportance: must knowfreq 56%

basics

~20 s

Find it empirically: sweep the shot count on a held-out set and ship the first point where added examples stop paying. Gains typically flatten after a handful, while every extra example keeps costing tokens and latency on every request.

open as a page

Why delimit untrusted document text in a prompt, and why is that not a security control?

level: middleimportance: must knowfreq 72%

basics

~20 s

Delimiting untrusted text tells the model which span is material to work on rather than instructions to follow, which removes ambiguity and makes prompts assemblable by code. It enforces nothing, though: the model still reads one undifferentiated token stream.

open as a page

In an LLM prompt, why is "respond only with JSON" unreliable, and what beats it?

level: middleimportance: must knowfreq 78%

basics

~20 s

A format instruction only biases the model; every token stays available, so it can still emit a preamble, a ```json fence, or a truncated object. Constraining decoding — a schema-based structured-output mode or a grammar — makes invalid output impossible rather than unlikely.

open as a page

When an LLM's JSON fails schema validation, how should the calling code respond?

level: middleimportance: must knowfreq 63%

basics

~20 s

Fail closed: reject the response, never coerce or scrape it. Then optionally send one targeted repair prompt quoting the exact validator error, cap retries at one or two, and fall back to a deterministic path or human review while logging the failure rate.

open as a page

In an LLM chat request, which instruction layer wins: system, user, or tool output?

level: middleimportance: must knowfreq 66%

basics

~20 s

Higher layers win. Provider policy outranks the application's system prompt, which outranks the end user's turn. Tool results and retrieved documents carry no instruction authority at all — they are data the model reasons about, not orders it follows.

open as a page

Why does system-prompt adherence decay over a long conversation?

level: middleimportance: must knowfreq 62%

basics

~20 s

Adherence decays because the standing rule sits at the far end of a growing history: it competes with thousands of tokens of newer, more concrete dialogue, and the model's own earlier replies — including any that already broke the rule — become the strongest example of how this conversation behaves.

open as a page

What must a prompt regression suite pin so a score change is attributable?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Pin everything except the prompt: the eval dataset version, the model version, decoding parameters, and — if scoring is automated by a model — the judge and its rubric. Anything unpinned becomes an alternative explanation for every score move.

open as a page

Why does a per-request retrieved exemplar block destroy prompt-cache hits?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Prompt caching matches an exact token prefix from the start of the request. Retrieved exemplars differ on every call, so everything from that block onward is uncached — placed near the top, the entire prompt is reprocessed every single request.

open as a page

A four-stage screening chain extracts fields then judges inclusion; how do you stop stage-2 errors poisoning stage 3?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Put a check at the seam. Validate the extracted fields against a schema, require each value to appear verbatim in the source, allow the stage to abstain rather than guess, and pass the source forward so the judging stage can disagree with the extraction.

open as a page

In an LLM support agent, why can't a system prompt enforce a $100 refund cap?

level: seniorimportance: must knowfreq 72%

basics

~20 s

A system prompt is guidance, not code. Nothing checks it at runtime, so a $100 cap written in prose holds only as often as the model chooses to hold it. The cap must be enforced in the refund tool or the payments service, which reject larger amounts regardless of what the model asks for.

open as a page

How do you re-anchor a drifting system-prompt constraint mid-conversation?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Restate the rule close to where generation happens — a short constraint block appended after the newest user message — rather than rewriting the system prompt. Re-anchor only the two or three rules that actually drift, keep the wording byte-identical every turn, and enforce anything business-critical with a programmatic check instead.

open as a page

How do you keep chain-of-thought reasoning out of the user-facing answer?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Ask for reasoning inside one delimiter and the final answer inside another, then strip the reasoning before display. It is still generated and can be logged for review — hiding it is a rendering step, not a prompt shortcut.

open as a page

In a prompt, when should you use XML-style tags instead of triple-backtick fences?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Named XML-style tags suit prompts holding several distinct or nested blocks: each block gets a name and an explicit close marker you can refer to. Triple-backtick fences suit short literal or code spans. Consistency matters more than the vocabulary.

open as a page

When should an LLM return Markdown instead of JSON, and why not mix both?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Ask for Markdown when a human reads the output directly, and JSON when code consumes it. Mixing them breaks parsing, since prose around an object stops the whole response from parsing. Put human-facing text inside a JSON field instead.

open as a page

How would you ablate a 900-token prompt to find its load-bearing instructions?

level: middleimportance: should knowfreq 42%

basics

~20 s

Remove one section at a time and rescore the identical eval items. Sections whose removal barely moves the score are not earning their tokens; sections that cause a clear drop are load-bearing. Then rebuild additively to catch redundant pairs.

open as a page

Your few-shot examples show a bare label but the task asks for a label plus a reason. What breaks?

level: middleimportance: should knowfreq 45%

basics

~20 s

The demonstrated shape usually wins over the written instruction. The model emits bare labels and drops the reason, or produces an unstable mix across calls. Fix the exemplars so every one carries both fields, in the order you want them produced.

open as a page

How can nearest-neighbour exemplar retrieval hurt accuracy on a borderline query?

level: middleimportance: should knowfreq 45%

basics

~20 s

Nearest neighbours are similar to the query, not representative of the task. For a borderline case they often all carry one label, or duplicate each other, so the prompt quietly argues for that one answer instead of showing the model where the decision boundary sits.

open as a page

How does contextual calibration use a content-free input like "N/A" to correct label bias?

level: middleimportance: should knowfreq 30%

basics

~20 s

Feed the prompt an input carrying no task evidence, such as "N/A" or an empty string, and read the label probabilities it returns. Those probabilities are the prompt's built-in prior; dividing real predictions by them and renormalising removes most of the bias.

open as a page

Should a fixed few-shot exemplar set favor diversity or typical cases?

level: middleimportance: should knowfreq 47%

basics

~20 s

Both, in a specific order: cover the distinct modes of the traffic you actually receive, then weight roughly toward the common ones. Cloning one typical case wastes tokens; filling the set with exotic cases teaches a distribution your users do not send.

open as a page

What is least-to-most prompting, and how does it differ from a fixed pipeline?

level: middleimportance: should knowfreq 45%

basics

~20 s

Least-to-most prompting asks the model to first list the subproblems a hard question contains, then solve them in dependency order, each answer feeding the next. The decomposition is produced at run time by the model, not authored in advance by an engineer.

open as a page

A system prompt says "always cite sources" and the user says "no citations" — what should the assistant do?

level: middleimportance: should knowfreq 48%

basics

~20 s

It depends on why the constraint exists. For a stylistic preference, satisfy the user's intent as far as the constraint allows and say what was kept. For a constraint protecting someone other than the user, keep it and refuse the removal — briefly, without lecturing.

open as a page

Chain-of-thought on every request costs tokens — how do you decide which ones need it?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Segment the traffic. Reason only on the slice whose errors are genuinely multi-step; answer directly on lookup and formatting work, where forced deliberation adds latency and can talk the model out of a correct answer. Gate on a cheap difficulty signal and measure each segment separately.

open as a page

showing 1–30 of 49