skip to content

In a prompt, when should you use XML-style tags instead of triple-backtick fences?

level: juniorimportance: should knowfreq 55%

answer

  1. structure a flat string imposes
  2. named blocks versus literal spans
  3. nesting is the deciding factor
  4. a close marker you can rely on
  5. backticks collide with code content

basics

~20 s

Named XML-style tags suit prompts holding several distinct or nested blocks: each block gets a name and an explicit close marker you can refer to. Triple-backtick fences suit short literal or code spans. Consistency matters more than the vocabulary.

solid answer

~50 s

XML-style tags such as `<contract>…</contract>` give a block two things a fence does not: a **name** the instruction can point at ("summarise the text in `<contract>`") and an unambiguous, explicitly-closed boundary that nests. That makes them the default for long prompts assembled from several parts — rules, examples, retrieved documents, the user's own text. Triple-backtick fences are better for a short literal span where whitespace and exact characters matter, typically code or sample output, and they are what a model is most used to seeing around code. Markdown headings are fine for *your* sections but have no close marker — a section only ends where the next heading starts — so they are a poor wrapper for untrusted content. The real decision drivers are: does the content nest, does it need a name, and could the content itself contain the marker you chose.

code

markdown · 14 lines
markdown
Extract the termination date and notice period from the text in <contract>.
Return one field per line.

<contract>
This agreement terminates on 2027-03-31. Either party may terminate
earlier with 60 days written notice.
</contract>

Use this helper when normalising the date:

```python
def normalise(d: str) -> str:
    return d.strip().replace("/", "-")
```

go deeper

for a junior

Know the three common ways to mark a block — XML-style tags, markdown headings, backtick fences — and be able to say that tags are named, nestable and explicitly closed, which is why they are the usual default for long prompts.

for a middle

Explain the choice by mechanism: nesting, whether the instruction needs to refer to the block by name, and whether the content can contain the marker you chose. Mention that the model is pattern-matching on the markers, not parsing them.

for a senior

Show that you standardise a vocabulary across a codebase rather than picking per prompt, and that you weigh token cost against clarity. Be ready to say what happens when a block sits thousands of tokens from the instruction that named it.

for a principal

Own the argument that delimiter choice is an interface decision for a prompt library — it constrains how prompts are assembled by code, tested and diffed. Be able to justify a house style and the escape hatches from it.

## Why a prompt needs delimiters at all A prompt is one flat string. Whatever structure you believe it has — an instruction section, a retrieved document, a user question — exists only as a convention you impose with punctuation. Nothing in the format marks where your instruction stops and pasted content begins. Delimiters are that convention: a marker that opens a span, ideally a name for what the span is, and a marker that closes it. The payoff is not only that the model can tell the parts apart. It is that **you can name the parts**, and an instruction that says "extract the termination date from the text in `<contract>`" is far less ambiguous than one that says "extract the termination date from the text below" when there are four things below. ## The vocabularies in common use **XML-style tags** — `<contract> … </contract>`. Named, nestable, and explicitly closed. Some providers' prompting guides recommend them specifically for structuring long prompts, and models handle them reliably because tag-shaped markup is abundant in the text they were trained on. Note that the model is not running an XML parser: the tags are a very strong statistical cue, not a grammar, so malformed or unclosed tags degrade gracefully rather than erroring. **Markdown headings** — `## Instructions`, `## Source document`. Highly readable and good for a prompt's own skeleton. Their weakness is that a heading has no closing marker: a section ends only where the next heading begins. That is fine for sections you control and bad as a wrapper around content that may itself contain headings. **Fenced blocks** — three or more backticks, optionally with a language tag. The right choice when the span is a literal: code, a stack trace, sample output, anything where exact characters and whitespace matter. Fences have a defined closing rule but collide with content that contains fences of its own. **Custom sentinels** — `#####BEGIN DATA#####`, or a random string generated per request. Verbose and token-hungry, but effectively collision-proof, which matters when you cannot inspect or modify the content you are wrapping. ## How to choose Three questions decide it: 1. **Does the content nest?** A retrieved document set where each document carries a title and a body wants tags, because tags nest and headings do not. 2. **Does the block need a name you will reference?** If the instruction has to point at the block, tags win — you get a noun for free. 3. **Can the content contain your marker?** Source text pulled from a wiki or a repository will contain backtick fences. A fence is then the worst choice available. A useful default: XML-style tags for structural blocks, fences reserved for literal code spans inside them. ## Practical rules Use lowercase, descriptive tag names and reuse the same names across the prompt library — `<document>` everywhere beats `<doc>`, `<source>` and `<input>` scattered across prompts. Always close what you open, and put the name at both ends so a truncated prompt is still diagnosable. Do not invent a new vocabulary per prompt; the point is a stable skeleton. And keep the delimiters out of the model's output format decision — how the answer is shaped is a separate concern from how the input is fenced. ## What delimiters do not do They do not make the enclosed text inert. The model reads one token stream, and text inside a tag is still text the model attends to; if that text contains something that reads like an instruction, wrapping it changes the odds but enforces nothing. They also do not guarantee parsing: the model can ignore a boundary, especially in a very long prompt where the block is thousands of tokens from the instruction that named it. Treat delimiters as strong, cheap structure that improves clarity and makes prompts assemblable by code — not as a mechanism with a guarantee behind it.

  • Why are markdown headings a weak wrapper for a pasted document?
    A heading opens a section but never closes one — the section ends only where the next heading appears. If the pasted document contains its own headings, the model has no way to tell which heading resumed your prompt and which belonged to the data. Tags close explicitly and nest, so the boundary survives content that looks structurally similar to your own skeleton.
  • Does the model actually parse XML tags in a prompt?
    No. There is no parser in the loop; the tags are tokens like any others. They work because tag-shaped markup is extremely common in training data, so the model has learned strong associations between an opening tag, its content, and the matching close. That is why malformed tags degrade gradually instead of failing, and why a tag thousands of tokens from the instruction that referenced it can still be missed.
  • Is there a cost to heavy tagging?
    Yes, two. Tags consume tokens on every request, which is measurable when a prompt wraps dozens of small blocks. And over-structuring can bury the actual task in ceremony, so an instruction competes with markup for attention. Tag the blocks you will refer to or that hold untrusted content; do not tag every sentence.

saying these in an interview costs you the question

  • Says XML tags are parsed by the model
  • Wraps code containing backticks in a triple-backtick fence
  • Invents a new tag vocabulary in every prompt
  • Uses markdown headings to wrap pasted untrusted content
  • Claims one delimiter style is universally correct

context