skip to content

How do you handle source text that contains the same delimiter your prompt uses?

level: seniorimportance: should knowfreq 45%

answer

  1. build-time problem, not prompting problem
  2. content can close your block
  3. escape, extend, or randomise
  4. verbatim content forbids escaping
  5. one function wraps all untrusted text

basics

~20 s

Handle it at interpolation time, before the prompt is built: pick a delimiter the content cannot contain, escape or normalise the offending marker in the span, or extend the fence beyond any run inside. Never concatenate raw content and hope.

solid answer

~50 s

Delimiter collision is a build-time concern, not a prompting concern. Three fixes exist and they compose. **Choose a collision-resistant marker** — a tag name your corpus does not use, or a per-request random sentinel — so the content cannot reproduce it. **Normalise the span** before interpolation: scan for the closing marker, and either escape it, replace it with a visibly-marked substitute, or reject the input, logging every hit. **Length-extend fences**: markdown's fence rule lets you open with more backticks than any run inside the content, so a document containing three-backtick fences can be wrapped in four or five. Which you pick depends on whether you may modify the content — a legal document you must quote verbatim rules out escaping, so you change the wrapper instead. What matters most is that the collision is detected in code and surfaced, rather than silently producing a prompt whose structure ends halfway through the data.

code

python · 12 lines
python
import re

def fence_for(content: str) -> str:
    runs = re.findall(r"`+", content)
    longest = max((len(r) for r in runs), default=0)
    return "`" * max(3, longest + 1)

def wrap_fenced(content: str) -> str:
    f = fence_for(content)
    return f"{f}\n{content}\n{f}"

print(wrap_fenced("see:\n```py\nx = 1\n```"))

go deeper

for a junior

Know that pasted content can contain the very marker you used to wrap it, and that the prompt will not error — it will just be structured wrongly from that point on.

for a middle

Be able to name the three fixes — a collision-resistant marker, escaping or rejecting the span, extending a fence beyond any run inside — and explain the markdown rule that a closing fence must be at least as long as the opening one.

for a senior

Show you make this a property of one interpolation function, with logging and counters on collisions, plus fixtures in the test suite. Explain why verbatim-quoting requirements rule out escaping and push you toward the wrapper instead.

for a principal

Own the trade between collision resistance and prompt stability: random sentinels solve the problem but cost tokens, diffability and prefix reuse. Set the house rule for which content classes get which treatment.

## The failure A pipeline summarises engineering documentation. Each page is pasted into a prompt inside a triple-backtick fence. One page is a tutorial that itself contains fenced code samples. The first inner fence closes the wrapper, and everything after it — half the page, plus your own trailing instructions — lands outside the block the instruction referred to. Output quality drops on exactly the pages that matter most, and nothing errors. The same thing happens with tags: a user profile whose "about" field literally contains `</document>` because the user pasted from somewhere, or wrote it deliberately. The prompt still renders. The structure is simply wrong. ## Three families of fix **1. Choose a marker the content cannot contain.** The cheapest version is a tag name that does not occur in your corpus — `<source_doc_9f2>` rather than `<document>`. The strongest version is a per-request random sentinel generated at build time and referenced in the instruction by that same generated name. Content written before the request cannot predict it. The costs are tokens and readability, and prompts become harder to diff — which matters when the rest of your library depends on a stable skeleton. **2. Normalise the untrusted span.** Before interpolation, scan the string for your closing marker and decide, explicitly, one of three outcomes: **escape** it (replace `</document>` with a marked substitute such as `<\/document>` and say in the instruction that the substitute is a literal), **strip** it, or **reject** the input. Escaping preserves meaning and is usually right for prose; rejecting is right when the content is expected to be machine-generated and a collision indicates a bug or an attempt. Whatever you choose, count it — a metric on collision hits tells you when your delimiter choice has gone stale. **3. Length-extend the fence.** Fenced blocks in CommonMark may be opened with three or more backticks, and the closing fence must be at least as long as the opening one. So a document containing runs of three backticks can be wrapped in four, and one containing four in five. Compute the fence length from the content instead of hard-coding three. This is the only fix that is fully faithful — nothing in the content changes — and it is the right default for code and documentation. ## Choosing between them Ask whether you are permitted to modify the content. A contract clause, an evidence excerpt, a diff under review, an audit log line: these must be quoted verbatim, so escaping distorts them and you must change the wrapper instead — extend the fence, or switch to a sentinel. For free-form user prose, escaping is fine and usually invisible in the result. Ask also how the content arrives. Content you control end-to-end can be normalised once on ingest. Content that arrives per request has to be handled in the interpolation path, which means the interpolation path must be a function, not a formatted string scattered across the codebase. That is the single most important structural consequence: **there should be exactly one place that wraps untrusted text**, and it is the place where this logic lives. ## Catching it before production Collision is trivially testable and rarely tested. Put a fixture in the suite for every wrapper you use — a document containing your closing tag, a document containing a fence, a document containing both — and assert that the assembled prompt still parses under whatever structural check you can apply (count of opening and closing markers, or a regex that the block is well-formed). Add the same fixtures to the eval set so you see the quality effect, not just the structural one. ## Do not fail silently The worst handling is a quiet `replace` buried in a template helper. Someone later debugging "why does the model ignore the last third of long pages" has no signal to follow. Log the collision with the field and the source identifier, emit a counter, and make the escaping visible in whatever prompt-inspection tooling you have. A collision is information: it tells you your delimiter vocabulary no longer fits your corpus.

  • When would you reject the input rather than escape the collision?
    When the content is machine-generated and a collision therefore signals a bug or a deliberate attempt — an API field expected to hold a plain identifier, a schema-constrained payload, an internal service's output. Rejecting turns a silent structural failure into a loud one at the boundary. For human-written prose, rejection is hostile and escaping is the better trade.
  • Why not just always use a random per-request sentinel and stop worrying?
    It solves collision completely but costs elsewhere: extra tokens on every call, prompts that no longer diff cleanly between runs, and a stable prefix that changes each request, which undermines any reuse of an unchanged prefix. Reserve sentinels for spans that must be quoted verbatim and come from genuinely untrusted sources; use fixed, corpus-checked tag names everywhere else.
  • What does the model do when the structure is broken — does it error?
    No. There is no parser, so a stray closing marker produces no error at all; the model simply has weaker evidence about where the block ends and often treats the remainder as a different part of the prompt. The result is a quality regression concentrated on the documents that contain the marker, which is why detection has to happen in your code.

saying these in an interview costs you the question

  • Hard-codes three backticks regardless of content
  • Silently strips the closing marker with no logging
  • Escapes content that must be quoted verbatim
  • Assumes the model errors on a broken block
  • Wraps untrusted text in several places across the codebase

context