Zero-shot or few-shot CoT for a bank's KYC risk memos — how do you choose?
answer
- billed on every request, not once
- PII lives inside the prompt
- synthetic exemplars from written policy
- template gives format without traces
- measure before paying for exemplars
basics
~20 sCost and data exposure decide it. Exemplars are re-sent and billed on every request, and exemplars written from real customer files put personal data into every call and every log. Prefer synthetic policy-derived examples, or zero-shot with a strict output template.
solid answer
~50 sTreat this as two separate decisions that people usually merge. First, **does the demonstrated reasoning earn its tokens?** Exemplars are part of the input on every single request, so on a high-volume path their cost is recurring; run both variants on a real eval slice and only pay if the quality gap is real. Prefix caching lowers the price of repeated tokens but does not remove them. Second, **what is inside the exemplars?** Hand-written KYC examples drawn from real files carry customer names, identifiers and adverse-media detail into every request, every provider log, every trace and every retained prompt — which pulls the prompt itself into the bank's data-handling and retention scope. The resolution most teams land on is synthetic exemplars derived from the written policy rather than from customer records, kept to the minimum that moves the metric, plus a strict memo template. If measurement shows the exemplars are not moving quality, zero-shot with that template is both cheaper and lower risk.
go deeper
Know that examples in a prompt are sent with every request — so they cost tokens each time, and anything sensitive written into them is transmitted each time too.
Be able to lay out the tradeoff explicitly: recurring input-token cost and context occupancy against format consistency and transferred method, and note that a template can buy format without full worked traces.
Show the measurement: an eval on historical cases comparing both variants on determination correctness, check coverage and format conformance, plus the recognition that exemplars written from real files put personal data in every call and log.
Own the standing decision — synthetic policy-derived exemplars, a named owner, version control and change review for the exemplar block, human review on adverse determinations, and a periodic re-test because a model upgrade can invalidate the original comparison.
## Why this is a judgment question There is no universally right answer, and an interviewer asking it is checking whether you can separate the strands: reasoning quality, recurring cost, output consistency, and regulatory data handling. A candidate who answers only on accuracy has missed most of the question. ## Strand one: does the demonstrated reasoning actually help? Risk-memo generation is a house-procedure task — a specific ordering of checks, specific thresholds, a specific way findings are worded. That is exactly the profile where worked examples usually help, because a trigger phrase cannot convey a bank's internal method. But "usually helps" is a hypothesis to test, not a conclusion. Build an eval set from real historical cases with analyst-written outcomes, run a zero-shot variant and a few-shot variant, and compare on the dimensions that matter separately: correctness of the risk determination, coverage of required checks, and format conformance. Frequently the exemplars turn out to be carrying format rather than method — in which case a written output template achieves the same result for a fraction of the tokens and none of the data risk. ## Strand two: the token economics Exemplar tokens are input tokens on every call. A prompt carrying four full worked memos is not a one-time setup cost; it is a per-request cost multiplied by volume for as long as the prompt lives. On a batch path processing overnight review queues, that is a budget line someone will eventually ask about. Prefix caching changes the arithmetic — a stable exemplar block shared across requests can be billed at a reduced rate after the first call — but it does not make the tokens free, does not help when the prompt prefix changes often, and does nothing at all about what the exemplars contain. Treat caching as a reason cost stops being the *deciding* factor, not as a reason to stop counting. There is a second, less visible cost: context occupancy. Every token of exemplar is a token unavailable for the case file, the policy extract, or the prior-review history. On long inputs that tradeoff bites before the invoice does. ## Strand three: what the exemplars contain This is the strand that makes the KYC setting different from a generic prompt-tuning exercise. The instinctive way to write a good exemplar is to take a real, well-handled case and write up its reasoning. That decision quietly means: - Customer personal data is transmitted to the model provider on **every** request, not only on requests about that customer. - It appears in provider-side logs, in your own request tracing, in eval fixtures, in prompt-version history in source control, and in any debugging capture. - The subject of those exemplars has no relationship to the request being processed, so there is no processing basis tied to the request itself. - Data-subject rights and retention rules now apply to a prompt template, which is not where anyone expects to look for personal data. The practical answer is **synthetic exemplars derived from the policy** — invented customers, invented identifiers, reasoning written to demonstrate the escalation logic in the written procedure. They are usually as effective, because what you needed to transfer was method, not the specific customer. Where realism genuinely matters, aggressive redaction plus a review step is the fallback, but redaction is a control that must be verified rather than assumed. ## The recommendation and how to defend it A defensible position for a regulated memo generator: start with a strict output template and zero-shot reasoning; measure; add the smallest number of *synthetic* worked examples that closes a measured gap; re-measure; and treat the exemplar block as reviewed, version-controlled artifact with a named owner, because it now encodes bank policy. Never put real customer records in it. Say the quiet part out loud in the interview: format consistency is not optional here — a memo that a reviewer, an auditor and a downstream system all have to read needs a fixed shape — but format consistency can be bought with a template, and it does not require handing the model somebody's file. ## What else this touches Auditability: if a determination is challenged, you need to be able to say what the prompt contained at that moment, which means versioning the exemplar block alongside the code. Change control: editing an exemplar changes decision behaviour and should go through the same review as changing the policy it encodes. And human review: for adverse determinations, the model's memo is an input to an analyst, not the decision itself — which is the control that makes the rest of the tradeoffs survivable.
- What if measurement shows the exemplars barely move quality?Drop them. That is the cheapest and lowest-risk outcome, and the one people resist because the exemplars took effort to write. Keep the output template, which is usually where most of the observed benefit was coming from, and re-run the comparison periodically — a model upgrade can change the answer in either direction, sometimes making a prompt that needed demonstrations no longer need them.
- How does prompt prefix caching change the cost side of this decision?It amortizes a stable exemplar block across requests at a reduced rate, so raw token cost often stops being the deciding factor. What it does not change: the exemplars still occupy context that could hold case material, they still travel to the provider on every call, and anything sensitive inside them is still transmitted and logged. Caching is a cost lever, not a privacy control.
- Are synthetic exemplars really as good as ones written from real cases?For transferring method, usually yes — what the model needs is the order of checks, the thresholds and the wording conventions, all of which come from the written procedure rather than from any particular customer. Where realism matters is surface texture: messy names, inconsistent document formats, partial records. You can reproduce that synthetically on purpose, and doing so is far cheaper than defending real records inside a prompt template.
- Who should own the exemplar block once it is in production?Whoever owns the policy it encodes, jointly with engineering. The exemplars are not prompt decoration — they define how borderline cases get reasoned about, so editing one changes decision behaviour. Version them with the code, put changes through the same review as a policy change, and keep the history so that a challenged determination can be reconstructed against the prompt that actually produced it.
saying these in an interview costs you the question
- Treats exemplar tokens as a one-time setup cost
- Says caching makes repeated exemplar tokens effectively free
- Puts real customer records into a prompt template for realism
- Chooses few-shot on instinct without measuring the quality gap
- Ignores that prompt content lands in provider and tracing logs