In a billing pipeline, when is a Monte Carlo algorithm's error probability unacceptable but a Las Vegas one's variable runtime fine?
answer
- two contracts, two things given up
- which is fixed: the answer or the cost?
- always right, unpredictable finish time
- bounded time, small chance of wrong
- which one can an auditor sign off?
basics
~20 sA Las Vegas algorithm is always correct with variable runtime; a Monte Carlo one has bounded cost and a chance of being wrong. Figures that must reconcile exactly can absorb a late finish, never a wrong number.
solid answer
~50 sThe two contracts give away different things. Las Vegas fixes correctness and lets runtime vary — a draw that shuffles then verifies, or a pass with randomly drawn pivots, always returns the right answer and only its finish time is a random variable. Monte Carlo fixes the budget and lets correctness be probabilistic: it finishes on a schedule you can promise, and is occasionally wrong. For an invoice line a customer disputes and an auditor reconciles, a one-in-a-million wrong total is not a rounding concern, it is a defect with a paper trail — so the Monte Carlo contract is off the table no matter how tight its bound. A variable finish time, by contrast, is an ordinary operations problem with ordinary answers: a wider window, a queue, a deadline with retry. The rule of thumb: buy Monte Carlo where the output is an estimate someone will act on statistically, and Las Vegas where the output is a number someone will be held to.
go deeper
Learn which side of each contract is guaranteed: Las Vegas guarantees the answer, Monte Carlo guarantees the cost. Be able to name one everyday example of each without mixing them up.
Explain the conversion between them. Verify-and-retry turns a probabilistic answer into a certain one at the price of a variable finish time, and a deadline turns a certain answer into a probabilistic one.
Bring the operational consequences: a variable runtime surfaces as tail latency and retries, an error probability surfaces as an incident nobody can reproduce. Say which of those your on-call rotation can actually handle.
Own the contract as a business commitment rather than an implementation detail. Decide where probabilistic correctness is permitted at all, write it down per pipeline, and get finance, audit and support to agree before the code ships.
## Two contracts, named after casinos A randomized algorithm has to give something up, and there are two choices about what. **Las Vegas**: the answer is always correct, and the *runtime* is a random variable. You may wait longer than expected, but what you get is right. A shuffle-then-verify draw is one — reshuffle until the arrangement satisfies whatever constraint you require, and you leave with a valid arrangement, eventually. So is a pass that draws its pivots at random: correct on every run, with a finish time that varies. **Monte Carlo**: the runtime is bounded and the *answer* is a random variable. You get a result on schedule, and with some probability it is wrong — either wrong in the sense of an estimate that is off, or wrong in the sense of a yes/no answer that is occasionally mistaken. A sampled count over a stream is the everyday instance. Both are respectable. Neither is a weaker version of the other. What decides between them is not elegance but what happens downstream when the guarantee is exercised. ## The conversion in both directions The contracts are not far apart, which is why they get confused. Given a **cheap way to verify** an answer, any Monte Carlo algorithm becomes a Las Vegas one: run it, check the result, repeat until the check passes. Correctness becomes certain and the uncertainty relocates into how many attempts you need. This is why verifiability is the pivotal question in the design review — if checking is as expensive as computing, the conversion is not available. Going the other way is trivial and is done accidentally all the time: cap a Las Vegas run at a deadline and return whatever you have, or a default. You have just bought a Monte Carlo contract, whether or not anyone wrote that down. A timeout on an "always correct" routine silently converts correctness into a probability — a conversion that belongs in the design document, not in an operations config nobody reads. Note that lowering an error probability, by more sampling or by running several independent attempts and taking a majority, never reaches zero. Small is not none, and the difference matters exactly when the output has legal or financial weight. ## Choosing by consequence, not by taste The deciding question is never "how small can we make the error?" It is **what happens the first time the guarantee is exercised.** Where the output is money someone is held to — an invoice total, a payout, a tax figure, a settlement — a wrong answer is not a small deviation. It is a customer dispute, an auditor's finding, a correction with a paper trail, and a support cost far exceeding whatever the fixed runtime saved. And it is worse than an ordinary bug because it is *by design* unreproducible: the same input can produce the correct answer when you go looking. On-call cannot debug a contract. Where the output is already an estimate — a dashboard's unique-visitor count, a sampled log query, a capacity forecast, a health metric — the Monte Carlo contract is the honest one. Insisting on exactness there buys precision nobody consumes at a cost everybody pays, and it is the more common failure in practice: exact answers computed over full data because "exact is safer", at ten times the infrastructure bill, for a number that gets rounded before anyone reads it. Meanwhile a variable runtime, the thing Las Vegas gives up, is a problem operations solves for a living. It becomes tail latency, a queue, a batch window that occasionally runs long, a retry. Those are budgeted, alerted on and absorbed. That asymmetry — one failure mode lands on people who know how to handle it, the other lands on finance and support — is usually the whole argument. ## Owning the decision The call is a commitment, not an implementation detail, and it belongs where consequences land. Say it in the language of a service level: *"always exact, finishes between four and eleven minutes"* versus *"always eleven minutes, and roughly one run in a million is off"*. Put both to the owner of the consequence. Do not bury either one in a code comment. Three practices follow. Write down which contract each pipeline holds, so nobody adds a timeout to a Las Vegas job without noticing what they changed. Record the seed with each run, so a Las Vegas result can be replayed and a Monte Carlo result explained. And revisit the choice when the output's use changes — a sampled metric that starts feeding a billing rule has silently changed contracts, and that migration, not the algorithm, is where the incident comes from.
- Can a Monte Carlo algorithm be turned into a Las Vegas one?Yes, whenever the answer is cheap to verify: run it, check, repeat until the check passes. Correctness becomes certain and the uncertainty moves into the attempt count. The conversion the other way is trivial and often accidental — cap a Las Vegas run with a deadline and return whatever you have, and you have bought an error probability without writing it down.
- Where in the same company would you happily accept the Monte Carlo contract?Anywhere the output is already an estimate: a dashboard's unique-visitor count, a sampled log query, a capacity forecast, a health signal. Match the contract to how the number is used, not to which guarantee sounds stronger. Computing exact answers over full data for a figure that gets rounded before anyone reads it is the more expensive mistake.
- How do you present this choice to a non-technical owner?As a service level, with both risks stated: "always exact, finishes in four to eleven minutes" against "always eleven minutes, and about one run in a million is wrong". Then let whoever owns the consequence pick. The failure mode to avoid is choosing on the engineer's taste and describing only the option you preferred.
One courier always delivers the right parcel but cannot promise the hour; the other always arrives at nine, and very occasionally with the wrong parcel. Which you hire depends entirely on what is in the box.
saying these in an interview costs you the question
- Las Vegas and Monte Carlo both mean the answer may be wrong
- Monte Carlo just means the algorithm uses random numbers
- A small enough error probability is the same as correct
- A variable runtime is always worse than a small error rate
- Adding a timeout does not change the correctness contract
- Answers can always be verified cheaply, so the distinction never matters