A colleague's Ruby micro-benchmark shows a 3x speed-up for their change; what would make you doubt it, and how do you set up a benchmark you can trust?
answer
- warm-up before measuring
- same process, order effects
- unrealistic input sizes
- variance, not a single number
- confirm end to end
basics
~20 sDoubt it without warm-up, with both versions in one process in a fixed order, toy input, no reported variance, or code that is not hot. Warm up, isolate runs, use real data, read the ± and confirm end to end.
solid answer
~40 sA micro-benchmark measures exactly what it ran, which is often not what production runs. I would check: **warm-up** - the first calls pay for loading code, filling caches and, with YJIT, compiling after the call threshold, so an unwarmed version looks slow; **order and shared state** - in one process the second report inherits a grown heap and warmed caches, so benchmark-ips' `hold!` or separate processes help; **input** - ten elements versus the real ten thousand changes which cost dominates; **GC** - a version that allocates more pays for it in collections that may fall outside a short run; **variance** - a ± of 15% makes a 1.2x difference meaningless; and **relevance** - a 3x win on code taking 0.1% of a request changes nothing. Then confirm with the real workload.
go deeper
Recall that a benchmark needs a warm-up and realistic input, and that one number without variance proves little.
Explain warm-up costs, order effects in one process, allocation versus time, and reading the ± column.
Show you challenge a claimed speed-up with specific questions and confirm it on the real workload before accepting it.
Define what evidence a performance change needs before merge, and when a micro-benchmark is acceptable evidence at all.
## Why micro-benchmarks lie A **micro-benchmark** times a small piece of code in isolation. It is the right tool for comparing two implementations of one operation, and a common source of false conclusions. The number is precise; the question it answers is often not the one you asked. ## The checklist 1. **No warm-up.** The first calls to Ruby code pay one-time costs: autoloading constants, populating method and constant caches, growing the object heap, and - when YJIT is enabled - running interpreted until the call threshold and then compiling. A benchmark that times those first calls measures start-up. benchmark-ips's warm-up phase (2 seconds by default) and `Benchmark.bmbm`'s rehearsal exist for this reason. 2. **Order effects in one process.** When version A runs first and B second in the same process, B inherits A's heap size, its garbage waiting to be collected and its warmed caches. Swap the order and the result can flip. Use benchmark-ips's `x.hold!("file")` to run each report in its own Ruby invocation, or run the whole script twice with the order reversed. 3. **Unrealistic input.** Timing on a 10-element Array when production handles 10,000 hides allocation and cache effects; timing on constant input can hide branches the real data takes. 4. **GC outside the window.** A version that allocates twice as much may look equal in a 5-second run and cost more under sustained load, when those allocations turn into extra collections. Count allocations as well as time. 5. **Noise.** A laptop on battery, a busy CI runner or thermal throttling moves numbers by 10% or more. Read the ± column; if it is as big as the difference, you have no result. 6. **Wrong layer.** Even a real 3x speed-up is irrelevant if the code is not on a hot path. A profile of the real workload decides what is worth benchmarking. ## Setting up one you can trust | Step | How in Ruby | |---|---| | Use an iterations-per-second tool | `Benchmark.ips` with `x.compare!` | | Warm up | keep the default warm-up, or lengthen it with `x.config(warmup: ...)` | | Isolate | `x.hold!` or separate processes; reverse the order once | | Match production | real data sizes and shapes; the same Ruby version and the same JIT setting as production | | Measure allocations too | an allocation report or object counts around the code | | Report variance | quote i/s with ±, and accept "same-ish" as an answer | | Confirm end to end | re-measure the request or job the change was for | ## Questions to ask the colleague - Which Ruby version and was YJIT on, and is that what production runs? - How big was the input, and where did it come from? - Did you run it more than once, and in the other order? - What was the ±, and did `compare!` say "slower" or "same-ish"? - Where does this code show up in a profile of the real workload? ## Summary Doubt a micro-benchmark that skipped warm-up, ran versions in one process in one order, used toy input, ignored allocations, hid its variance, or targeted code that is not hot. Trust one that controls those, and then check the real workload anyway.
- Why can enabling YJIT change which of two implementations wins a benchmark?YJIT speeds up Ruby-level code, such as loops and method calls, much more than work already done in C. An implementation that leans on a C-implemented core method may win when interpreted and lose once the Ruby-level alternative is compiled. Benchmark with the same JIT setting production uses, and warm up long enough for compilation to happen.
- How do you check that an order effect is distorting a Ruby benchmark?Run it again with the reports in the opposite order, or use benchmark-ips's `hold!` so each report runs in its own Ruby process. If the ranking or the gap changes, the shared process state, such as heap size or warmed caches, was part of the result.
saying these in an interview costs you the question
- The first run is the most accurate because caches are cold
- Running both versions in one process always gives a fair comparison
- A small input is fine because the speed-up ratio stays the same at scale
- A 3x micro-benchmark win guarantees a faster request
- Variance does not matter as long as the average is lower