skip to content

How does the benchmark-ips gem's Benchmark.ips measure Ruby code differently from Benchmark.bm, and how do you read its compare! output?

level: middleimportance: must knowfreq 45%

answer

  1. iterations per second, not seconds
  2. warmup 2s, time 5s by default
  3. ± standard deviation per report
  4. compare!: slower vs same-ish
  5. order: :baseline, hold! for separate runs

basics

~20 s

Benchmark.ips runs each block repeatedly for a fixed time after a warm-up (2 s and 5 s by default) and reports iterations per second with a standard deviation. compare! ranks the reports as "Nx slower" or "same-ish: difference falls within error".

solid answer

~40 s

`Benchmark.bm` runs a block once and prints seconds, so you must guess an iteration count. `Benchmark.ips` from the benchmark-ips gem (`require "benchmark/ips"`) first **warms up** each report (2 seconds by default) to estimate how many iterations fit in 100 ms, then runs it for a fixed **time** (5 seconds) and reports **iterations per second** with a ± **standard deviation** and time per iteration. `x.compare!` then prints each report against the fastest as `1.13x slower`, or `same-ish: difference falls within error` when the ranges overlap; `x.compare!(order: :baseline)` compares against the first report instead. `x.config(warmup:, time:, iterations:)` changes the phases, the `|times|` block form removes per-call overhead for tiny code, and `x.hold!("file")` runs each report in a separate Ruby process.

code

ruby · 10 lines
ruby
require "benchmark/ips"

ids = (1..1_000).to_a

Benchmark.ips do |x|
  x.config(warmup: 2, time: 5)
  x.report("select.map") { ids.select(&:even?).map { it * 2 } }
  x.report("filter_map") { ids.filter_map { it * 2 if it.even? } }
  x.compare!(order: :baseline)
end

go deeper

for a junior

Recall that Benchmark.ips reports iterations per second after a warm-up and that compare! ranks the reports.

for a middle

Explain the warm-up and calculation phases, the ± deviation, same-ish versus Nx slower, and order: :baseline.

for a senior

Show you use hold! or separate runs to isolate reports and refuse to report differences inside the error range.

for a principal

Decide which micro-benchmarks belong in review evidence and require them alongside an end-to-end measurement.

## Why iterations per second The standard `Benchmark.bm` runs each block **once** and prints seconds. For small pieces of code that is noise, so people wrap blocks in `100_000.times`, and then argue about whether the count was right. The **benchmark-ips** gem inverts the question: instead of "how long do N runs take?", it asks "how many runs fit in a fixed amount of time?". ```ruby require "benchmark/ips" Benchmark.ips do |x| x.report("interpolate") { "#{1},#{2}" } x.report("join") { [1, 2].join(",") } x.compare! end ``` ## The two phases 1. **Warm-up** (`warmup: 2` seconds by default). Each report runs repeatedly to estimate how many iterations fit into 100 ms; the output shows it as `i/100ms`. This phase also absorbs one-time effects: loaded code, filled caches and, on a JIT, compilation of the hot path. 2. **Calculation** (`time: 5` seconds by default). The report runs in 100 ms cycles for the configured time, and the gem computes **iterations per second**, the **standard deviation** as a percentage, and the time per iteration. `x.config(warmup: 2, time: 5)` changes both; `iterations:` (default 1) repeats the warm-up and calculation stages and keeps the last result, which the README recommends for implementations whose optimiser needs longer to settle. The total run then takes `iterations * (warmup + time)` seconds. ## Reading compare! `x.compare!` adds a Comparison section after the reports: | Output | Meaning | |---|---| | `join: 7.1M i/s` | the fastest report, listed first | | `interpolate: 5.9M i/s - 1.20x slower` | slower by that factor, and outside the error range | | `... - same-ish: difference falls within error` | the difference is inside the measured variation; treat them as equal | With `x.compare!(order: :baseline)` the **first** report is the reference and every other one is compared with it, which suits "old versus new" comparisons. ## Variants worth knowing - **Tiny code**: `x.report("name") { |times| i = 0; while i < times; ...; i += 1; end }` receives the iteration count and loops itself, removing the per-call overhead of the block. - **Separate processes**: `x.hold!("results.json")` runs one report per Ruby invocation and stores results in the file until all have run, so one report's garbage or warmed caches cannot affect another. - **Quick comparisons**: `Benchmark.ips_quick(:upcase, :downcase, on: "hello")` compares methods by name; its README warns it understates differences for code faster than about a million iterations per second. - **GC control**: a custom `suite:` object can run `GC.start` between reports; the README shows one. ## Reading the numbers honestly - A ± of 10% or more means the environment is noisy; rerun before drawing conclusions. - "1.05x slower" with overlapping ranges is not a result. - Micro-benchmarks answer micro-questions: a 2x win on a line that takes 0.1% of a request is invisible in production. ## Summary benchmark-ips replaces guessed iteration counts with a warm-up plus a fixed measuring time, reports iterations per second with variance, and `compare!` tells you whether a difference is real or "same-ish".

  • benchmark-ips says two reports are "same-ish: difference falls within error"; what should you conclude?
    That this run cannot distinguish them: the difference in iterations per second is inside the measured variation. Rerun on a quieter machine or with a longer `time:` if the difference matters; otherwise keep the more readable version and do not claim a speed-up.
  • When would you use the |times| block form of x.report?
    When the code under test is so small that calling the block each iteration is a noticeable share of the time. The block receives the iteration count and loops itself with `while`, so the measurement contains mostly the code, not block-call overhead.

saying these in an interview costs you the question

  • Benchmark.ips runs each block exactly once, like Benchmark.bm
  • same-ish means the first report is slightly faster
  • The warm-up phase results are included in the final i/s figure
  • compare! always compares against the first report
  • A lower i/s figure means the code is faster