skip to content

Profilers & Benchmarks

Benchmark.measure and benchmark-ips time Ruby code, while stackprof, vernier and memory_profiler show where CPU time and allocations go. Interviewers ask how you prove a change is faster.

on this pageshow

explore

questions

5

In Ruby 4.0, how do Benchmark.measure, Benchmark.realtime, Benchmark.bm and Benchmark.bmbm differ, and what changed about loading the library?

level: juniorimportance: must knowfreq 55%

answer

  1. require "benchmark" first
  2. measure returns a Benchmark::Tms
  3. realtime returns elapsed seconds
  4. bm labels rows, bmbm rehearses first
  5. bundled gem since 4.0: Gemfile entry

basics

~20 s

Benchmark.measure returns a Benchmark::Tms of CPU and real time, Benchmark.realtime returns elapsed seconds, bm prints a labelled table and bmbm adds a rehearsal pass. Since Ruby 4.0 benchmark is a bundled gem, so Bundler apps must list it.

solid answer

~40 s

All four come from `require "benchmark"`. `Benchmark.measure { }` returns a `Benchmark::Tms` whose `utime`, `stime`, `total` and `real` separate CPU time from wall-clock time. `Benchmark.realtime { }` returns only the elapsed seconds as a Float, handy for logging one duration. `Benchmark.bm(label_width) { |x| x.report("name") { } }` runs each block once and prints a labelled table under `Benchmark::CAPTION`. `Benchmark.bmbm` runs every block twice: a *rehearsal* pass first, then the measured pass, so one-time costs such as loading code or growing the heap do not land on whichever block runs first. In Ruby 4.0 `benchmark` became a **bundled gem**: plain `ruby` still finds it, but under Bundler `require "benchmark"` fails with `LoadError` unless the Gemfile lists `gem "benchmark"`.

code

ruby · 14 lines
ruby
require "benchmark"

rows = Array.new(100_000) { |i| [i, "name#{i}"] }

t = Benchmark.measure { rows.map { |id, name| "#{id},#{name}" }.join("\n") }
puts format("cpu=%.3fs real=%.3fs", t.total, t.real)

Benchmark.bmbm(8) do |x|
  x.report("map+join") { rows.map { |r| r.join(",") }.join("\n") }
  x.report("buffer") do
    out = +""
    rows.each { |r| out << r[0].to_s << "," << r[1] << "\n" }
  end
end

go deeper

for a junior

Recall require "benchmark", what measure and realtime return, and that bm prints a labelled table.

for a middle

Explain Tms's utime, stime, total and real, why bmbm rehearses first, and the 4.0 bundled-gem change under Bundler.

for a senior

Show you read real versus total to spot waiting, and know when Benchmark's single runs are too noisy to decide anything.

for a principal

Set the team's bar for performance claims: which tool, how many runs, and what variance must be reported with a result.

## The four entry points Ruby's `Benchmark` library answers "how long did this take?". Load it with `require "benchmark"`. | Method | Returns / prints | Use it for | |---|---|---| | `Benchmark.measure { }` | a `Benchmark::Tms` object | one block, with CPU and wall time separated | | `Benchmark.realtime { }` | a Float, elapsed seconds | a single duration to log or compare | | `Benchmark.bm(width) { \|x\| ... }` | prints one row per `x.report` | a quick side-by-side of a few blocks | | `Benchmark.bmbm(width) { \|x\| ... }` | prints a rehearsal table, then the real one | the same, with a warm-up pass | ### Benchmark::Tms `Benchmark.measure` returns a `Benchmark::Tms` with these readers: - `utime` - user CPU time of this process; - `stime` - system CPU time (time the kernel spent on the process's behalf); - `cutime`, `cstime` - the same for finished child processes; - `total` - the sum of the CPU times; - `real` - elapsed wall-clock time. Comparing `total` with `real` is the useful habit: if `real` is much larger, the block spent its time **waiting** (on I/O, a lock, a sleep), not computing, and making the Ruby code faster will not help much. ### bm and bmbm `Benchmark.bm` takes a label width and yields a reporter; each `x.report("label") { ... }` runs its block once and prints a row with user, system, total and real columns under the header `Benchmark::CAPTION`. `Benchmark.bmbm` exists because order matters. The first block often pays one-time costs: autoloading, method caches, the heap growing to its working size. `bmbm` therefore runs the whole set twice - a **rehearsal** pass whose numbers are printed but not meant to be trusted, then the measured pass. It reduces order effects; it does not make a single run statistically meaningful. ## What changed in Ruby 4.0 Ruby 4.0 turned `benchmark` from a default gem into a **bundled gem** (version 0.5.0 ships with 4.0). The practical effect: 1. Running `ruby script.rb` without Bundler: `require "benchmark"` still works, because bundled gems are installed with Ruby. 2. Inside a Bundler-managed app (`bundle exec`, or `Bundler.setup`): only gems in the Gemfile are loadable, so `require "benchmark"` raises `LoadError`, and Ruby prints a hint that benchmark "is not part of the default gems since Ruby 4.0.0" and should be added to the Gemfile or gemspec. 3. Fix: add `gem "benchmark"` to the Gemfile (or the gemspec for a library). The same move happened to `ostruct`, `logger`, `pstore`, `irb` and others in 4.0, and to `csv`, `bigdecimal` and `base64` in 3.4. ## Limits of the Benchmark module - It runs each block **once** per pass, so a block that takes microseconds gives noise; wrap it in `n.times` or use an iterations-per-second tool. - It reports no variance, so two numbers that differ by 5% may not differ at all. - It measures, it does not explain: it tells you *how long*, never *where* the time went. That is a profiler's job. ## Summary `measure` for one block with CPU versus wall time, `realtime` for a bare duration, `bm` for a quick table, `bmbm` to take order effects out of it. On Ruby 4.0 under Bundler, list `benchmark` in the Gemfile first.

  • Benchmark.measure reports total 0.2s and real 3.1s for a block; what does that tell you?
    The block spent about 0.2 seconds on the CPU and the rest waiting: on I/O, a network call, a lock or a sleep. Optimising the Ruby code in it can save at most part of those 0.2 seconds, so look at what it waits for instead.
  • Why does `require "benchmark"` work in `ruby -e` but fail in the same app under `bundle exec` on Ruby 4.0?
    Since 4.0 benchmark is a bundled gem, installed with Ruby but not a default gem. Plain `ruby` can activate any installed gem, while Bundler restricts loading to the gems in the lockfile, so `require "benchmark"` raises `LoadError` until the Gemfile lists `gem "benchmark"`.

saying these in an interview costs you the question

  • Benchmark.realtime returns a Benchmark::Tms object
  • bmbm runs each block once, just with nicer formatting than bm
  • Benchmark is still a default gem in Ruby 4.0, so no Gemfile entry is needed
  • A large gap between real and total means the Ruby code itself is slow
  • One Benchmark.bm run is enough to prove a 5% speed-up
open as a page

How does the benchmark-ips gem's Benchmark.ips measure Ruby code differently from Benchmark.bm, and how do you read its compare! output?

level: middleimportance: must knowfreq 45%

basics

~20 s

Benchmark.ips runs each block repeatedly for a fixed time after a warm-up (2 s and 5 s by default) and reports iterations per second with a standard deviation. compare! ranks the reports as "Nx slower" or "same-ish: difference falls within error".

open as a page

A Ruby CSV export endpoint is slow; how would you use stackprof or vernier to find where the time goes, and which sampling mode would you choose?

level: seniorimportance: must knowfreq 40%

basics

~20 s

Sample a real export with stackprof in :wall mode, its default, to see waits as well as CPU; use :cpu for computation only. vernier adds threads, GVL and GC activity. Look for the widest self-time frames.

open as a page

In Ruby, what does a memory_profiler report show about allocated versus retained objects, and when do you reach for it instead of a sampling profiler?

level: middleimportance: should knowfreq 30%

basics

~20 s

MemoryProfiler.report { code }.pretty_print lists every object the block allocated and those still retained after a forced GC, grouped by gem, file, location and class, with counts and bytes. Use it to find allocation hot spots and leaks, not for CPU time.

open as a page

A colleague's Ruby micro-benchmark shows a 3x speed-up for their change; what would make you doubt it, and how do you set up a benchmark you can trust?

level: seniorimportance: should knowfreq 35%

basics

~20 s

Doubt it without warm-up, with both versions in one process in a fixed order, toy input, no reported variance, or code that is not hot. Warm up, isolate runs, use real data, read the ± and confirm end to end.

open as a page