skip to content

In CRuby 4.0, when does an &block parameter actually allocate a Proc object, and which uses of it avoid that allocation?

level: seniorimportance: nice to knowfreq 22%

answer

  1. lazy Proc allocation since 2.5
  2. a proxy for block.call (2.6)
  3. forwarding with &block is free
  4. storing or returning it allocates
  5. yield never builds a Proc

basics

~20 s

CRuby creates the Proc lazily: forwarding with other(&block), calling block.call and testing if block run without allocating. Using the parameter as a value, such as storing, returning or passing it positionally, builds the Proc once per call.

solid answer

~50 s

A literal block is passed as a lightweight block handler, not an object. Since Ruby 2.5, declaring `&block` no longer converts it into a `Proc` up front; CRuby waits until the parameter is used as a value. Its compiler recognises cheap shapes and reads the parameter through a proxy instead: forwarding with `other(&block)`, calling `block.call(...)` (optimised since 2.6) and a truthiness test such as `if block`. Anything that needs the object itself, for example `@handler = block`, returning it, passing it as a positional argument, or calling another method on it like `block.arity`, creates the `Proc` then and reuses it for the rest of that call. Making a `Proc` can also move the caller's local variables to the heap so they outlive the frame. In hot paths, prefer `yield`, which never builds a `Proc`, or keep `&block` to the proxy-friendly shapes; capture freely where storing the block is the point.

code

ruby · 17 lines
ruby
class EmailExport
  def each_email(users, &block)
    users.map(&:email).each(&block)   # forwarded: no Proc built
  end

  def on_skip(&handler)
    @on_skip = handler                # stored: Proc built here
    self
  end

  def notify_skip(user)
    @on_skip&.call(user)
  end
end

puts RubyVM::InstructionSequence.of(EmailExport.instance_method(:each_email)).disasm
# look for getblockparamproxy (no allocation) vs getblockparam (allocates)

go deeper

for a junior

Recall that yield runs a block without creating an object, while &block can turn the block into a Proc when you use it as a value.

for a middle

Explain which &block uses stay cheap (forwarding, block.call, if block) and which create a Proc (storing, returning, other methods on it).

for a senior

Show how you would confirm an allocation hotspot with GC.stat or the disassembled bytecode before changing a hot method's block handling.

for a principal

Weigh interpreter-specific micro-optimisations against readable APIs, and set a policy that such changes need a measured hot path to justify them.

## Blocks are not objects until they must be When you write `users.each { |u| ... }`, CRuby does not create a `Proc`. It passes a **block handler**: a pointer to the compiled block plus a reference to the caller's frame. `yield` runs that handler directly, which is why plain blocks are cheap. A **`Proc`** is a heap object. Creating one costs an allocation and, because the proc may outlive the call, CRuby can also have to move the caller's local variables from the stack into a heap-allocated environment. Doing that on every call of a hot method adds garbage-collection pressure. ## Lazy allocation for &block Older Rubies converted the block into a `Proc` as soon as a method declared `&block`. Two releases changed that: - **Ruby 2.5** introduced **lazy Proc allocation** for block parameters: the handler stays a handler until the parameter is actually needed as an object. - **Ruby 2.6** made **`block.call`** on a block parameter fast, without materialising the `Proc`. In CRuby 4.0 the compiler reads the parameter through a special **block-parameter proxy** in these shapes: 1. **Forwarding**: `other_method(&block)` passes the original handler straight through. 2. **Calling**: `block.call(user)`, where `block` is the method's own block parameter, invokes the handler. 3. **Testing**: `if block` or `unless block` checks presence without building an object. ## What forces the allocation Any use that needs the real object creates the `Proc`, stores it in the parameter, and marks it so later uses in the same call reuse it: - assigning it: `@on_skip = block`, `handlers << block`; - returning it or passing it as a **positional** argument: `register(block)`; - calling a method other than `call` on it: `block.arity`, `block.lambda?`, `block[user]`. | Code inside `def run(users, &block)` | New `Proc`? | |---|---| | `users.map(&block)` | no | | `block.call(users.first) if block` | no | | `yield users.first` | no | | `@callback = block` | yes, once per call | | `block.arity` | yes, once per call | Two more cases never allocate a *new* proc: if the caller passed an existing `Proc` with `&`, the handler already is that object; and for `&:email`, CRuby passes the `Symbol` as the block handler and keeps a small cache of symbol procs, so `users.map(&:email)` does not allocate per call. ## Tracing a callback registry Consider an exporter that registers a skip handler once and then processes thousands of users: 1. `export.on_skip { |u| warn "skipped #{u.name}" }` calls `on_skip(&handler)`. 2. `@on_skip = handler` needs the object, so CRuby builds one `Proc` and, if necessary, moves the caller's locals to the heap. 3. Each later `@on_skip&.call(user)` calls an ordinary `Proc` stored in an instance variable; nothing new is allocated for the handler itself. 4. By contrast, a method such as `each_email(users, &block)` that forwards with `.each(&block)` on every call never builds a `Proc` at all. The lesson is that allocation cost follows **how often** the materialising line runs, not whether `&block` appears in the signature. ## Measuring rather than guessing These are **interpreter details**, not language guarantees; they have changed across releases and may change again. To check a hot path: - compare `GC.stat(:total_allocated_objects)` before and after a loop that calls the method many times; - read the bytecode with `RubyVM::InstructionSequence.of(method(:run)).disasm` and look for `getblockparamproxy` (proxy) versus `getblockparam` (materialise). ## Practical guidance 1. Use **`yield`** when the method only runs the block while it executes. 2. Use **`&block`** when you must forward, store or introspect the block; forwarding and calling stay cheap. 3. When the whole point is to **store** the block, such as a callback registry, the allocation is the cost of the feature; it happens once per registration, not per event. 4. Do not rewrite readable code for this without a measurement showing the method is hot.

  • Why is block[user] not treated like block.call(user) for this optimisation?
    The CRuby compiler special-cases only a `call` sent to the method's own block parameter. `[]` is a different method name, so the compiler emits the ordinary parameter read, which materialises the `Proc` first. Both calls do the same thing once the object exists; only `call` skips building it.
  • Does users.map(&:email) allocate a Proc for every call?
    Not normally. CRuby passes the `Symbol` itself as the block handler when `Symbol#to_proc` has not been redefined, and when a proc is needed it comes from a small per-symbol cache in the main Ractor. That keeps the shorthand about as cheap as a literal block.

saying these in an interview costs you the question

  • Declaring &block always allocates a Proc on every call
  • yield allocates a Proc just like block.call
  • Forwarding with other(&block) wraps the block in a second Proc
  • block.arity is as free as block.call
  • These optimisations are guaranteed by the Ruby language specification