skip to content

In Ruby, how do you run code before and after every fork with Process._fork, and why is overriding Kernel#fork not enough?

level: seniorimportance: nice to knowfreq 18%

answer

  1. internal method behind every fork
  2. prepend to Process.singleton_class
  3. super returns 0 in the child
  4. fork, Process.fork, IO.popen with -
  5. Process.daemon bypasses it

basics

~20 s

Process._fork is the internal method that Kernel#fork, Process.fork and IO.popen("-") all call. Prepend a module to Process.singleton_class that overrides _fork, calls super and branches on the result: 0 in the child, the child's pid in the parent.

solid answer

~40 s

`Process._fork` performs the actual fork for `Kernel#fork`, `Process.fork` and `IO.popen("-")`, and exists so that libraries can hook fork events; its documentation says not to call it directly. To hook it, define a module with `def _fork`, run your before-fork code, call `super`, then branch: `super` returns **0 in the child** (unlike `fork`, which returns `nil` there) and the child's pid in the parent; return that Integer. Install it with `Process.singleton_class.prepend(ForkHook)`. Overriding `Kernel#fork` alone misses `Process.fork` and `IO.popen("-")`, and gems cannot rely on everyone using one entry point. The hook is where a child reopens sockets and database connections inherited from the parent and restarts background threads, since only the forking thread survives. `Process.daemon` does not go through `_fork`, so hook it separately if you use it.

go deeper

for a junior

Recall that a forked child inherits open connections and only one thread, so some setup must be redone after fork.

for a middle

Explain which entry points call Process._fork, how to prepend a hook to Process.singleton_class, and why super returns 0 in the child.

for a senior

Write a fork hook that reconnects sockets and restarts threads in the child, handles errors, and accounts for Process.daemon not using it.

for a principal

Decide which layer owns after-fork recovery, the server config, each library, or a shared _fork hook, so no inherited resource is missed.

## The problem hooks solve A forked child inherits the parent's open file descriptors and memory, but only the thread that called `fork`. For a preforking job runner that means: - A **database or cache connection** opened in the parent is one socket now used by several processes; their requests and responses interleave and corrupt the protocol. - A **background thread** (a metrics flusher, a heartbeat) started in the parent does not exist in the workers. - A **monitoring agent** may need to know about the new process to label its data. Each of those needs code that runs around every fork, whichever library or method triggered it. ## Process._fork: one place for every fork `Process._fork` is the core method that performs the fork. `Kernel#fork`, `Process.fork` and `IO.popen` with the `"-"` command all call it, and its documentation describes it as a hook point for application monitoring libraries rather than something to call directly. | Entry point | Goes through `_fork`? | |---|---| | `Kernel#fork` | yes | | `Process.fork` | yes | | `IO.popen("-")` | yes | | `Process.daemon` | no; may use fork(2) internally but bypasses the hook | This is why overriding `Kernel#fork` is not enough: `Process.fork` and `IO.popen("-")` would slip past it, while a hook on `_fork` sees all three. ## Writing a hook ```ruby module ForkHook def _fork JobRunner.before_fork # runs in the parent pid = super if pid.zero? JobRunner.reconnect! # runs in the child JobRunner.start_heartbeat end pid end end Process.singleton_class.prepend(ForkHook) ``` Rules that matter: 1. **Prepend to `Process.singleton_class`.** `_fork` is a singleton method of `Process`, so the module must sit in front of it there. 2. **Branch on 0, not `nil`.** `_fork` returns an Integer in both processes: **0 in the child**, the child's pid in the parent. `Kernel#fork` converts that 0 to `nil` for its own caller. 3. **Return the Integer from `super`.** Returning anything else makes the calling `fork` raise `TypeError`. 4. **Keep failures in mind.** An exception raised by the hook in the child terminates the child; an exception in the parent propagates to whoever called `fork`. ## What to do in each half **Before `super`, in the parent**: stop or pause background threads that must not be copied mid-operation, and flush buffers so data is not written twice by parent and child. **After `super`, in the child**: - **Reopen connections.** Close or discard the inherited sockets and open fresh ones, so parent and child never share a connection. - **Restart threads** the parent was running, because they did not come across. - **Reset per-process state** such as identifiers or counters that should be unique per worker. **After `super`, in the parent**: resume what was paused and record the new pid. ## Checking that the hook works A few quick checks catch most mistakes: - **Fork and compare identities.** In a test, fork with a block, and inside it assert that `Process.pid` differs from the parent's and that the connection object is not the one the parent held. - **Count forks.** Increment a counter in the parent half of the hook, then call `fork`, `Process.fork` and `IO.popen("-")` once each; the count should be three. - **Force a failure.** Make the child half raise and confirm the child exits instead of running work on a half-initialised process. ## How this relates to server hooks Preforking servers expose their own configuration hooks around forking workers. Those run only when that server forks. A `_fork` hook is lower level: it runs for any fork in the process, including one started by a gem you did not write, which is why instrumentation and connection-pool libraries prefer it.

  • Your `_fork` override returns `nil` instead of the value of `super` by mistake. What happens?
    `Kernel#fork` and `Process.fork` expect `_fork` to return an Integer pid, so the call raises `TypeError`. Always end the override with the value returned by `super`, after running your child-side or parent-side code.
  • A worker keeps using the database connection that the parent opened before forking. What goes wrong, and where do you fix it?
    Parent and children now share one socket's file descriptor. Their queries and replies interleave on the same connection, so processes read each other's results or break the protocol. Fix it in the child half of a `_fork` hook, or the library's own after-fork callback: drop the inherited connection without sending a close handshake and open a new one.

saying these in an interview costs you the question

  • Overrides only Kernel#fork and expects Process.fork to be hooked
  • Checks for nil to detect the child inside a _fork override
  • Calls Process._fork directly from application code
  • Assumes Process.daemon triggers the _fork hook
  • Lets forked workers keep using the parent's database socket