skip to content

You inherit a fleet of long-lived Linux servers configured by a large set of Chef cookbooks, maintained by a team fluent in Ruby. How do you decide whether to keep converging with chef-client or move that configuration somewhere else?

level: principalimportance: nice to knowfreq 28%

answer

  1. measure before you decide
  2. zero updated resources means good cookbooks
  3. who can debug compile versus converge?
  4. the code, the tool, and the team are three questions
  5. freeze and shrink beats big-bang rewrite

basics

~20 s

Judge the cookbooks, not the tool's reputation: whether they are declarative and idempotent, whether anyone left can debug a compile-versus-converge bug, and whether the Chef Infra Server plus agent estate is worth operating. Working, quiet cookbooks are a poor migration candidate.

solid answer

~60 s

Start from evidence rather than fashion. Run the fleet and look at the converge reports: if most nodes report zero updated resources and failures are rare, the cookbooks are genuinely declarative and are doing their job — rewriting them buys risk, not correctness. The real costs of staying are specific to Chef's model: recipes are Ruby, so the compile-versus-converge trap and hand-written custom resources need people who actually know it, and the hiring pool for that is thinner than it was. The operational surface is a Chef Infra Server, node objects, and an agent on every host, plus a Policyfile migration if the estate still runs Berkshelf and environment pins. Weigh that against where the configuration *should* live long term — much of what a cookbook does on a long-lived host stops being needed when the host is replaced rather than patched. My usual answer is to freeze cookbook growth, move new configuration to the target model, and let the Chef surface shrink by attrition rather than schedule a big-bang rewrite.

go deeper

for a junior

You will not be asked to make this call, but know that Chef converges long-lived hosts on an interval and that inherited cookbooks often encode fixes nobody remembers the reason for.

for a middle

Be ready to describe what you would look at first: how often converges report changes, how much bare Ruby the recipes contain, and whether the estate is still on Berkshelf pins rather than Policyfiles.

for a senior

Demonstrate that you would gather evidence before recommending anything, and that you can sequence a migration by host role with a clear owner per file so the two tools never fight over the same configuration.

for a principal

Own the framing: the tool, the code written in it, and the team's fluency are three separate decisions. Argue the freeze-and-shrink path against a rewrite in terms of risk, hiring, and the operational surface you are choosing to keep running.

## Start with measurements, not opinions The weak answer to this question is "Chef is legacy, migrate to X". The strong answer begins by establishing what you actually have. **Are the cookbooks idempotent?** Chef reports, per run, how many resources were updated. A fleet where healthy nodes report zero updates on a steady-state converge has cookbooks that genuinely declare state. A fleet where every node updates a dozen resources every 30 minutes has cookbooks full of unguarded `execute` blocks pretending to be declarations — and that is a code-quality problem you would carry into any tool you migrated to. **How much of it is real Ruby?** Count the cookbook libraries, custom resources, and bare Ruby in recipes. A cookbook that is 95% built-in resources with a few templates is close to portable. One that computes its configuration in library methods and defines a dozen custom resources encodes real logic that a rewrite must reproduce and re-test. **How often does it change?** Configuration that has not been touched in two years is not costing you engineering time; it is costing you only the operational surface. Configuration under weekly change is where the friction is, and that is where a migration pays back. **Who can debug it?** This is the decisive one for Chef specifically. The compile-then-converge model, attribute precedence across cookbook, role, environment and node, and the resource/provider split are all learnable — but somebody has to have learned them before the pager goes off. If the Ruby fluency you inherited is real and stable, that is a genuine asset. If it is one person, the risk is not Chef, it is the bus factor. ## The costs that are specific to Chef - **The Ruby DSL cuts both ways.** It is why Chef can express complex logic that a pure-data format cannot, and it is why recipes can be written imperatively by people who do not know the two-phase model. Cookbooks written that way behave differently on first and subsequent runs, which is exactly the class of bug that is expensive to find. - **There is a server and an agent.** A Chef Infra Server to run, back up and upgrade; node objects that accumulate; an agent on every host converging on an interval. Every one of those is a thing that can page you and a thing that must be patched. - **Licensing has moved.** Chef Infra Client has required explicit license acceptance since version 15, and the commercial terms under Progress have changed over the years. Whatever the current terms are, confirm them before committing to a multi-year plan — this is a due-diligence item, not a technical one. - **A Policyfile migration may still be pending.** An estate on Berkshelf plus environment pins has a known reproducibility gap; closing it is a real project in its own right, and it competes for the same time a migration would take. ## The costs of migrating A cookbook estate is not just code, it is accumulated knowledge of every host quirk the fleet has hit — the sysctl someone set after an incident, the package pinned because a newer one broke the app. Rewriting throws that away and rediscovers it in production. A rewrite also has no natural stopping point: you are not done until the last cookbook is gone, and half-migrated estates mean two tools, two sets of on-call knowledge, and ambiguity about which one owns a given file on a given host. ## The shape of a defensible answer For a working, quiet, low-churn estate: **keep it, and stop it growing.** Freeze new cookbook development, route all new configuration to the target model, invest in the Policyfile migration if it is outstanding because that is cheap relative to a rewrite, and make sure more than one person can read a recipe. The Chef surface then shrinks as hosts are retired rather than as a project. For a noisy, imperative estate maintained by nobody: the cookbooks are not an asset. Do not port them — re-derive the configuration from what the hosts actually need, in whatever the organisation's standard is, one host role at a time, with the old and new paths verifiable against each other before you cut over. The answer an interviewer is listening for is that you separated **the tool** from **the code written in it** and from **the organisation that maintains it**, and that you would not start a rewrite on the strength of the tool's age alone. ## Where you would still choose Chef today Rarely for a greenfield estate — but honestly rather than dismissively: if the fleet is long-lived hosts that must be continuously enforced rather than replaced, the team is fluent in Ruby, and the configuration genuinely needs programmatic logic, a Chef estate with Policyfiles is a coherent, well-understood system. What has changed is not that Chef stopped working; it is that fewer estates now consist of servers that live long enough for continuous convergence to be the interesting problem.

  • What single metric from the existing Chef runs would most change your recommendation?
    The proportion of nodes reporting zero updated resources in steady state. High and stable means the cookbooks are genuinely declarative and idempotent, so they are an asset worth keeping and a poor use of rewrite budget. Persistently non-zero means recipes are re-doing work every converge — unguarded execute resources and imperative Ruby — and that code has little value to port anywhere.
  • How would you handle the period where new configuration goes to the new tool and old configuration stays in Chef?
    Split by ownership boundary, never by file. Each host role, and ideally each file or service on a host, is owned by exactly one tool, written down. Two tools converging the same file is the failure mode that makes people distrust both. Where overlap is unavoidable, remove the resource from the Chef cookbook in the same change that adds it to the new tool, and verify a converge afterwards reports no updates.
  • The team says they cannot migrate because the cookbooks encode years of host-specific fixes. How do you respond?
    Take it seriously — that knowledge is the estate's real value and it is undocumented. Mine it before touching anything: walk the cookbooks and record what each non-obvious resource is for, with a git-blame and incident trail where one exists. That inventory is worth doing whether you migrate or not, and it usually reveals that a good share of the fixes are for platforms or versions the fleet no longer runs.

saying these in an interview costs you the question

  • Recommends migrating because Chef is old, without examining the cookbooks
  • Treats a working idempotent cookbook estate as technical debt by definition
  • Ignores whether anyone left can debug a Chef run
  • Plans a big-bang rewrite with no per-host-role cutover
  • Lets two tools converge the same files during migration

context