skip to content

A Laravel seeder creates 500 appointments whose factory sets 'doctor_id' => Doctor::factory() — why does it also create 500 doctors, and how do you fix it?

level: seniorimportance: should knowfreq 26%

answer

  1. definition() runs once per record
  2. a Factory value is created, then keyed
  3. recycle() takes a model or collection
  4. random pick from the pool per use
  5. reaches nested factories and for()

basics

~10 s

definition() runs once per appointment, and a nested Doctor::factory() is created each time to supply the key. Create a pool first and chain recycle($doctors), so the factory picks a random existing doctor instead.

solid answer

~40 s

When the factory expands attributes, any value that is itself a `Factory` is resolved by calling `create()` on it and taking its key, once per record, so 500 appointments mean 500 doctors (and 500 patients if `patient_id` works the same way). The fix is to create the pool once, `$doctors = Doctor::factory()->count(20)->create()`, and chain `->recycle($doctors)`: before creating a model of that class, the factory picks a random one from the recycled pool. `recycle()` accepts a model, several models or a collection, groups them by class, and is passed down into nested factories and `for()` / `has()` relations, so a department referenced by both the doctor and the appointment is shared too. Use `for($doctor)` to pin every record to one parent, and `withoutParents()` when the key may stay null.

code

php · 23 lines
php
<?php

namespace Database\Seeders;

use App\Models\Appointment;
use App\Models\Doctor;
use App\Models\Patient;
use Illuminate\Database\Seeder;

class AppointmentSeeder extends Seeder
{
    public function run(): void
    {
        $doctors = Doctor::factory()->count(20)->create();
        $patients = Patient::factory()->count(200)->create();

        // 500 appointments, zero new doctors or patients
        Appointment::factory()->count(500)
            ->recycle($doctors)
            ->recycle($patients)
            ->create();
    }
}

go deeper

for a junior

Recall that a factory call in definition() creates a new related model for every record it builds.

for a middle

Explain how recycle(), for() and withoutParents() change what happens to a nested factory attribute.

for a senior

Diagnose a seeder from its row counts, then pick a pool, a fixed owner or a sequence per relation to get a realistic, fast demo dataset.

for a principal

Set team conventions for factories so implicit parent creation stays convenient for one-off use without silently inflating large seeders.

## Why the parents multiply A Laravel **model factory** evaluates `definition()` once for **each** record it builds. After merging states, it walks the attribute array and **expands** it: - a value that is another `Factory` is created with `create()`, and its primary key replaces the factory; - a value that is a `Model` is replaced by its key; - a closure is called with the attributes so far. So `'doctor_id' => Doctor::factory()` is not "some doctor"; it is "a new doctor for this row". A seeder that runs `Appointment::factory()->count(500)->create()` inserts 500 doctors, and if `patient_id` and `department_id` are wired the same way, hundreds more rows follow. The data is valid but useless as a demo: no doctor has more than one appointment, and seeding is several times slower than it needs to be. ## recycle(): reuse a pool `recycle()` hands the factory models to reuse instead of creating new ones: 1. Create the pool once: `$doctors = Doctor::factory()->count(20)->create();`. 2. Chain it: `Appointment::factory()->count(500)->recycle($doctors)->create();`. 3. During expansion, before creating a nested model of a class, the factory checks its recycled pool for that class and, if found, uses a **random** member's key. Details that matter: - `recycle()` accepts a single model, several models as arguments, an array or a collection, and can be called repeatedly; models are grouped by class, so `->recycle($doctors)->recycle($patients)` covers both parents; - the pool is passed down into nested factories and into `for()` and `has()` relations, so if `DoctorFactory` itself references `Department::factory()`, recycled departments are shared across the whole tree; - the choice is random per use, so the spread is uneven; use a sequence when you need an exact distribution. ## Other tools and how they differ | Approach | Parents created | Distribution | |---|---|---| | `'doctor_id' => Doctor::factory()` in `definition()` | one per record | one appointment each | | `->for(Doctor::factory())` | one for the chain | all on one doctor | | `->for($doctor)` | none | all on that doctor | | `->recycle($doctors)` | none (pool pre-made) | random across the pool | | `->withoutParents()` | none | key left null | | `sequence(fn ($s) => ['doctor_id' => ...])` | none | exactly as you compute it | `withoutParents()` switches off the expansion of nested factories (or, given class names, only those), so the foreign key stays null; it suits `make()` for in-memory objects but not inserts into a `NOT NULL` column. ## Diagnosing it in a real seeder The symptom is usually a row count that does not match the plan: `SELECT COUNT(*)` on the parent table equals the child count, or seeding a demo takes minutes. Read the child factory's `definition()` for nested `::factory()` calls, then decide per relation whether the demo needs a shared pool (`recycle`), one owner (`for`) or none (`withoutParents`). Keep the nested factory in `definition()` anyway: it is what makes `Appointment::factory()->create()` work on its own when nobody supplied a doctor. ## A worked fix For the hospital demo, the planned shape is 20 doctors, 200 patients and 500 appointments. The seeder that produced 500 doctors is rewritten in three steps: 1. Create the parents first and keep the collections: `$doctors` and `$patients`. 2. Build the appointments with `->recycle($doctors)->recycle($patients)`, so every nested `Doctor::factory()` and `Patient::factory()` in `definition()` is answered from the pools. 3. Check the counts after `migrate:fresh --seed`: 20, 200 and 500 rows, with each doctor holding roughly 25 appointments. If the demo needs a predictable distribution instead, for example every doctor exactly 25 appointments, loop over `$doctors` and call `Appointment::factory()->count(25)->for($doctor)->recycle($patients)->create()` for each one. `for()` pins the doctor while `recycle()` still spreads patients. ## Cost at scale `create()` still saves each appointment with its own `INSERT`, fires model events and runs `afterCreating` callbacks. For very large demo sets the factory's `insert()` method builds the models and writes them with one multi-row insert instead, skipping model events, `afterCreating` callbacks and `has()` children; nested parent factories in `definition()` are still resolved first, so `recycle()` matters there too.

  • When would you choose for() over recycle()?
    Use `for($doctor)` when every record must belong to one specific parent, such as a doctor's day list. Use `recycle($doctors)` when records should spread across a pool of parents; it chooses a random member each time a parent of that class is needed.
  • What does the factory's insert() method skip compared with create()?
    `insert()` builds the models and writes them in one multi-row insert through the query builder, so there is no `save()` per model: no model events, no `afterCreating` callbacks, no `has()` children, and nothing is returned. Nested parent factories are still created first.

saying these in an interview costs you the question

  • A nested factory in definition() is resolved once and shared
  • recycle() always assigns the first model in the pool
  • recycle() only affects has() children, not definition() keys
  • for(Doctor::factory()) creates a new doctor per record
  • withoutParents() deletes the extra parents after saving