How does a mapper's declared cascade of saves and deletes actually execute, and how does that differ from a schema-level cascade?
answer
- a walk in memory, not in the engine
- one statement per reached object
- only what the layer is tracking
- bulk and direct writes bypass it
- one graph, one transaction
basics
~10 sA declared cascade is an in-memory graph walk at flush: the layer follows marked links and emits one ordinary statement per object. A schema cascade runs inside the engine and covers every writer.
solid answer
~50 sThe layer walks from the objects it is tracking along links declared to cascade, and for each object it reaches it emits the same statement it would emit if you had written that object yourself — ordered so the schema stays satisfied at each step: parents before children on insert, children before parents on delete. Three consequences follow. It is **many statements**, so deleting a parent with thousands of children is thousands of deletes plus the reads to reach them. It covers **only the object graph**, so a bulk statement, direct SQL or another process bypasses it entirely — that is what a foreign-key cascade in the schema is for. And because the statements go out one at a time, a failure partway leaves a half-written graph unless the whole walk sits in one transaction. Removing a child from a collection also only clears the link; deleting it is a separate declaration.
go deeper
Hold on to the basic split: the mapping's cascade is work your application does object by object, while a cascade declared on the foreign key is work the database does. Removing a child from a collection unlinks it; deleting it is a separate setting.
Describe the walk — reachable objects along marked links, one statement each, ordered to satisfy constraints — and name at least one path that bypasses it, such as a statement issued over a predicate.
Show operational judgment: estimate the statement count before enabling a cascade delete, keep the whole graph inside one transaction, and know when to replace the walk with a narrow statement and accept the stale tracked objects that follows.
Decide where each guarantee lives. Integrity that must hold for every writer belongs in the schema; graph bookkeeping belongs in the mapping. Declaring both is normal — declaring only the mapping's version is a data-integrity gap waiting for the first out-of-band write.
A cascade declared in a mapping is not a database feature with a mapper-shaped wrapper around it. It is an **in-memory graph walk** the layer performs when it decides what to write: starting from the objects it is tracking, it follows the links marked to cascade, and for each object it reaches it emits the same ordinary statements it would emit if the code had written that object itself. Everything that follows is a consequence of that one fact. ## How the walk runs 1. At flush, the layer walks from each tracked object along links declared to cascade the operation in question — saving, deleting, or both. 2. Each reached object joins the work: a new one becomes an `INSERT`, a changed one an `UPDATE`, a deleted one a `DELETE`. 3. The layer orders the statements so the schema stays satisfied at each step: parents inserted before the children that point at them, junction rows removed before the rows they reference, children deleted before their parent. 4. The statements go out one at a time, or in batches, over the same connection and inside the same transaction as everything else in the unit of work. ## Cascade versus the schema's own rule | | Cascade declared in the mapping | Cascade declared on the foreign key | |---|---|---| | Who executes it | the layer, in application memory | the engine, inside the statement | | What it covers | objects reachable from what the layer is tracking | every row that references the parent, from any client | | Statements produced | one per affected object | one, expanded internally by the engine | | Visible to the tracked set | yes; the layer knows those objects are gone | no; the layer keeps objects for rows that no longer exist | | Paths it misses | bulk statements, direct SQL, other applications | none for that foreign key | The two are not alternatives so much as different guarantees: one keeps the object graph honest, the other keeps the data honest no matter who writes it. Many systems declare both, and then the mapper's deletes run first and the schema rule finds nothing left to do. ## What the walk cannot see - A **bulk statement issued through the layer** — a delete or update over a predicate — is sent to the engine as written. It removes rows without instantiating objects, so no walk happens and no declared cascade runs. It also leaves the tracked set holding objects for rows that are gone. - **Direct SQL**, another service, or an operator at a console are further outside the walk entirely. - Objects the layer never loaded and cannot reach from what it has loaded are not visited, so what the cascade covers depends on what the graph in memory happens to contain. ## Statement count is the operational risk Deleting a parent with ten thousand children through a declared cascade is at least ten thousand and one statements, and the layer must usually load those children to walk to them, so the read cost comes first. For a handful of children this is invisible; at scale it is a long transaction holding locks over a large row set. The remedy is not to fight the cascade but to bypass it for that case: delete the children with one predicate-based statement, then the parent, and accept that the tracked set now holds stale objects — which is fine if the unit of work ends there. ## Unlink is not delete Removing a child from a parent's collection means, by default, **the link is cleared** — the child row survives with a null or unchanged key, depending on which end owns the link. Deleting the detached child is a **separate declaration** (removal of orphans). Getting this wrong in either direction is common: expecting rows to disappear and finding a growing tail of unreferenced children, or expecting a temporary detach and finding the row deleted. And if the link column is not nullable, the plain unlink cannot even be written — the flush fails on a constraint the code never mentioned. ## Half-saved graphs Because the walk emits statements one at a time, a failure partway through leaves part of the graph written. That is harmless when the whole unit of work is one transaction: the rollback removes all of it. It stops being harmless when the code commits inside the loop that builds the graph, when a long import flushes in stages and treats each stage as durable, or when a nested call opens its own transaction for part of the work. Then a constraint violation on the tenth child leaves nine children and a parent committed, and the retry re-inserts what already exists. The rule that keeps this simple: **one graph, one transaction** — and if the graph is too big for one transaction, split it deliberately into units that are each valid on their own, rather than letting the flush boundary decide. ## What an interviewer is listening for That the cascade is the layer's walk, not the engine's; that its coverage is exactly the object graph and therefore excludes bulk and direct writes; that its cost is a statement per object; that unlinking and deleting the orphan are different declarations; and that a half-saved graph is a transaction-boundary bug rather than a mapper defect.
- You delete a parent with fifty thousand children through a declared cascade and the transaction runs for minutes. What do you change?Bypass the object path for that case: delete the children with one predicate-based statement, then the parent, and end the unit of work rather than trusting the tracked objects afterwards. Keep the schema-level rule as the backstop for writers that never go through the layer.
- Why does a bulk delete issued through the layer skip the declared cascade?It is sent to the engine as a statement over a predicate, so no objects are instantiated and there is no graph to walk. It also removes rows for objects the layer may still be tracking, leaving stale instances behind for the rest of the unit of work.
- If the schema already declares a cascade on the foreign key, is the mapping's cascade redundant?No — they guarantee different things. The schema rule keeps the data consistent for every writer, including those that never touch the layer. The mapping's cascade keeps the in-memory graph honest and lets the layer's own bookkeeping know those objects are gone.
saying these in an interview costs you the question
- Thinks the mapping's cascade is executed by the database engine
- Expects a cascade delete of a large collection to be one statement
- Assumes bulk statements and direct SQL honour the declared cascade
- Treats removing a child from a collection as deleting the child row
- Commits inside the graph walk and blames the layer for partial data