In PowerShell, when should you use the `ForEach-Object` cmdlet rather than the `foreach` statement, and how do the two differ in the way they consume their input?
answer
- one collects, one streams
- statement materialises the collection first
- cmdlet sees $_ one item at a time
- return skips an item, break kills the pipeline
- -Parallel exists only on the cmdlet
basics
~20 sThe foreach statement evaluates its whole collection into memory first, then loops over it quickly. ForEach-Object is a cmdlet that streams: it runs its script block on each item as it arrives, so it composes inside a pipeline and holds constant memory, at a higher per-item cost.
solid answer
~50 sThey solve the same problem at different points in the pipeline. The `foreach` statement is a language construct: the collection expression is evaluated and materialised before the first iteration, so `foreach ($line in Get-Content big.log)` pulls the whole file into memory, but each iteration is cheap and `break`/`continue` behave the way you expect. `ForEach-Object` is a cmdlet in the pipeline, so it receives one object at a time through `$_` and passes results straight downstream — memory stays flat over an arbitrarily long stream, and it can sit between `Where-Object` and `Sort-Object`. The tradeoffs are per-item overhead (the cmdlet invokes a script block per object, measurably slower on large in-memory collections) and surprising control flow: inside the script block `return` just ends the current item, while `break` and `continue` are not loop keywords and can terminate the pipeline. In PowerShell 7 `ForEach-Object -Parallel` adds runspace-based concurrency, which the statement has no equivalent of.
code
powershell · 5 lines# Statement: the whole file is materialised before the first iteration
foreach ($line in Get-Content .\big.log) { if ($line -match 'ERROR') { $line } }
# Cmdlet: streams line by line, flat memory, composes with other stages
Get-Content .\big.log | ForEach-Object { if ($_ -match 'ERROR') { $_ } } | Select-Object -First 20go deeper
Know both forms and be able to write each. Say clearly that the cmdlet form works in a pipeline with $_ while the statement form loops over a collection you already have.
Explain the consumption difference: the statement materialises the whole collection before iterating, the cmdlet processes one object at a time and keeps memory flat. Mention the per-item overhead that makes the statement faster on in-memory data.
Demonstrate the control-flow trap — return skips an item while break can kill the pipeline — and pick the right form for a large stream versus a materialised collection, including when -Parallel actually pays.
Frame it as a throughput-versus-footprint decision for pipelines your team will run on unknown data sizes, and set a convention that keeps streaming shapes the default so a bigger input never becomes an outage.
## Two things with almost the same name `foreach` in PowerShell is overloaded, which is half the confusion: - At the **start of a statement**, `foreach` is a language keyword: `foreach ($item in $collection) { ... }`. - **After a pipe**, `foreach` resolves to an alias for the `ForEach-Object` cmdlet: `$collection | foreach { ... }`. Same word, two different execution models. Spelling out `ForEach-Object` in shared code removes the ambiguity for readers. ## How each consumes input The `foreach` **statement** evaluates its collection expression fully before the loop body runs even once. The results are collected, then iterated: ```powershell foreach ($line in Get-Content .\big.log) { if ($line -match 'ERROR') { $line } } ``` `Get-Content` runs to completion and every line is held before the first `if` executes. On a 5 GB log that is a memory problem, not a style problem. `ForEach-Object` is a cmdlet occupying a pipeline stage. Objects arrive one at a time, the script block runs with `$_` (equivalently `$PSItem`) bound to the current item, and anything the block emits flows immediately to the next stage: ```powershell Get-Content .\big.log | ForEach-Object { if ($_ -match 'ERROR') { $_ } } ``` Here `Get-Content` streams, the script block runs per line, and memory stays flat regardless of file size. The same shape is what lets `ForEach-Object` sit in the middle of a longer pipeline. `ForEach-Object` also accepts `-Begin` and `-End` script blocks that run once before and after the stream, which is how you set up and tear down state around a streaming operation — the statement form has no equivalent because you simply write code before and after the loop. There is a shorthand worth knowing: `ForEach-Object PropertyName` (the `-MemberName` parameter, positional) projects a member without a script block, and `ForEach-Object MethodName` calls a method on each item. ## Speed versus memory On a collection that already fits in memory, the `foreach` statement is faster — often several times faster — because the cmdlet pays pipeline and script-block-invocation overhead per object. On a stream you cannot or should not materialise, `ForEach-Object` wins outright because the statement's collect-first behaviour is exactly what you are trying to avoid. That is the whole tradeoff: **statement = speed on materialised data; cmdlet = streaming and composability**. A middle path many people forget: you can pipe into `ForEach-Object` for the streaming half and use a plain `foreach` statement inside it for a small nested collection. ## Control flow is not the same This is the trap that separates a rehearsed answer from a real one. In the `foreach` statement, `break` exits the loop and `continue` skips to the next item — the ordinary meaning. Inside a `ForEach-Object` script block there is no loop, only a script block invoked repeatedly. Consequently: - `return` ends processing of the **current item** and moves to the next. It does not return from the enclosing function. - `break` and `continue` are not scoped to the cmdlet. They look for an enclosing loop; with none present, `break` terminates the pipeline rather than skipping one item. People write `continue` expecting "skip this object" and get behaviour they did not intend. So: to skip an item inside `ForEach-Object`, use `return` (or restructure with a `Where-Object` filter upstream, which is usually clearer and faster). ## Parallelism (PowerShell 7) `ForEach-Object -Parallel { ... }` runs the script block for multiple items concurrently in separate runspaces, with `-ThrottleLimit` capping concurrency. Because each runspace is isolated, variables from the caller must be referenced as `$using:varName`, and ordering of output is not guaranteed. It is a genuine win for I/O-bound work — hitting many hosts or URLs — and usually a loss for cheap CPU-light work, where runspace setup costs more than the work itself. There is no parallel form of the `foreach` statement. ## Choosing, in practice - Data already in a variable, and you want speed → `foreach` statement. - The input is a stream, a huge file, or you are mid-pipeline → `ForEach-Object`. - You need `break` to abandon the loop early on a materialised collection → statement (clean semantics). - You need per-item work fanned out concurrently → `ForEach-Object -Parallel` on PowerShell 7. - You just want a property or a method call on each item → `ForEach-Object Name` shorthand, or `Select-Object -ExpandProperty`.
- Inside a `ForEach-Object` script block, how do you skip the current item and move to the next?Use `return`. It ends the current invocation of the script block only, so processing continues with the next object. `continue` and `break` are loop keywords with no loop to bind to here — `break` can terminate the whole pipeline instead of skipping one item. Cleaner still is to filter upstream with `Where-Object` so the item never reaches the block.
- When is `ForEach-Object -Parallel` a bad idea?When the per-item work is cheap. Each concurrent item runs in its own runspace, and creating runspaces plus marshalling data costs far more than a trivial CPU-light operation, so the parallel version is often slower. It also loses output ordering and requires `$using:` to see caller variables. It pays off for I/O-bound fan-out — many hosts, many web calls — with a sensible `-ThrottleLimit`.
- Why do people say `$collection | foreach { }` and `foreach ($x in $collection) { }` are different commands?Because they are. At the start of a statement `foreach` is the language keyword and collects its collection first; after a pipe it resolves to the alias for the `ForEach-Object` cmdlet, which streams. Same word, different execution model — which is exactly why shared code should spell out `ForEach-Object`.
saying these in an interview costs you the question
- Says the foreach statement streams its input lazily
- Uses continue inside ForEach-Object expecting to skip one item
- Claims ForEach-Object is always faster because it is a cmdlet
- Thinks $_ is available inside the foreach statement body
- Assumes -Parallel speeds up any loop