skip to content

You need to run a PowerShell health-check script against several thousand servers and collect the results. How would you design the fan-out, and what limits shape the design?

level: principalimportance: should knowfreq 30%

answer

  1. the cmdlet already parallelises
  2. default throttle is 32
  3. batch, stream, bound memory
  4. some hosts are always down
  5. when to stop pushing and start pulling

basics

~20 s

Batch the fleet and let Invoke-Command fan out with a tuned -ThrottleLimit (default 32) rather than hand-rolling parallelism. Reuse sessions for multi-step work, run batches as jobs, tag every result with PSComputerName, and treat unreachable hosts as expected data rather than failures.

solid answer

~50 s

`Invoke-Command -ComputerName` already fans out in parallel, throttled by `-ThrottleLimit`, which defaults to 32; the first design decision is what that number should be given your client's CPU and memory, the network, and how heavy the remote script is. Beyond that, batch: process the fleet in chunks so you never hold thousands of concurrent connections or an unbounded result set in memory. Use `-AsJob` (or PowerShell 7's `ForEach-Object -Parallel`) to keep batches moving and stream results out to storage as they arrive rather than accumulating an array. Keep the payload small — project with `Select-Object` on the remote side, since everything is serialized. Design for partial failure as the normal case: at that scale some hosts are always down, so unreachable hosts must be recorded and retried, not allowed to abort the run. And be honest about the ceiling — past a certain size, an agent-based or purpose-built orchestration system beats a client fanning out live runspaces.

code

powershell · 10 lines
powershell
$check = { [pscustomobject]@{ FreeGB = [math]::Round((Get-PSDrive C).Free / 1GB, 1) } }
$batchSize = 250

for ($i = 0; $i -lt $servers.Count; $i += $batchSize) {
    $end   = [math]::Min($i + $batchSize - 1, $servers.Count - 1)
    $batch = $servers[$i..$end]
    Invoke-Command -ComputerName $batch -ThrottleLimit 64 -ScriptBlock $check -ErrorAction SilentlyContinue |
        Select-Object PSComputerName, FreeGB |
        Export-Csv -Path .\health.csv -Append -NoTypeInformation
}

go deeper

for a junior

Know that Invoke-Command accepts many computer names and contacts them in parallel, and that -ThrottleLimit controls how many at once, defaulting to 32.

for a middle

Explain batching, why results should be projected small before crossing the serialization boundary, how PSComputerName identifies each result, and why sessions must be removed rather than leaked.

for a senior

Design for the real run: measured throttle, bounded batches, results streamed to storage, per-host failures captured as data, idempotent remote work, one retry pass over the failed subset, and a canary batch before the fleet.

for a principal

Own the tradeoff between push and pull. Argue where remoting fan-out stops being appropriate — credential concentration, single point of failure, runtime bounded by the slowest host — and what an agent or convergence-based system buys instead, with a defensible line for your fleet's size.

## Start with what the cmdlet already does The common mistake is hand-rolling parallelism that PowerShell already provides. `Invoke-Command -ComputerName $servers -ScriptBlock {...}` contacts the hosts concurrently, up to `-ThrottleLimit`, whose default is **32**. So the first question is not 'how do I parallelise' but 'what should the throttle be, and what is the batch size'. ```powershell Invoke-Command -ComputerName $batch -ThrottleLimit 64 -ScriptBlock { [pscustomobject]@{ Uptime = (Get-Date) - (Get-CimInstance Win32_OperatingSystem).LastBootUpTime FreeGB = [math]::Round((Get-PSDrive C).Free / 1GB, 1) } } ``` Raising the throttle is not free: each concurrent connection costs client memory, a thread, an authentication handshake and network capacity. The right number is found by measurement on the machine that will actually run this, not copied from a blog post. ## Batch, and stream the results out Thousands of hosts in one call means thousands of simultaneous connections and one enormous in-memory result set. Chunk instead: ```powershell $batchSize = 250 for ($i = 0; $i -lt $servers.Count; $i += $batchSize) { $batch = $servers[$i..([math]::Min($i + $batchSize - 1, $servers.Count - 1))] Invoke-Command -ComputerName $batch -ThrottleLimit 64 -ScriptBlock $check | Export-Csv -Path .\health.csv -Append -NoTypeInformation } ``` Writing each batch out as it completes bounds memory and, just as importantly, means a run that dies at 80% still leaves you 80% of the data. `-AsJob` plus `Receive-Job` is the variant when you want batches in flight while you process earlier results; PowerShell 7's `ForEach-Object -Parallel` is another way to drive per-host work with its own `-ThrottleLimit`. ## Keep the payload small Everything returned crosses the serialization boundary, so a fat object costs CPU on both ends and memory in the middle. Project on the remote side — return a small `[pscustomobject]` with exactly the fields you will analyse. Across 5,000 hosts the difference between a 200-byte record and a 20 KB object dump is the difference between a file you can open and one you cannot. The results arrive tagged with `PSComputerName`, which is what makes aggregation possible. Do not rely on ordering; group on that property. ## Sessions: reuse or not For a single self-contained script block, let `-ComputerName` create and destroy the temporary session — simplest, and nothing to leak. For multi-phase work (import a module, then run several checks against the same host), create sessions with `New-PSSession` per batch, reuse them, and `Remove-PSSession` in a `finally`. At fleet scale, leaked sessions are not a tidiness issue: WinRM enforces per-user concurrent-shell quotas, so a leaking script eventually breaks every other automation running under the same account. ## Partial failure is the steady state With thousands of targets, some are always rebooting, decommissioned but still in inventory, or behind a broken firewall rule. A design that treats any failure as fatal will never complete. Practically: - Capture per-host failures alongside successes so the output distinguishes *unhealthy* from *unreachable* — an unreachable host is a data point, not a gap. - Retry the failed subset once at the end rather than retrying inline, so slow-failing hosts do not stall a batch. - Make the remote script **idempotent** if it changes anything, because a retry will run it twice. - Bound the blast radius: run against a canary batch first, verify the shape of the results, then proceed. ## Know the ceiling The honest principal answer includes when *not* to do this. A client fanning out live runspaces is a pull model with a single point of failure, a credential that can reach everything, and a runtime bounded by the slowest hosts. At a few thousand nodes, or for anything run on a schedule rather than by a human, an agent- or pull-based configuration system — where each node reports in and converges on its own — scales better and fails more gracefully. Ad-hoc remoting fan-out is excellent for investigation and one-off queries, and increasingly the wrong shape for standing operations. Being able to say that, and say where the line sits for your fleet, is what distinguishes this from a cmdlet-parameter answer.

  • What is the default -ThrottleLimit for Invoke-Command, and what happens if you raise it a lot?
    It is 32. Raising it increases concurrent connections, each costing client memory, a thread and an authentication handshake, so past some point the client, the network, or an authentication service becomes the bottleneck and total throughput falls. Find the value by measuring on the machine that will run the job, and remember the remote hosts are also paying for the work.
  • Why write each batch's results to storage as it completes rather than collecting everything and exporting once?
    It bounds memory — thousands of objects held at once is a real cost — and it makes the run resumable in the practical sense: if it dies at 80%, you still have 80% of the data and can retarget the remainder. Streaming out also lets analysis start before the run finishes.
  • How should the design distinguish an unhealthy host from an unreachable one?
    Record both explicitly. Successes carry their `PSComputerName` and the health fields; connection failures should be captured per host and written to the same output with a distinct status, so the final dataset accounts for every target. Otherwise a missing row is ambiguous — it might mean healthy-but-filtered, down, or simply never attempted.
  • At what point would you stop using remoting fan-out and choose something else?
    When the work becomes standing operations rather than investigation: scheduled, fleet-wide, and needing convergence rather than a snapshot. A push model concentrates a credential that reaches everything into one client and is bounded by the slowest hosts, whereas an agent or pull-based system has each node converge and report independently, degrading gracefully when parts of the fleet are unavailable.

saying these in an interview costs you the question

  • Hand-rolls threads instead of using the cmdlet's fan-out
  • Sends all several thousand hosts in one Invoke-Command call
  • Returns whole objects instead of projecting fields remotely
  • Treats any unreachable host as a fatal error
  • Assumes a higher throttle always means faster completion

context