On Windows, how would you decide whether new functionality — say, intercepting or instrumenting file I/O — belongs in a kernel-mode driver or a user-mode component, and what does choosing kernel mode commit you to?
answer
- ask what user mode cannot do
- being in the path, not beside it
- no isolation above the line
- raised IRQL narrows what you may call
- keep the kernel piece small
basics
~20 sDefault to user mode and move to kernel mode only when you must sit in the I/O path or observe what no user-mode interface exposes. Kernel mode means a bug bugchecks the whole machine, and it commits you to IRQL and pool discipline, code signing, and lockstep servicing with the OS.
solid answer
~60 sStart from the question "what can I not do from user mode?" Existing user-mode surfaces — event tracing, documented Win32 and native APIs, a service consuming those feeds — satisfy more requirements than people expect, and a crash there kills one restartable process. You need kernel mode when you must be *in* the path rather than beside it: a file system filter driver that can block or transform an operation before it completes, or a device driver for hardware nothing else exposes. Choosing it brings real obligations. Your code shares the kernel address space, so a bad pointer bugchecks the machine instead of faulting your process. You inherit IRQL rules — at raised IRQL you may not touch pageable memory or block — and pool discipline, where a leak is machine-wide. You must be signed to load and stay compatible with code-integrity enforcement. And you are coupled to OS servicing: every feature update becomes a compatibility event. Name the middle grounds too — a user-mode driver framework component, or a thin kernel collector feeding a user-mode service that holds the logic.
go deeper
Know that Windows code runs either in user mode or in kernel mode, that drivers are the kernel-mode kind, and that ordinary applications and services are written in user mode for good reason.
Be able to name what forces kernel mode — being in the I/O path with authority to block, or servicing hardware with no existing stack — and which user-mode alternatives, such as event tracing, already cover observation.
Speak to the operational obligations: a shared address space so faults bugcheck, IRQL and pool rules constraining what your code may call, signing before it loads, and compatibility testing against every OS update.
Own the placement strategy for a product: keep the kernel footprint minimal and the logic in user mode, prefer supported extensibility frameworks over bespoke interception, and pair the decision with rollout gating, kill switches and crash telemetry so a kernel component never becomes an unrecoverable fleet risk.
## Frame the decision as "what is impossible in user mode?" The answer that impresses is not a preference for one mode; it is a disciplined elimination. Kernel mode is the place with no isolation, no second chances and the heaviest deployment obligations, so it should be entered only for capabilities that genuinely do not exist above the line. Capabilities that usually *do* exist in user mode: - **Observation.** Event Tracing for Windows exposes process, thread, image-load, registry and file activity to a consuming service, with the kernel doing the collection. If you only need to know what happened, you rarely need to be a driver. - **Policy at an API boundary.** Many controls can be applied where the operation is requested rather than where it is serviced. - **Hardware access through an existing stack.** If a class driver already exposes the device, a user-mode component talking to it is enough. Capabilities that require kernel mode: - **Being in the path, synchronously, with authority to fail or transform an operation.** A file system filter driver sees the request before the file system does, may modify or block it, and does so for every process on the machine including ones that started before you did. No user-mode hook has that reach or that authority — a user-mode hook is per-process, evadable and racy. - **Servicing new hardware or a new bus** for which no stack exists. - **Work at raised IRQL** — the timing-critical part of interrupt handling. ## What kernel mode commits you to **No fault isolation.** There is one kernel address space. A stray write does not raise an access violation inside your component; it corrupts someone else's data structures and the machine bugchecks, possibly minutes later and somewhere unrelated. Diagnosis is crash-dump analysis, not a stack trace in your log. **IRQL and pool discipline.** Kernel code runs at an interrupt request level, and the rules at raised IRQL are strict: at `DISPATCH_LEVEL` you may not touch pageable memory and you may not block on a dispatcher object, because the scheduler itself is unavailable to you. Allocations must come from the right pool — non-paged memory for anything touched at raised IRQL — and a pool leak in a driver is a machine-wide resource leak that survives every process restart. None of this has an equivalent in user-mode programming, and it is where most first-time driver bugs live. **Loading is a privilege granted by signing.** A kernel-mode driver must be signed to load on modern 64-bit Windows, so you take on certificate management and a submission process, and you must remain compatible with hypervisor-enforced code integrity where it is enabled. "Just ship a hotfix" is not available the way it is for a user-mode service. **Servicing coupling.** You now live inside the OS's release cadence. Kernel structures and driver interfaces evolve, and each feature update is a compatibility test event. A user-mode service consuming documented APIs simply does not carry that risk at the same magnitude. **Blast radius owns your support cost.** When your driver is on a fleet and a bugcheck spike appears after an update, you are the prime suspect regardless of fault, because kernel code is what stops machines. ## The middle grounds A good senior answer stops at "driver or service"; a good principal answer names the hybrids: - **User-mode driver framework components.** For many device classes a driver can run in user mode, gaining crash isolation and easier debugging at the cost of some performance and some capability. Where the class supports it, that is often the right default. - **Thin kernel collector, thick user-mode brain.** Put the smallest possible amount of logic in the kernel — only the part that must be in the path — and ship everything analytical, updatable and risky to a user-mode service. Your parsing, your policy language, your network client and your update mechanism have no business in kernel memory, and keeping them out means a bug in them cannot bugcheck a customer. - **Supported filtering frameworks rather than bespoke hooking.** Where the platform offers an extensibility point — a file system filter framework with managed ordering, or the network filtering platform — use it. The framework handles ordering, teardown and unload semantics that hand-rolled interception gets wrong. ## How to state the decision Say what capability forces kernel mode, say how little of the system will live there, and say how you will contain the risk: staged rollout gated on crash telemetry, verifier-driven testing before shipping, a kill switch that disables enforcement without a driver update, and a boundary drawn so that most of your change velocity happens in user-mode code you can roll back without a reboot.
- Why is a user-mode API hook a poor substitute for a file system filter driver?Because it is per-process, opt-in and evadable. You must inject into every process, you miss anything that started before you or that you cannot inject into, and the hooked code can be restored or sidestepped by calling a lower layer directly. A filter in the I/O path sees every request on the volume regardless of caller, cannot be bypassed from user mode, and can fail or transform an operation authoritatively.
- What does the IRQL rule at DISPATCH_LEVEL actually forbid, and why?Touching pageable memory and waiting on a dispatcher object with a nonzero timeout. Both need the scheduler and the page-fault path, and at that level the current processor cannot be rescheduled — a page fault would be unresolvable and a wait would deadlock the processor. It is why drivers pre-allocate from non-paged memory and defer real work to a lower level.
- How would you keep a required kernel component from becoming a fleet-wide risk?Make it as small and as dumb as possible: it should collect or gate, never parse untrusted data or hold updatable policy. Put the logic in a user-mode service you can update and roll back without a reboot, give the pair a kill switch that disables enforcement without unloading, test under the driver verifier, and roll out in stages with crash telemetry as the gate.
saying these in an interview costs you the question
- Reaches for a driver whenever tracing would do
- Thinks a user-mode API hook gives the same coverage as a filter
- Assumes a driver crash just terminates the driver
- Treats code signing and servicing coupling as paperwork
- Puts parsing and policy logic in kernel memory