When would you still run full-parameter SFT instead of an adapter-based one?
answer
- quality is no longer the deciding axis
- capacity, not fidelity
- one artifact per task, forever
- can you undo it in a minute
- large shift vs teaching form
basics
~20 sRarely, as of mid-2026: a well-configured adapter matches full fine-tuning on typical SFT data at a fraction of the compute. Full-parameter updates earn their cost only when the data volume exceeds adapter capacity or the target behaviour is a large distribution shift.
solid answer
~50 sThe default has flipped. Work published in late 2025 and since absorbed into mainstream training tooling showed that a properly configured adapter matches full fine-tuning on post-training-scale SFT data at roughly two-thirds of the compute, so treating full-parameter training as the quality option is out of date. Full-parameter SFT still wins in two situations: when the dataset is large enough to exceed what a low-capacity adapter can absorb, and when you are pushing a genuinely large distribution shift — a new language, a new modality-like output regime — rather than teaching form on top of behaviour the base half-does. Against that you weigh real operational costs: optimizer state for every parameter, a full model artifact per task rather than a small one, no cheap rollback because the original weights are gone, and a larger forgetting and safety-regression surface. In practice I default to an adapter, prove it insufficient on held-out quality, and only then spend the full-weight budget.
go deeper
Know the distinction: full-parameter fine-tuning updates every weight, while an adapter trains a small added set and leaves the base frozen. Know that the adapter route is the common default.
Explain the resource difference — gradients and optimizer state for every parameter versus a small trained set — and that a well-configured adapter matches full fine-tuning on typical SFT data, so quality is not the deciding factor.
Reason about operations: artifact count per task, merged versus unmerged serving, rollback speed when a regression appears, and the empirical signal that a dataset has exceeded adapter capacity.
Set the organizational default and the exception criteria. Argue adapters as the standing choice on reversibility and cost grounds, define what evidence justifies a full-weight run, and account for the storage, regression-testing and governance obligations that decision creates.
## The question behind the question An interviewer asking this is testing whether your priors are current and whether you reason about training decisions with serving and rollback in mind, not just quality. The naive answer — "full fine-tuning is better, adapters are the budget option" — was defensible in 2023 and is wrong now. ## What changed By late 2025, careful head-to-head work established that a well-configured adapter reaches the same loss as full fine-tuning on supervised fine-tuning datasets of typical post-training scale, at meaningfully lower compute — roughly two-thirds. The earlier folklore that adapters were a lossy approximation came largely from under-configured runs. The important consequence for this decision is that **quality is no longer the axis that separates them** for ordinary SFT. Once that is true, the choice is made on capacity, cost and operations. The degradation regime is real but specific: an adapter has finite capacity, and when the dataset carries more information than that capacity can hold, the adapter falls behind. That is a large-dataset phenomenon, in the territory where you are approaching continued-pretraining scale rather than instruction tuning. ## When full-parameter SFT genuinely earns its cost **The dataset exceeds adapter capacity.** Millions of examples, or a corpus whose breadth approaches the pretraining distribution, will saturate a constrained update. If held-out quality with a well-configured adapter plateaus below what you need and does not improve as you give the adapter more capacity, that is the empirical signal. **A large distribution shift.** Teaching a model a language it barely represents, or an output regime structurally unlike anything it produces, changes representations broadly rather than steering existing behaviour. Narrow instruction tuning — "emit this structured summary from this transcript" — is emphatically not this case; it is form on top of an ability the base model half-has. **Serving constraints you cannot change.** Some inference stacks will only load a single monolithic checkpoint. An adapter can be merged into the base weights to produce exactly such an artifact, so this argument is weaker than it looks — but if merging is not available in your toolchain, it can force the decision. **Deliberate, permanent re-basing.** If the tuned model is the new base for everything downstream and you never want the original behaviour back, the reversibility advantage of an adapter is worth nothing to you. ## What full-parameter SFT costs **Memory and compute.** Every parameter needs gradients and optimizer state, several times the model's own footprint. This is what pushes runs onto larger clusters and is the direct source of the compute gap. **Artifacts.** One full model per task. Ten tasks is ten full checkpoints to store, version, scan and deploy. Adapters are small files against one shared base, which changes the storage and distribution story qualitatively, not marginally. **No cheap rollback.** This is the underrated one. With a full-weight update, the tuned model *is* the weights; reverting means redeploying a different multi-gigabyte artifact. With an adapter, the base is untouched and disabling the adapter restores prior behaviour immediately. When the thing you are rolling back is a safety regression found in production, minutes versus a redeploy cycle is a governance property, not a convenience. **A larger drift surface.** Updating every parameter gives forgetting and alignment erosion more room to operate than a constrained update does. Neither approach is immune, and both need base-versus-tuned regression testing, but the full-weight update starts from a worse position. **Iteration speed.** Cheaper runs mean more experiments per week, and on SFT the number of dataset iterations you can afford usually dominates the choice of parameterisation in its effect on final quality. ## The decision I would actually defend Default to an adapter. Fix the data first — templates, masking, turn boundaries — because formatting defects explain more failed SFT runs than parameterisation ever will. Evaluate against a held-out set and against the base model on general capability and refusals. Only if quality is genuinely capacity-bound, and giving the adapter more capacity keeps helping, do you consider a full-parameter run — and then you budget for the storage, the regression testing and the rollback story up front. Stated as a principle: the parameterisation is an operational decision with a quality constraint, not a quality decision with an operational footnote. Anyone who opens with "full fine-tuning for best quality" has not looked at the evidence since it changed.
- If adapters match full fine-tuning on quality, why does full-parameter training still exist for SFT?Because parity holds on post-training-scale data, not universally. An adapter has bounded capacity, so a dataset large or broad enough to exceed it will train better with a full update. Large distribution shifts — a new language, an unfamiliar output regime — also change representations broadly rather than steering existing behaviour. Outside those regimes, full-parameter SFT mostly buys cost.
- Does serving need to differ between an adapter and a full fine-tune?Not necessarily. An adapter can be merged into the base weights to produce a single ordinary checkpoint, so serving stacks that only load one monolithic model are still served. Keeping it unmerged buys instant rollback and lets several task adapters share one loaded base, which is the operational reason to prefer it when the stack supports it.
- How does the rollback difference change your risk posture?With a frozen base, a safety or quality regression found in production is undone by disabling the adapter — effectively immediate. With a full-weight update the tuned model is the weights, so reverting is a redeploy of a large artifact. That difference is why I treat adapters as the default for anything customer-facing, independent of the compute argument.
- What would make you stop and fix data rather than change parameterisation?Almost always the first response to disappointing SFT quality. Broken chat templates, mis-masked loss, missing stop tokens and inconsistent turn boundaries account for more failed runs than the choice between full and adapter training. If a run underperforms, I decode rendered examples and check the supervised spans before spending a larger compute budget on the same defective data.
saying these in an interview costs you the question
- Says full fine-tuning is always the higher-quality option
- Ignores per-task artifact storage and versioning cost
- Treats rollback as equivalent for full-weight and adapter updates
- Reaches for full-parameter training to fix a data-formatting problem
- Assumes an adapter cannot be served as a single merged checkpoint