How would you structure model selection across providers in a LangChain service?
answer
- build models from config, not imports
- name tiers, not model ids
- the interface hides real differences
- a config switch is a behaviour change
- evaluation is the gate, rollback is config
basics
~20 sConstruct models from configuration with init_chat_model rather than importing provider classes at call sites, keep a small set of named tiers per task, and gate any switch behind evaluation — because capability differences leak through the shared interface even when the code compiles.
solid answer
~50 sUse `init_chat_model("openai:gpt-4o-mini", temperature=0)` so the provider and model come from config, not from an import in business logic; `configurable_fields` plus a `config` payload lets a single instance be re-pointed per request when you need routing. Above that, define a handful of named tiers — a cheap model for classification and extraction, a stronger one for synthesis, a fallback for outages — and let features reference tiers rather than model ids. Then be honest about what the abstraction does not equalize: strict JSON-schema structured output, tool-calling fidelity, streamed usage reporting, reasoning-token behaviour, and context windows all differ, so a config-only switch is a code-level no-op and a behaviour-level change. That makes evaluation the gate: an offline suite plus a canary on live traffic before any tier is re-pointed, with cost and latency budgets tracked per tier.
code
python · 12 linesfrom langchain.chat_models import init_chat_model
# provider chosen by config string, no provider class imported
fast = init_chat_model("openai:gpt-4o-mini", temperature=0)
# one instance, re-pointed per request
router = init_chat_model(temperature=0, configurable_fields=("model", "model_provider"))
answer = router.invoke(
"Classify this ticket: refund request",
config={"configurable": {"model": "gpt-4o-mini", "model_provider": "openai"}},
)go deeper
Know that init_chat_model builds a model from a config string like "openai:gpt-4o-mini", so provider choice lives in configuration rather than in imports across the codebase.
Explain configurable_fields and per-call config routing, and be able to say why a single instance re-pointed at runtime beats constructing new model objects per request.
Show what breaks on a switch — structured-output strictness, tool-call fidelity, streamed usage reporting, context limits — and how canary rollout plus parse-failure and cost metrics catch it.
Own the whole policy: named tiers with a single owner, evaluation as the gate for any re-point, budgets and rollback as config, and an explicit decision about where you deliberately couple to one provider.
## The mechanism first `init_chat_model` builds a chat model from a string rather than an import: `init_chat_model("openai:gpt-4o-mini")`, or `init_chat_model("gpt-4o-mini", model_provider="openai")`. The provider package must be installed, but nothing in your business logic mentions `ChatOpenAI`. That alone removes the most common coupling: provider class names scattered across dozens of call sites. For runtime selection, `configurable_fields` makes named fields settable per call through the `config` dict — so one instance can serve a request routed to a cheap model and the next routed to an expensive one, without rebuilding objects. `config_prefix` keeps those keys namespaced when several configurable models coexist in one chain. ## Tiers, not model ids The design that survives contact with a growing codebase is a small, named set of tiers owned centrally: - **fast/cheap** — classification, routing, extraction, guardrail checks. High volume, low stakes. - **strong** — synthesis, long-context reasoning, anything user-facing and quality-sensitive. - **fallback** — a different provider entirely, for outages and quota exhaustion. Features request a tier. The mapping from tier to concrete model lives in config, so re-pointing a tier is one change, reviewed by whoever owns cost and quality. Without this, model ids proliferate: a year in, nobody can answer "which models are we actually calling in production" and a deprecation notice becomes an archaeology project. ## Where the abstraction leaks The interview substance is knowing that a uniform interface is not uniform behaviour. Things that differ under the same `invoke()` call: - **Structured output.** Strict JSON-schema mode exists on some providers and not others; `"function_calling"` is broader but weaker. A schema that validates everywhere under one provider may be rejected or loosely honoured under another. - **Tool calling.** Parallel tool calls, argument fidelity and adherence to descriptions vary materially between models of the same nominal quality. - **Usage reporting.** Whether streamed calls report usage, and whether cached-input and reasoning tokens are broken out, is provider-specific — so your cost ledger can lose fidelity on a switch. - **Reasoning models.** They may spend large hidden output budgets and ignore some sampling parameters, changing both latency and price shape. - **Context window and truncation behaviour**, and what happens when you exceed it. - **Prompt sensitivity.** The same system prompt does not produce the same behaviour across model families; prompts are model-coupled artefacts even though the code is not. Because of these, treat a model change like a dependency upgrade with behavioural risk, not like a config typo fix. ## The gate: evaluation and rollout A credible answer includes how you would actually make a switch safely: 1. **Offline suite** — a fixed set of representative inputs with graded outputs per tier, run before any re-point, scoring correctness, schema-validity rate, tool-call accuracy and refusal rate. 2. **Shadow or canary** — route a small share of live traffic to the candidate and compare outcome metrics, not just latency and cost. Structured-output parse failure rate is an unusually sensitive early signal. 3. **Budgets** — per-tier cost and p95 latency targets, alerting on regression, so a cheaper model that triples retries is caught. 4. **Fast rollback** — because the mapping is config, rollback is a config change, which is the main operational payoff of the whole design. ## When to drop the abstraction Be willing to say where the framework stops earning its keep. If a feature depends on one provider's distinctive capability — a specific caching mechanism, a bespoke response format, a batch API with different economics — pinning that call to the provider SDK or provider-specific class is honest engineering. The failure mode to avoid is the middle ground: pretending to be portable while every prompt and schema is tuned to one vendor, so the switch that the architecture promises has never been tested and would not work. Equally, portability is a means, not a goal. If you have one provider, one model and no plan to change, `init_chat_model` still helps by keeping construction in one place, but building a routing layer for a hypothetical migration is speculative complexity. ## Interview framing Open with the mechanism in one sentence, spend most of the answer on the leaks and the evaluation gate, and close with a clear statement of when you would deliberately couple to a provider. That sequence shows you understand that model selection is a cost, quality and risk decision that happens to be expressed in configuration — which is exactly the judgment a principal is being asked for.
- A team wants to re-point the 'strong' tier to a cheaper model. What do you require before approving?An offline run of the tier's evaluation suite showing no regression on correctness, structured-output validity and tool-call accuracy, then a canary on a small traffic share comparing the same metrics plus p95 latency and realized cost. Retry and parse-failure rates matter most: a cheaper model that fails schema validation more often can cost more than the one it replaced.
- Which behaviours are not equalized by the shared chat-model interface?Strict JSON-schema structured output, parallel and high-fidelity tool calling, whether streamed calls report token usage, reasoning-token behaviour and hidden output budgets, context-window size, and prompt sensitivity. The code compiles identically across providers, so these differences surface as quality and cost incidents rather than as errors — which is why a switch needs evaluation, not just a config edit.
- When would you deliberately bypass the framework and use a provider SDK directly?When a feature depends on a capability the abstraction does not express — a provider-specific caching or batch API with materially different economics, or a bespoke response format. Pin that one path explicitly and document it as non-portable, rather than degrading the abstraction for everyone or pretending the whole service is provider-agnostic when it is not.
- What is the risk of building a routing layer before you need one?Speculative complexity: a configurable indirection with one real configuration, tested only in the path you actually use. Untested portability is worse than none, because it invites the assumption that a switch is safe. Start with centralized construction via init_chat_model — cheap and useful immediately — and add routing when a second model has a genuine, evaluated purpose.
saying these in an interview costs you the question
- Assuming a config-only model swap is behaviour-neutral
- Scattering provider class imports across business logic
- Switching models without an evaluation suite
- Believing prompts transfer unchanged between model families
- Building multi-provider routing with no second provider in use