skip to content

Can you ship Mistral's Codestral weights inside a commercial product?

level: seniorimportance: should knowfreq 38%

answer

  1. not every Mistral model is permissive
  2. the specialist coder has its own licence
  3. dev and test yes, production no
  4. internal use is still commercial
  5. hosted API sidesteps the weight licence

basics

~20 s

Not from the open weights alone. Codestral's downloadable weights carry the Mistral Non-Production License, which permits development, testing and research but excludes production and commercial use. Commercial paths are a paid Mistral licence, the hosted Codestral API, or an Apache-2.0 model instead.

solid answer

~50 s

Codestral is Mistral's code-specialist model, and its open weights ship under the **Mistral Non-Production License (MNPL)** rather than Apache 2.0. The MNPL lets you download, run, modify and evaluate the model for development, testing and research; it explicitly excludes production deployment and commercial use, including internal commercial use. So self-hosting Codestral behind a paid IDE plugin or a customer-facing completion service is a licence breach, not a grey area. There are three lawful routes. **Buy a commercial weight licence** from Mistral if you need it on your own hardware. **Call the hosted Codestral API** on la Plateforme, where use is governed by Mistral's commercial terms rather than the weight licence — that is the normal path. Or **pick an Apache-2.0 model** and give up the specialist edge: the Mistral Small line and Mistral's Apache-2.0 coding/agentic releases such as Devstral Small are strong code models with no such restriction. Also remember derivatives: a fine-tune or quantised build of Codestral inherits the MNPL.

go deeper

for a junior

Know that some Mistral models are freely usable and others are not, and that you check the specific release's LICENSE file before putting any downloaded model into a product.

for a middle

Name the Mistral Non-Production License, state that it permits development and testing but not production or commercial use, and know that the hosted API is governed by different terms.

for a senior

Lay out the three lawful paths with tradeoffs, flag that fine-tunes and quantisations inherit the restriction, and put the licence check at the dev-to-production promotion boundary.

for a principal

Weigh the total cost: a negotiated weight licence carries procurement lead time and ongoing audit burden, so an Apache-2.0 alternative that is slightly weaker on benchmarks is often the cheaper organisational decision.

## What the MNPL actually says The Mistral Non-Production License is a source-available licence written specifically so developers can get hands-on with a model before paying for it. It grants a broad right to use, reproduce, modify and create derivatives of the weights **for development, testing, evaluation and research**, and it withholds the right to use the model or its derivatives in production or in connection with commercial activity. Mistral offers a separate commercial licence for anyone who needs the withheld rights. The practical reading is: your laptop, your CI, your internal benchmark harness, your prototype demo — fine. Anything that serves real users or supports revenue-generating work — not fine, whether the users are external customers or your own engineers. ## Why teams get this wrong Three failure modes recur. **"It's on Hugging Face, so it's open source."** Restricted and permissive weights sit side by side on the hub, download through the same API call, and load with the same code. Nothing in the tooling enforces the licence. **"It's internal, so it isn't commercial."** An internal developer-productivity tool at a for-profit company is commercial activity. "Non-production" is not a synonym for "not customer-facing". **"We fine-tuned it on our own code, so it's ours now."** A fine-tune is a derivative work. So is a merged model, and so is a GGUF or AWQ quantisation someone else uploaded — third-party quantisations frequently carry a sloppy or simply wrong licence tag while the underlying restriction persists. ## The three lawful paths, and how to choose **1. Hosted API.** Calling Codestral through Mistral's platform moves you out of weight-licence territory entirely: hosted use is governed by Mistral's commercial terms of service, which permit production use on a paid plan. This is the intended commercial route and by far the least friction. The tradeoffs are the usual hosted ones — per-token cost, network latency on a workload (code completion) that is exquisitely latency-sensitive, and sending source code to a third party, which is precisely the objection that pushes some teams toward self-hosting in the first place. **2. Negotiated commercial weight licence.** If your constraint is that code must never leave your network — regulated industry, defence, strict IP policy — this is the path that gets Codestral onto your own GPUs lawfully. Budget procurement lead time; this is a contract, not a checkbox. **3. An Apache-2.0 alternative.** Mistral's permissive catalogue includes strong code-capable models — the Mistral Small line and Apache-2.0 coding/agentic releases such as Devstral Small — that you may self-host commercially with no agreement at all. You give up whatever specialist edge Codestral holds on fill-in-the-middle completion quality; often the gap is small enough that the licensing simplicity wins outright. This is the option senior engineers should raise first, because it removes a whole class of ongoing compliance work. ## Running the compliance process For any self-hosted model, not just Codestral: - Record the exact release you deployed and archive its LICENSE file with the deployment, not just a link. - Keep a per-model licence inventory alongside your dependency inventory, and classify each entry as permissive / research-only / non-production / API-only. - Gate on the boundary between environments: a model that is legitimately in your dev cluster may be unlawful in prod, so the promotion pipeline is where the check belongs. - Re-verify on upgrade. Mistral has changed licences between generations, so last year's classification is not evidence about this year's release. - Treat derivatives — fine-tunes, merges, quantisations, distillations — as inheriting the base terms, and record the base model in the artefact's metadata. ## What a strong answer sounds like A weak answer is "Codestral is open, so yes" — it fails on the licence and on the habit of checking. A strong answer names the MNPL, states the specific restriction (non-production, and internal commercial use counts), lists the three lawful paths with their tradeoffs, mentions that derivatives inherit the restriction, and points out that the Apache-2.0 alternative usually costs less total effort than the compliance work of the restricted one.

  • Your team fine-tunes Codestral on your internal codebase. Does the resulting adapter escape the restriction?
    No. A fine-tune or LoRA adapter is a derivative work of the base weights and inherits the Mistral Non-Production License, no matter how much proprietary data went into training it. The adapter is also useless without the restricted base, so the base licence governs the deployment either way. Fine-tune an Apache-2.0 base if you need to ship the result.
  • How does calling Codestral through Mistral's hosted API change the legal position?
    Completely. The weight licence governs weights you download and run; hosted inference is governed by Mistral's commercial terms of service, which permit production use on a paid plan. That is the intended commercial route for teams that do not want to negotiate a weight licence. What you trade away is network latency on a latency-critical completion workload, per-token cost, and sending source code off your network.
  • What would you propose if legal blocks both a paid licence and sending code off-network?
    Move to an Apache-2.0 model you can self-host without any agreement — the Mistral Small line or Mistral's permissively licensed coding/agentic releases such as Devstral Small. You lose whatever specialist completion edge Codestral holds, which is worth measuring on your own repositories rather than assuming, but you remove an entire category of ongoing compliance risk and procurement delay.

saying these in an interview costs you the question

  • Assuming any downloadable Mistral model is safe to ship
  • Reading non-production as merely not customer-facing
  • Believing a fine-tune resets the base model's licence
  • Trusting a third-party quantisation's licence tag on the hub
  • Confusing the hosted API's terms with the weight licence

context