What does a model card record about a model's intended use and training population?
answer
- the label, not the chemistry
- written for the people downstream
- who was in the training data, with counts
- what it must NOT be used for
- versioned with the model, gate before launch
basics
~20 sA model card is a short document shipped with a model stating what it is for, what it must not be used for, who and what the training data represented, how it performs on relevant subgroups, and its known limits.
solid answer
~50 sA model card is the model's label. The sections that carry the weight are: **intended use** — the decision it supports and the users it supports, stated narrowly; **out-of-scope use** — what it must not be used for; **training and evaluation data** — where the data came from and who is represented in it; **performance**, reported overall and broken out by the slices that matter; and **known limitations**. For a dermatology triage model, that reads as: intended use is prioritising which cases a clinician reviews first, not producing a diagnosis; the training population's skin-tone distribution is stated with counts; and the known-limitations section records the measured sensitivity drop on tones that are under-represented. The card is written before launch, versioned with the model, and reviewed as part of release — not assembled after somebody misuses the thing.
go deeper
Be able to name the main sections on demand: intended use, out-of-scope use, training and evaluation data, performance including slices, known limitations, and ownership.
Explain why the training population belongs in the card at all — every reported metric is conditional on it, and a downstream team cannot judge transfer to their own population without it.
Demonstrate that you would publish measured slice performance with sample sizes, including the unflattering numbers, and keep the card versioned with the model so a retrain forces a revision.
Own the card as a release gate rather than a document: who signs it, what an incomplete limitations section blocks, and how the out-of-scope list is agreed before launch instead of after the first misuse.
## What a model card is A model card is a short, structured document that travels with a trained model, introduced by Mitchell and colleagues in 2019 as a reporting convention. Think of it as the label on a medication: not the chemistry, but what it treats, who it was tested on, what it must not be combined with, and what side effects were observed. Its audience is not the team that built the model — they have the notebooks. Its audience is everyone downstream: the product team wiring it into a workflow, the clinician or underwriter relying on it, the reviewer who has to approve the release, and the engineer who inherits it in two years. ## The sections that carry weight **Model details.** Name, version, date, owning team, contact, model family at a level a non-specialist can parse. Enough to know which artefact this card describes. **Intended use.** The decision the model supports, the users it supports, and the deployment context — stated narrowly and concretely. `Triages incoming dermatology referral photographs so that clinicians review higher-risk cases first` is a usable statement. `Skin lesion classification` is not: it does not say what happens to the output or who acts on it. **Out-of-scope use.** The explicit list of things the model must not be used for. For the same model: not a diagnosis, not a substitute for clinician review, not for body sites absent from the training data, not for patient populations outside the ones evaluated. This section is discussed further below. **Training and evaluation data.** Where the data came from, over what period, how it was collected and labelled, and — critically — who is in it. For an image triage model that means the distribution of skin tones, ages and capture devices, with counts, not adjectives. `Diverse dataset` is a claim nobody can check. A table showing that one tone band holds 3% of images is a fact a downstream reader can reason about. **Performance.** Headline metrics, and then a breakdown across the slices that matter for the use case, each with its sample size so the reader can judge how much to trust the number. An overall figure alone hides exactly the failures the card exists to surface. **Known limitations and ethical considerations.** Where the model is weakest, what happens when it is wrong in each direction, and what mitigations are in place — for a triage model, that a missed high-risk case is far more costly than an over-prioritised benign one, and what the fallback path is. **Maintenance.** Who owns it, when it is reviewed, how to report a problem. ## Why the training population is the load-bearing section Everything a model card says about performance is conditional on the population it was measured on. A downstream team deploying into a clinic whose patient mix differs from the training mix cannot infer the headline number applies to them — but they can only notice that if the card tells them what the training mix was. Recording the skin-tone distribution, with counts, converts an invisible assumption into a stated precondition. When the card also carries the measured drop on under-represented tones, the receiving team can decide whether to deploy, to restrict, or to collect more data first. Without those two facts, that decision is made in the dark and usually made optimistically. Stating the gap does not fix it. It does something more modest and more important: it moves the gap from something discovered in production into something weighed at the decision point. ## Writing the out-of-scope section before launch The hardest section to write honestly is what the model must not be used for, and the temptation is to defer it. Written before launch, it is a design constraint: it forces the team to say out loud that a triage model is not a diagnostic model, and it gives a reviewer something concrete to approve. Written after the first misuse, it is an incident report with a document attached. A workable habit is to write out-of-scope use at the same time as intended use, by asking three questions. What is the nearest neighbouring decision someone might plausibly point this at? Which populations or conditions did we never evaluate? What would a user assume the output means, that it does not mean? A triage score reported as a percentage will be read as a probability of malignancy by somebody unless the card says otherwise. ## Common failure modes - **A card that reads like a training log** — learning rates, architecture trivia, experiment identifiers. Those belong in internal documentation; the card is about use and limits. - **Adjectives where counts belong** — `large and diverse` instead of a distribution table. - **Overall metrics only**, with no slice breakdown and no sample sizes. - **An empty or generic limitations section**, which reads as either dishonesty or inattention. - **A card written once and never versioned**, so it describes a model that has since been retrained on different data. ## What the interviewer is checking That you can name the sections without prompting, that you understand intended use and training population are the two that constrain downstream deployment, and that you treat the card as a release gate written before launch rather than documentation produced afterwards.
- Why does a model card state what the model must not be used for?Because the likeliest harm is a plausible neighbouring use, not a bizarre one — a triage model read as a diagnosis, or applied to a population never evaluated. Naming those exclusions before launch turns them into a design constraint a reviewer can approve against. Written after the first misuse, the section is just an incident report.
- The card records a measured performance drop on an under-represented group. Doesn't publishing that invite criticism?It invites informed decisions. The gap exists whether or not it is written down; the card only controls whether the deploying team learns about it before or after patients are affected. Stating it with counts lets them restrict the deployment, add a human check, or fund more data. Silence converts a known limitation into a surprise.
- How is a model card different from a dataset datasheet?A datasheet documents a dataset — how it was collected, labelled, and what it should and should not be used to train. A model card documents a trained artefact — its intended use, evaluated population, slice performance and limits. They overlap on data provenance, and a good card cites the datasheet rather than restating it.
- When should a model card be written and updated?Drafted alongside the evaluation plan, finished before the release review, and treated as a gate for it. It is versioned with the model: any retrain that changes the training population, the evaluated slices or the measured performance invalidates the card and requires a revision. A card describing a superseded model is worse than none, because it is trusted.
A model card is the label on a medication: what it treats, who it was tested on, what it must not be combined with, and the side effects observed.
saying these in an interview costs you the question
- Fills the card with architecture and hyperparameter trivia
- Describes training data as diverse without any counts
- Reports only overall accuracy with no slice breakdown
- Leaves known limitations blank or generic
- Writes the card after launch as documentation cleanup
- Never revises the card when the model is retrained
- States intended use so broadly it excludes nothing