Home/Blog/What a model card should say and usually does not

What a model card should say and usually does not

A model card was meant to be an evidentiary document: what a model was trained on, where it fails, how it was measured. Regulation has now made most of it mandatory, and the gap between the template and its honest use is the whole subject.

The model card was proposed in 2018 as a short document accompanying a trained model: what it is for, what it was trained on, how it performs across different conditions, and where it fails.

The idea was modest and the intent was specific. A model card is an evidentiary artifact, not a marketing one. It exists so that someone deciding whether to use a model can find out what it does and does not do before deploying it, rather than after an incident.

Eight years later most of it is law. The EU AI Act requires providers of high-risk systems to document intended use, training-data summary, evaluation across relevant subgroups, known limitations, and version history, with penalties reaching tens of millions of euros or a share of global turnover. Enforcement began in August 2026. US state law has created parallel obligations for high-risk uses in employment, housing, healthcare and lending.

A model card that satisfies the template can still fail at the thing the template was for. Compliance audits repeatedly find cards that are compliant in form and non-compliant in substance, and the difference is exactly the difference between documenting that you evaluated a model and documenting what you found. The fields that matter most are the ones a template cannot force you to fill honestly.

What belongs in one, and what each field is actually for

The regulated template converges on a stable set of sections. Worth reading each as the question it answers rather than the label it carries, because the label is what gets filled and the question is what gets skipped.

Intended use. Not what the model can do, but what it is for and, critically, what it is not for. The most useful sentence in this section is the one ruling something out, and it is the one most often absent because ruling a use out closes a door someone wanted open.

Training data summary. What it was trained on, how current it is, and where it has representation gaps. A good version reads like a limitation: trained on this market, this period, these categories, with named gaps. A useless version says the data was large and diverse, which is true of everything and informative about nothing.

Evaluation, disaggregated. Performance across relevant subgroups rather than a single aggregate. The aggregate is what a model card cannot get away with any longer, because an average that looks acceptable can hide a subgroup where the model is unusable, and the entire reason for disaggregation is to surface exactly what the average conceals.

Known limitations. Where it fails. This is the section that separates an evidentiary document from a brochure, and its length is the fastest read on whether anyone was honest. A limitations section under a few sentences is reliable evidence that nobody looked hard, which is a finding in itself.

Human oversight. Who monitors it, on what cadence, and how someone intervenes when it is wrong, including how fast it can be stopped. The specific question that exposes a hollow card: what is the kill switch and how long does it take. If that answer does not exist, the oversight is aspirational.

Version history. What changed and when. Boring, and it is what an incident investigator needs to reconstruct what the system was doing at the time, which is the moment a version history stops being bureaucracy.

The two things a template cannot force

Everything above can be filled to the letter and still fail, because two properties are not fields and cannot be made into fields.

Whether the limitations are real. A limitations section can be written to satisfy an auditor or to warn a user, and these produce different documents. The satisfying version lists generic caveats that apply to any model: may reflect biases in training data, should not be used as sole basis for decisions, performance may vary. The warning version names the specific inputs on which this specific model fails. Only the second is useful, and no template distinguishes them, because both fill the same box.

Whether the evaluation was adversarial. Disaggregated metrics can report the subgroups you chose to measure, which are frequently the ones you expected to pass. An honest evaluation includes the subgroups you feared, the edge cases you found during development, and the failures you would rather not publish. The template asks for disaggregation and cannot ask whether you disaggregated along the axis where the problem actually is.

This is the substance-versus-form gap the compliance audits keep finding. The card is complete. Every field is populated. And it documents that the model was evaluated without documenting what the evaluation should have been afraid of.

Model card, datasheet, system card

Three documents get conflated and they answer different questions, which matters because a requirement for one is often met with another.

A model card documents the trained model: architecture, performance, intended use, limitations.

A datasheet documents the training data: where it came from, how it was collected, who is in it, what preprocessing was applied, what it may not be used for. This is a separate artifact proposed separately, and it answers the provenance question a model card only summarises. Building one model can touch dozens of data sources, and the datasheet is where their origin and legal usability are established.

A system card documents the deployed system: the model plus its guardrails, its oversight, its disclosure to users, and its behaviour in context. A model can be documented perfectly and deployed inside a system that changes everything about how it behaves.

The regulatory stack for general-purpose models now expects several of these layers together: internal technical documentation, a public model card, a downstream deployer package, and a published training-data summary. A requirement satisfied with the wrong layer is a common and expensive form of non-compliance in substance.

The provenance problem underneath

The training-data section is where the hardest gap sits, and it is not primarily a documentation problem.

Data provenance is proof of origin: where data came from, who collected it, and whether it is legally usable. It is distinct from data lineage, which maps how data flows through systems, and it connects directly to where that data is allowed to live. You can have perfect lineage, a complete map of every transformation, and no provenance, no proof that the data at the start was collected lawfully and licensed for this use.

The regulatory direction is toward machine-readable provenance for generated content, and the practical difficulty is that provenance has to be captured at collection time. A model trained on data whose origin was not recorded cannot have that origin reconstructed afterward, which means the gap in an existing model's card is frequently unfixable rather than merely unfilled. The honest entry in that case is that provenance was not tracked, which is itself the disclosure.

Why cards are written badly even when they are written

The failures are structural rather than negligent, and naming them explains why mandating the document does not by itself fix the problem.

Honesty is competitively costly. A card that names specific failure modes hands a competitor a list of where to attack and a customer a reason to hesitate. The incentive is to write the generic caveat that satisfies the requirement and reveals nothing, and no field on the form counteracts that incentive.

The card is written by the people who built the model. They evaluated along the axes they thought to check, and the axes they did not think to check are exactly the ones missing from the card, invisibly, because you cannot document a blind spot you have.

Intellectual property cuts against disclosure. Training-data curation and architecture are treated as competitive advantages, and there is a real and unresolved tension between transparency and protecting them. A card can be honest about capabilities and limitations while disclosing nothing about how the model was built, which is where most of them settle.

And the deadline was the driver. Much of the current documentation exists because enforcement began, which means it was produced to a compliance schedule rather than a diagnostic one. A document written to be audit-ready optimises for the auditor, and the auditor checks that fields are populated, not that they are true.

How to read one, and how to write one

As a reader. Skip to limitations first, and measure its length and specificity against the capability claims. A long, specific limitations section on a modestly-claimed model is a good sign. A short, generic one on an ambitiously-claimed model is the tell. Then check whether the evaluation is disaggregated along an axis that could actually fail, rather than along convenient ones. Then look for the kill switch. If oversight is described without a stopping mechanism and a time, it is aspirational.

As a writer. The card is worth more than its compliance value, and this is the argument for taking it seriously beyond the fine. Writing down what a model does and does not do, before deployment rather than after an incident, forces a precision that surfaces problems while they are still cheap. One documented case: a lender building its first card discovered it had no record of its training-data cutoff, a gap found during writing and fixed before the audit rather than during an incident.

The discipline that produces an honest card is to write the limitations section first, from the failures found during development, before writing anything about capability. A card built in that order documents what you were afraid of. A card built in the usual order documents what you were proud of, and adds caveats at the end to satisfy the form.

What is unresolved

Whether mandated documentation improves outcomes or produces theatre. Making cards compulsory guarantees they exist and does not guarantee they are honest, and the substance-versus-form gap suggests compliance can be satisfied without the transparency the requirement was for. Whether enforcement evolves to evaluate substance, or settles for form, will determine which of these the regime produces.

How to document a model nobody understands. A model card assumes the authors can state how the model behaves. For large models whose failure modes are discovered continuously after release, the card is out of date the moment it is published, and there is no established practice for a living document that keeps pace.

Whether provenance can be required retroactively. Data whose origin was not recorded cannot have it reconstructed, so a provenance requirement is a requirement to have started years ago. What standard applies to models already trained on data of unknown origin is unsettled, and the honest answer of "not tracked" sits awkwardly with a regime that expects the record to exist.

Whether transparency and IP can be reconciled. The tension between disclosing enough to be useful and protecting what is treated as competitive advantage has no clean resolution, and where the line falls is being negotiated case by case rather than settled in principle.

The counter-argument

Mandated cards are better than no cards. Before regulation, most models shipped with nothing, and a compliant-in-form card still forces some documentation and creates an artifact that can be audited and improved. The critique that cards can be hollow should not slide into the position that requiring them was pointless.

The generic caveats are sometimes correct. "May reflect biases in training data" is vague and it is also true and worth stating. Dismissing all standard caveats as evasion risks discarding genuine and general limitations that a reader does need to know, even if they apply broadly.

Honesty has real costs that are not just cover for evasion. A company that fully documented every failure mode of its model would hand advantages to competitors and ammunition to litigants, and pretending this cost is illusory is not fair to the people writing the cards. The tension between useful disclosure and self-protection is real on both sides.

And the deadline achieved something. Documentation produced to a compliance schedule is imperfect and it exists, which is more than existed before. A diagnostic culture may grow from a compliance requirement, and the current imperfect cards are the substrate it would grow from.

The short version

The model card was proposed in 2018 as an evidentiary document: what a model is for, what it was trained on, how it performs across conditions, and where it fails. Regulation has since made most of it mandatory, with the EU AI Act requiring intended use, training-data summary, disaggregated evaluation, known limitations and version history for high-risk systems, enforced from August 2026 with substantial penalties, and parallel US state obligations.

The template converges on a stable set of sections, each best read as the question it answers: intended use including what the model is not for, training data including its gaps, evaluation disaggregated across subgroups rather than averaged, limitations as the section whose length reveals whether anyone was honest, human oversight including the kill switch and its timing, and version history as what an incident investigator will need.

Two properties cannot be made into fields and are where cards fail. Whether the limitations are real, since a generic caveat and a specific warning fill the same box and only the second is useful. And whether the evaluation was adversarial, since disaggregation can report the subgroups you expected to pass rather than the ones you feared. This is the substance-versus-form gap that compliance audits keep finding: complete cards that document evaluation without documenting what the evaluation should have feared.

Model cards, datasheets and system cards answer different questions, and a requirement for one is often met with another. Underneath the training-data section sits provenance, proof of origin distinct from lineage, which must be captured at collection time and cannot be reconstructed afterward, making many existing gaps unfixable rather than merely unfilled.

The discipline that produces an honest card is to write the limitations first, from the failures found during development, before writing anything about capability. A card built that way documents what you were afraid of; a card built in the usual order documents what you were proud of and adds caveats to satisfy the form. No template can force the first, which is why the fields that matter most are the ones it cannot make you fill honestly.

Common questions

What is a model card? A short document accompanying a trained model that states its intended use, what it was trained on, how it performs across different conditions, and where it fails. Proposed in 2018 as an evidentiary artifact rather than a marketing one, it exists so that someone deciding whether to deploy a model can learn what it does and does not do beforehand. Most of its contents are now legally required for high-risk systems.

What should a model card include? Intended use, including what the model is explicitly not for. A training-data summary with its currency and representation gaps. Evaluation disaggregated across relevant subgroups rather than a single average. Known limitations describing specific failure modes. Human oversight covering who monitors it, on what cadence, and how it is stopped, including the kill switch and its timing. And version history, which an incident investigator needs to reconstruct what the system was doing. Regulation now mandates most of these for high-risk uses.

What is the difference between a model card and a datasheet? A model card documents the trained model: its architecture, performance, intended use and limitations. A datasheet documents the training data: where it came from, how it was collected, who is represented, what preprocessing was applied, and what it may not be used for. They answer different questions, and a system card is a third artifact documenting the deployed system including its guardrails and oversight. A requirement for one is frequently and mistakenly satisfied with another.

Why do model cards fail even when they are complete? Because two things cannot be made into fields. Whether the limitations are real, since a generic caveat that applies to any model and a specific warning about this model fill the same box, and only the second is useful. And whether the evaluation was adversarial, since disaggregated metrics can report the subgroups you expected to pass rather than the ones you feared. Compliance audits repeatedly find cards that are complete in form and non-compliant in substance for exactly this reason.

Are model cards legally required? For high-risk systems, yes, in a growing number of jurisdictions. The EU AI Act requires documentation of intended use, training data, disaggregated evaluation, limitations and version history for high-risk AI, enforceable from August 2026, with penalties reaching tens of millions of euros or a share of global turnover. US state laws including Colorado's create parallel obligations for high-risk uses in employment, housing, healthcare and lending. Enterprise buyers now demand the documentation in procurement regardless of jurisdiction.

What is data provenance and why does it matter for model cards? Provenance is proof of origin: where training data came from, who collected it, and whether it is legally usable. It differs from lineage, which maps how data moves through systems. You can have complete lineage and no provenance. It matters because it must be captured at collection time and cannot be reconstructed afterward, so a model trained on data of unrecorded origin has a gap that is unfixable rather than merely unfilled, and the honest card entry in that case is that provenance was not tracked.

How do I read a model card critically? Go to the limitations section first and weigh its length and specificity against the capability claims. A long, specific limitations section on a modestly-claimed model is a good sign; a short, generic one on an ambitiously-claimed model is the tell that nobody looked hard. Check whether evaluation is disaggregated along an axis that could fail rather than convenient ones. Then find the kill switch: oversight described without a stopping mechanism and a time is aspirational.

How do I write a good model card? Write the limitations section first, from the failures found during development, before writing anything about capability. A card built in that order documents what you were afraid of; the usual order documents what you were proud of and appends caveats to satisfy the form. Beyond compliance, the exercise forces precision about what the model does and does not do before deployment rather than after an incident, which is where its real value is: one lender discovered an undocumented training-data cutoff while writing its first card and fixed it before the audit.

Learn the concepts

← All posts