[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"glossary-model-card::en":3,"gloss-cluster-model-card::en":26,"gloss-next-model-card::en":9},{"slug":4,"category":5,"name":6,"definition":7,"meta_desc":8,"faq":9,"schema_markup":9,"related":10},"model-card","mlops","Model Card","A model card is a short structured document that travels with a trained model and states what it is for, what it was trained and evaluated on, how it performs across relevant subgroups, and where it should not be used. The format was proposed to make the limits of a model as visible as its headline score, on the grounds that an aggregate accuracy number hides exactly the information a deployer needs: which populations, inputs or conditions the model was never tested on. A useful card is specific rather than reassuring. It names the intended use and the out-of-scope uses; describes training data provenance at least in kind, including known gaps; reports evaluation results broken down by the groups that matter for the application rather than as one number; lists known failure modes; and records the version, date and contact owner so a reader can tell whether it still describes the model in production. Cards that omit the negative sections are marketing documents with a technical layout. For teams building on third-party models, the card is one of the few artefacts that supports a real vendor assessment, and its absence or vagueness is itself a finding. For teams shipping their own models — including fine-tuned ones, which inherit the base model's limits and add their own — writing a card is cheap and pays off in procurement, security review and incident response, where the first question after a bad output is always what the model was supposed to be good at.","A model card documents a model's intended use, training data, subgroup results and known failure modes — the artefact that makes its limits as visible as its score.",null,[11,14,17,20,23],{"slug":12,"name":13},"benchmark-contamination","Benchmark Contamination",{"slug":15,"name":16},"golden-dataset","Golden Dataset",{"slug":18,"name":19},"llm-benchmark","LLM Benchmark",{"slug":21,"name":22},"model-registry","Model Registry",{"slug":24,"name":25},"trust-center","Trust Center",[27,31,35,38,41,45,48,51,54,57,60,63],{"slug":28,"category":5,"name":29,"updated_at":30},"annotation-guidelines","Annotation Guidelines","2026-08-24T03:30:02+00:00",{"slug":32,"category":5,"name":33,"updated_at":34},"baseline-model","Baseline Model","2026-08-24T02:46:38+00:00",{"slug":36,"category":5,"name":37,"updated_at":34},"batch-inference","Batch Inference",{"slug":39,"category":5,"name":40,"updated_at":34},"canary-prompt","Canary Prompt",{"slug":42,"category":5,"name":43,"updated_at":44},"champion-challenger","Champion-Challenger (A\u002FB Model Testing)","2026-08-24T02:46:37+00:00",{"slug":46,"category":5,"name":47,"updated_at":34},"class-imbalance","Class Imbalance",{"slug":49,"category":5,"name":50,"updated_at":34},"continuous-batching","Continuous Batching",{"slug":52,"category":5,"name":53,"updated_at":34},"cross-validation","Cross-Validation",{"slug":55,"category":5,"name":56,"updated_at":34},"data-labeling","Data Labeling",{"slug":58,"category":5,"name":59,"updated_at":44},"drift-detection","Drift Detection",{"slug":61,"category":5,"name":62,"updated_at":44},"eval-harness","Eval Harness",{"slug":64,"category":5,"name":65,"updated_at":44},"experiment-tracking","Experiment Tracking"]