Taxonomy and Label Schema
The defined set of classes and rules used to assign labels consistently across a dataset.
A taxonomy is the formal structure behind any labeling task: the finite set of categories an item can be assigned, their definitions, and the rules for choosing between them when an item could plausibly fit more than one. The label schema is the technical encoding of that taxonomy — how it's represented in the data itself (flat categories, hierarchical trees, multi-label sets).
A well-designed taxonomy has categories that are mutually exclusive and collectively exhaustive for the domain it covers; a poorly designed one produces constant ambiguity because two categories overlap or a valid input has no home at all.
Taxonomies aren't static — as a dataset grows and new patterns appear, teams run schema migration to add, split, or merge categories, which requires re-labeling affected historical data to keep the dataset internally consistent.
What this means for trainers
When a taxonomy feels like it doesn't fit the data you're seeing, that's worth flagging upstream — it's often a sign the schema needs to evolve, not that you're applying it wrong.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs