Skip to content
aitrainer.work - AI Training Jobs Platform
Taxonomy and Label Schema definition
annotation core concepts

Taxonomy and Label Schema

The defined set of classes and rules used to assign labels consistently across a dataset.

What this means for trainers

When a taxonomy feels like it doesn't fit the data you're seeing, that's worth flagging upstream — it's often a sign the schema needs to evolve, not that you're applying it wrong.

A taxonomy is the formal structure behind any labeling task: the finite set of categories an item can be assigned, their definitions, and the rules for choosing between them when an item could plausibly fit more than one. The label schema is the technical encoding of that taxonomy, meaning how it is represented in the data itself, whether as flat categories, hierarchical trees, or multi-label sets.

A well-designed taxonomy has categories that are mutually exclusive and collectively exhaustive for the domain it covers. A poorly designed one produces constant ambiguity because two categories overlap or a valid input has no home at all in the existing structure.

Taxonomies are not static. As a dataset grows and new patterns appear, teams run schema migration to add, split, or merge categories, which requires re-labeling affected historical data to keep the dataset internally consistent. A category added late still has to be applied retroactively wherever it belongs, or the dataset ends up with a silent gap for everything labeled before the change.

Related terms

Put this into practice

Browse open data annotation and evaluation roles from Mercor, Micro1, Outlier, and more.

Browse AI training jobs