Class Imbalance
A dataset condition where some labels appear far more often than others.
What this means for trainers
If you notice you're rarely seeing examples of a category you know exists, that imbalance is worth surfacing — targeted sourcing of rare cases is often a separate, sometimes better-paid task.
Class imbalance describes a dataset where the distribution of labels is skewed, so one category might make up the large majority of examples while another makes up a tiny fraction. This is common and often unavoidable, since real-world data on harmful content, rare defects, or unusual user intents is naturally scarce compared to common cases.
The practical consequence is that models trained on imbalanced data tend to perform well on frequent classes and poorly on rare ones, since there is simply less signal to learn from for the underrepresented category. Teams address this through targeted collection of underrepresented examples, hard negative mining, or reweighting during training, but the annotation side of the fix starts with schema coverage analysis to identify which classes are underrepresented in the first place.
Class imbalance also complicates evaluation. Raw accuracy on an imbalanced dataset can look high while performance on the rare but important class is poor, which is why precision and recall broken out per class matter more than a single aggregate score.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs