Skip to content
aitrainer.work - AI Training Jobs Platform
Hate Speech Taxonomy definition
safety and alignment

Hate Speech Taxonomy

A defined set of classes and scope rules for labeling abuse targeted at protected groups.

A hate speech taxonomy is a specialized label schema that defines which protected characteristics (race, religion, gender, and similar) are in scope, what counts as targeting versus general discussion of a group, and how severity is tiered from coded language to explicit calls for harm.

This is narrower than general toxicity annotation: toxicity covers abuse of any kind, while a hate speech taxonomy specifically requires a protected-group target, which means annotators need clear guidance on distinguishing genuine hate speech from adjacent categories like general insults, political criticism, or reclaimed language used within a community.

Because legal and cultural definitions of hate speech vary significantly by region and platform, taxonomies in this space tend to be more detailed and more frequently revised than most other content categories.

What this means for trainers

This is one of the categories where regional and cultural context matters most — guidelines here are usually more detailed than average specifically because the boundary cases are genuinely contentious.

Related terms

Put this into practice

Browse open AI training roles from Alignerr, Mercor, Outlier, and more.

Browse AI training jobs