Misinformation Labeling
Flagging claims that are unsupported, deceptive, or manipulated.
Misinformation labeling identifies content that makes false or misleading factual claims, distinguishing between honest error, deliberate deception, and manipulated media (see also fact-checking for the claim-verification mechanics underlying this task). Unlike straightforward fact-checking, misinformation labeling often needs to account for intent and framing — a technically true statement presented in a deliberately misleading way can still warrant a misinformation flag.
This task has grown alongside synthetic and AI-generated media, where manipulated images, audio, or video add a layer beyond textual claim-checking — annotators increasingly need to assess whether content itself has been fabricated or altered, not just whether the claims within it are accurate.
Misinformation datasets feed both platform-level content moderation and the safety policy enforcement training used to keep models from generating or amplifying false claims.
What this means for trainers
Distinguish clearly between "this is false" and "this is deceptively framed" in your reasoning — guidelines usually treat these as different severity tiers, and conflating them is a common source of inconsistent labels.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs