Confidence Scoring
A rating an annotator provides to indicate how certain they are about their own labeling decision.
What this means for trainers
Report confidence honestly rather than defaulting to "high" — inflated confidence on genuinely uncertain items removes the signal reviewers rely on to catch real problems in the data.
Confidence scoring asks annotators to attach a certainty level, often a simple scale like low, medium, or high, to each label they produce, separate from the label itself. Two annotators might agree on a label but disagree sharply on how confident they are in it, and that gap is useful signal on its own.
Low-confidence labels are natural candidates for adjudication or additional review, and aggregated confidence scores across a dataset help identify systemic edge cases that the guidelines do not cover well. A cluster of low-confidence labels in one area usually points to a taxonomy gap rather than to weak individual annotators.
Confidence scoring is also the human-side analog of model uncertainty sampling, which uses a model's own low-confidence predictions to prioritize what gets sent to humans for labeling. Both are built on the same idea: knowing how sure a judgment is can matter as much as the judgment itself.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs