Calibration
The process of aligning annotators on a shared interpretation of the guidelines before and during production work.
What this means for trainers
Calibration rounds are usually unpaid or low-paid onboarding steps, but performing well in them is often what unlocks access to the better-paying production queue.
Calibration is what happens before a large annotation effort starts producing real data. A small group of annotators labels the same sample set, compares results, discusses disagreements, and converges on a shared interpretation of the guidelines. Only once agreement on this sample reaches an acceptable threshold does the team move to full production.
Calibration is not a one-time event. Guidelines evolve, new edge cases surface, and annotator populations turn over, so mature pipelines re-run calibration periodically to catch drift before it degrades the dataset. A guideline that everyone agreed on in January can be interpreted differently by a new cohort of annotators in June.
Without calibration, the first batch of production labels effectively becomes an uncontrolled experiment in how differently people read the same instructions. That is why platforms treat calibration performance as a gate rather than a formality, even when the round itself is unpaid or low-paid.
Related terms
Related guides
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs