Inter-Annotator Agreement (IAA)
A measure of how consistently different human raters assign the same labels or ratings to the same examples.
Inter-Annotator Agreement (IAA) is a statistical metric that measures how often multiple human raters make the exact same decision when evaluating the identical piece of data. In the context of AI alignment, IAA is the ultimate barometer for the clarity of annotation guidelines and the quality of the rating workforce.
If five different evaluators are asked to rank two model responses, and all five agree that Response A is better, the IAA is 100%. If the votes are split, the IAA drops. High IAA indicates that the evaluation task is objective, the guidelines are perfectly clear, and the raters are highly calibrated. Low IAA is a massive red flag. It suggests that the prompt is highly subjective, the rubric is confusing, or the raters are relying on their personal biases rather than the project rules.
AI labs aggressively monitor IAA. They frequently inject 'benchmark' or 'honeypot' tasks—questions where the correct answer has already been determined by senior researchers—into a rater's queue. If a rater consistently disagrees with the consensus or the benchmark, their data is discarded, and they are typically removed from the project.
Achieving high IAA on complex reasoning, coding, and safety tasks is incredibly difficult but absolutely necessary. A reward model cannot learn human preferences if the humans themselves cannot agree on what they prefer.
What this means for trainers
To maintain high IAA and keep your position on premium projects, you must shed your personal biases. Whenever you encounter an ambiguous response, don't ask 'What do I think is best?' ask 'What do the guidelines dictate is best?' Calibrating your mindset to the project's consensus is the secret to longevity in this industry.
Related terms
Related guides
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs