Preference Ranking
Comparing two or more model outputs and selecting the better one according to a defined rubric.
Preference ranking is the task-level mechanics behind preference labeling: given two (or more) candidate responses to the same prompt, an annotator applies a rubric — usually covering accuracy, helpfulness, safety, and tone — and produces an ordering rather than an isolated score.
When only two items are compared, this is called pairwise ranking; larger candidate sets require more careful, often multi-pass comparison since human judgment gets noisier past a handful of options at once.
The resulting rankings feed directly into reward model training for RLHF and DPO pipelines — this is one of the most common and highest-volume annotation task types on modern AI training platforms.
What this means for trainers
Consistency across sessions matters more than any single judgment call — reviewers track whether your rankings of similar pairs stay stable over time (see ranking consistency), and drifting standards show up as quality flags even when each individual choice seems reasonable.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs