Pairwise Ranking
Comparing exactly two candidate outputs and choosing the better one.
What this means for trainers
Resist the urge to rate absolute quality in your head — the task is comparative: which of these two is better, not how good is each one on its own.
Pairwise ranking is the simplest and most common form of preference ranking: two responses to the same prompt, one judgment call for which is better. Its simplicity is the point. Humans are much more reliable at comparing two things directly than at scoring items independently on an absolute scale, since a comparative judgment does not require agreeing on what a given numeric score means in the abstract.
Large-scale preference datasets for RLHF and DPO are built almost entirely from pairwise comparisons rather than absolute scores, precisely because pairwise judgments hold up more consistently across annotators and over time than a scale-based rating does.
When more than two candidates need ranking, systems typically decompose the problem into multiple pairwise comparisons and aggregate them, for example through a tournament or an Elo-style system, rather than asking annotators to sort a full list at once. Sorting more than a handful of items directly tends to produce noisier, less reliable results than a series of simple comparisons.
Related terms
Put this into practice
Browse open RLHF, fine-tuning, and preference-labeling roles from Mercor, Alignerr, and more.
Browse AI training jobs