Refusal Quality
Evaluating whether a model declines an unsafe request clearly, safely, and without being needlessly restrictive.
What this means for trainers
Watch specifically for over-refusal, not just under-refusal — annotators new to this task tend to reward caution unconditionally, but a model refusing a clearly benign request is a real quality failure too.
Refusal quality judges a model's decision to decline a request along multiple dimensions at once: did it correctly identify the request as one it should refuse, did it explain the refusal clearly rather than curtly, and did it avoid refusing a request that was legitimate but merely resembled an unsafe one.
This last point is what makes the task genuinely hard. A model that refuses too eagerly, known as over-refusal, is graded as a failure just like one that complies with something it should not, since excessive caution makes the assistant less useful without making it meaningfully safer.
Refusal quality is the natural complement to jailbreak detection and safety policy enforcement. Those tasks identify what should be refused, while this one grades how well the model executes that refusal in practice, feeding into policy-compliant refusal writing standards used to train future responses.
Related terms
Related guides
Put this into practice
Browse open red-teaming, safety evaluation, and model-alignment roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs