Skip to content
aitrainer.work - AI Training Jobs Platform
Refusal Quality definition
safety and alignment evaluation

Refusal Quality

Evaluating whether a model declines an unsafe request clearly, safely, and without being needlessly restrictive.

Refusal quality judges a model's decision to decline a request along multiple dimensions at once: did it correctly identify the request as one it should refuse, did it explain the refusal clearly rather than curtly, and — critically — did it avoid refusing a request that was actually legitimate but merely resembled an unsafe one.

This last point is what makes the task genuinely hard: a model that refuses too eagerly (over-refusal) is graded as a failure just like one that complies with something it shouldn't, since excessive caution makes the assistant less useful without making it meaningfully safer.

Refusal quality is the natural complement to jailbreak detection and safety policy enforcement — those tasks identify what should be refused, while this one grades how well the model actually executes that refusal in practice, feeding into policy-compliant refusal writing standards.

What this means for trainers

Watch specifically for over-refusal, not just under-refusal — annotators new to this task tend to reward caution unconditionally, but a model refusing a clearly benign request is a real quality failure too.

Related terms

Put this into practice

Browse open AI training roles from Alignerr, Mercor, Outlier, and more.

Browse AI training jobs