Response Safety Grading
Scoring model outputs across defined safety risk dimensions.
Response safety grading assigns a score or tier to a model output based on how much risk it poses across defined dimensions — physical harm, psychological harm, legal risk, and similar — rather than a single pass/fail safety judgment. This granularity lets teams distinguish between a mildly risky response and a severely dangerous one, which matters for both training signal quality and for setting deployment thresholds.
This differs from toxicity annotation or policy enforcement in scope: those focus on content categories (is this abusive, does this violate policy X), while safety grading focuses on consequence severity across whatever category applies.
The resulting scores feed into harmlessness training objectives alongside helpfulness signals, since a genuinely well-aligned model needs both — being safe alone isn't the goal if it comes at the cost of being useless.
What this means for trainers
Grade severity, not just presence — a response that's marginally risky and one that's severely dangerous should land in clearly different tiers, even though both technically involve some risk.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs