Skip to content
aitrainer.work - AI Training Jobs Platform
Rubric-Based Evaluation definition
core concepts evaluation

Rubric-Based Evaluation

Scoring outputs across clear, predefined dimensions such as correctness, safety, and completeness.

Rubric-based evaluation breaks a holistic quality judgment into explicit, separately scored dimensions rather than asking for one overall rating. A response might be scored independently on correctness, helpfulness, safety, and tone, with each dimension defined clearly enough that different annotators land on similar scores for the same output.

This is the general framework underneath many specific evaluation types covered elsewhere in this glossary — summarization evaluation, code correctness evaluation, and instruction-following evaluation are all rubric-based evaluation applied to a specific domain with its own dimensions.

Well-designed rubrics anchor each score level with concrete examples rather than abstract descriptions ("a 3 looks like this" rather than "a 3 is moderately good"), since vague anchors are one of the fastest ways to introduce rubric drift across a large annotator pool.

What this means for trainers

Score each dimension independently rather than letting an overall impression bleed into every category — a response can legitimately be high on safety and low on helpfulness at the same time.

Related terms

Put this into practice

Browse open AI training roles from Alignerr, Mercor, Outlier, and more.

Browse AI training jobs