Instruction Following Evaluation
Checking whether a model's output satisfies the explicit constraints given in the prompt.
What this means for trainers
Read the prompt for constraints first, response second — annotators who skim the prompt and focus only on response quality routinely miss the exact failures this task exists to catch.
Instruction-following evaluation isolates one specific question: did the model do what it was asked, including every constraint such as word count, format, tone, list length, and required sections, not just the general topic. A response can be well-written and accurate and still fail this evaluation if it ignores a stated constraint like answering in exactly three bullet points.
This differs from broader helpfulness scoring in that it is checklist-driven rather than holistic. Annotators typically extract every explicit and implicit constraint from the prompt and verify each one independently before forming an overall judgment, so a violation of a minor formatting rule cannot get lost inside a generally positive impression of the response.
Models are notoriously inconsistent at following compound instructions, meaning several constraints given at once, which makes this one of the more diagnostic evaluation types for catching real capability gaps. A model that handles single constraints reliably can still drop one of three when asked to satisfy all three together.
Related terms
Put this into practice
Browse open AI training roles from Alignerr, Mercor, Outlier, and more.
Browse AI training jobs