Skip to content
aitrainer.work - AI Training Jobs Platform

RLHF and AI Evaluator Jobs: Get Paid to Rate Model Responses

Paid work to judge AI answers: read a prompt and two or more model responses, pick the better one, score it against a rubric, and explain why. Listings call it AI rater, evaluator, reviewer, task auditor or RLHF annotator; the session is the same. It is remote contract work you do from home, usually part-time and in sessions you schedule. Every field and most languages are represented, from generalist rating to physicians grading clinical answers. Pay on this page runs $35 to $100/hr for the middle half of roles.

285 open jobs
Typical pay
$60/hr
Most roles pay
$35 to $100/hr
Based on
210 posted rates
Updated
October 4, 2026

About RLHF & AI Evaluation work

RLHF evaluation is paid judgment: you decide which answer is better and say why

A session opens with a prompt someone sent to a model, followed by two to four responses to it. You read all of them, pick the best one or rank them, score each against a short rubric (did it follow the instruction, is it correct, is it safe, how is the tone), and write a rationale in your own words. Your choices become the preference data that reinforcement learning from human feedback (RLHF) trains the next model version on, which is why the listing title can say annotator, evaluator, AI rater, reviewer or task auditor and describe the same afternoon of work.

The skill being bought is calibration, not machine learning. Nobody in this queue needs to know what a reward model is. What matters is whether your 5 out of 7 means the same thing as the next rater's 5 out of 7, whether you spot that a confident answer invented a citation, and whether your rationale lets a reviewer check your reasoning in thirty seconds. The word rater comes from search-engine quality rating, and the habits transfer: read the guidelines, apply them the same way every time, flag what the guidelines do not cover.

Three tiers show up in the listings on this page. Generalist rating takes any careful reader with strong written English or another required language, and most language-paired roles sit here. Expert rating asks for a credential or a career (clinicians grading medical advice, lawyers grading legal reasoning, engineers grading code) and pays accordingly. Benchmark authoring is the hardest version: writing questions a frontier model cannot yet answer, with a reference solution and a grading key, so a lab can measure progress.

One example task, start to finish

This is an original example written for this page, not a question from any platform. The prompt: a user asks for a three-day packing list for a November work trip to Oslo, carry-on only, with one client dinner.

Response A runs to forty items, includes a beach towel and sunscreen, and never mentions the carry-on limit. Response B is twelve items, grouped by day, reminds the user that liquids over 100 ml will not clear security, and adds a wool layer for the dinner. It also states that Oslo averages 15°C in November, which is wrong by about ten degrees.

The task asks you to pick the better response, score each from 1 to 7 on helpfulness, accuracy and instruction following, and write a rationale. B is the pick: it followed the carry-on constraint and the dinner, and it is usable as written. The 15°C claim costs it on accuracy, and your rationale has to say so, because an unflagged factual error is the thing a reviewer will catch. A good rationale names the specific failure and the specific win; a weak one restates the scores.

Weak rationaleStrong rationale
B is better because it is more helpful and better organized. A is too long.B is the pick: it respects carry-on only (liquids note) and covers the dinner. A ignores both constraints and lists beach items for Oslo in November. B loses an accuracy point for "15°C average", which is far too warm for the month and should not have been stated without a hedge.

Most queues run tasks like this at ten to thirty minutes each, with a quality reviewer re-checking a sample of your work. The free Academy module Your First Annotation Task walks through a full pairwise comparison with the rationale written out.

It suits careful readers, bilinguals and credentialed experts, in that order of volume

Three groups get hired from this page. The largest is people who read closely and write clearly in the language the listing names: many roles here are language-paired, which means a native Korean, Danish or Brazilian Portuguese reader with good English is more employable than a monolingual English speaker with the same judgment. The language pages list those roles by language.

The second group is domain experts. A hospitalist grading a model's discharge summary, a tax preparer scoring a workflow, a security engineer auditing a code change: these listings name the credential in the title and pay at the top of the range. The third group is engineers and researchers writing benchmarks and evaluation harnesses, where the work is closer to test design than to rating.

It does not suit people who want to skim. Rating queues track agreement with gold-standard answers and with other raters, and accounts that drift below the threshold stop receiving work without much explanation. If you tend to finish tasks fast and move on, transcription or agentic task work rewards speed more than this does.

Screening tests whether you grade the way the guidelines do

The usual sequence is a profile, a qualification test, and a paid trial batch. The qualification test gives you sample prompts and responses with the platform's guidelines and compares your scores with a gold standard, so the question it answers is how calibrated you are, not how much you know. Some platforms add an AI-led video interview before the test, and expert tiers add a credential check or a domain task. Where a listing on this page states its screening step, it is almost always an interview; where it says nothing, assume a test.

Preparation that works: read the guidelines twice before the first sample, score slowly, and write every rationale as if a stranger will audit it. The Academy modules Passing the Screening Process and Reading the Rubric Like a Grader cover the stages and the anatomy of a rubric, and the guide on what rubrics are explains the vocabulary. Interview questions on this site are never quoted from a platform, because candidates sign NDAs.

Pay comes from the listings above, and it splits by tier more than by platform

The pay block at the top of this page is computed from the open listings below it: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. It is recomputed as listings change, so this page never quotes a range that a listing did not state. The method is on the pay methodology page.

Listed rates on this page currently run $35 to $100/hr for the middle half of roles.

Within that spread, generalist rating sits at the bottom, language-paired rating in the middle, and credentialed expert rating and benchmark authoring at the top. A few listings pay per task rather than per hour; those are excluded from the hourly figures. The notes beside this section say when one platform or our own estimates dominate the numbers. For the ranking across every category, see the highest-paying fields and platforms.

How to start this week

  1. Apply to two or three platforms at once. Rating queues fill and empty by project, and one platform rarely gives steady hours. Sort the list below by newest and pick listings whose language and expertise match yours exactly.
  2. Do the free walkthrough before the test. Foundations of RLHF (50 minutes) explains where your ratings go, and Your First Annotation Task has you make a real pairwise decision and write the rationale. Both are readable without an account.
  3. Write your profile in the listing's words. If the title says "AI Response Evaluator (Native Speaker)", your profile should say native speaker, evaluation, and the language, with any credential spelled out.
  4. Take the qualification test rested and unhurried. Agreement with the gold standard decides it, and most people fail by rushing the last third.
  5. Treat the first paid batch as the second test. Quality review on the first fifty tasks decides whether you see more work. Advanced Instruction Following covers the failure reviewers look for most: a response that looks compliant and is not.

Which companies are hiring raters right now, with live counts, is kept up to date in this guide. If you would rather have the board filtered to your background and minimum rate, the matcher does it in three questions and no account.

Adjacent kinds of work: red teaming and safety review grades the model on what it should have refused, agentic tasks grade what a model did rather than what it said, and transcription produces the speech data that voice models train on.

Type:
Popular:

289 Jobs

Page 1 of 10

AfterQuery expert AI training platform

ML Research Expert — Published First Author (Benchmark Authoring)

$130-170

/hr

AfterQuery • PhD • 13d ago
High demand
Ethos expert network platform

Software Engineer - AI Reviewer Expert

$130-170

/hr

Ethos • 54d ago
AfterQuery expert AI training platform

Full-Stack Software Engineer - Testing & Evaluation

$70-90

/hr

AfterQuery • 116d ago
New job posted today
Terac paid research studies platform

Software Engineers: Paid Code Review for AI Agent Evaluation

$60-80

/hr

Terac • 🆕 Today
Mercor AI hiring platform

Baseball Fan – Live MLB Game AI Evaluator (US)

$100-120

/hr

Mercor • 2d ago
English Expert
Newsletter subscription
AI Training Jobs

Weekly AI Training Intelligence

Trending jobs + new guides + real pay data. Every Tuesday.

✅ 500+🔒 No spam

Unsubscribe anytime.

Mercor AI hiring platform

Clinical Mental Health Expert — AI Conversation Evaluation

$90-120

/hr

Mercor • 3d ago
Ethos expert network platform

Data Scientist - AI Reviewer Expert

$130-170

/hr

Ethos • 54d ago
Mercor AI hiring platform

Dermatologist (MD) — Clinical Image Interpretation & AI Evaluation

$240-300

/hr

Mercor • 12d ago
Mercor AI hiring platform

Radiologist (MD) — Medical Imaging AI Annotation & Evaluation

$200-400

/hr

Mercor • 12d ago
Micro1 AI training platform

Gmail & Google Calendar AI Assistant Evaluator

$15-30

/hr

Micro1 • 8d ago
Terac paid research studies platform

Accounting Professionals: Workflow Evaluation Study

$90-120

/hr

Terac • 2d ago
Accounting Coding React
Terac paid research studies platform

Brand Evaluators: Image Evaluation Task

$80-100

/hr

Terac • 2d ago
AI Training Review & QA
Terac paid research studies platform

Photography Evaluators: Paid Task on Image Quality and Brand Guidelines

$50-70

/hr

Terac • 1d ago
AI Training Review & QA
AfterQuery expert AI training platform

Content Writer and Evaluator - Global

$10-20

/hr

AfterQuery • 116d ago
Writing AI Training
Mercor AI hiring platform

Generalist Annotator — Health AI Conversation Quality Evaluation

$20-160

/hr

Mercor • Bachelor's • 15d ago
Terac paid research studies platform

LATAM Software Engineers: Coding Tasks for AI Evaluation

$35-45

/hr

Terac • 1d ago
Terac paid research studies platform

South Asian Software Engineers: Coding Tasks for AI Evaluation

$30-40

/hr

Terac • 1d ago
Handshake AI fellowship program

AI Evaluation Specialist

$30-40

/hr

Handshake • Associate • 7d ago
Expert AI Training
Micro1 AI training platform

Generalist — U.S. Tax Workflow Evaluation

$30-110

/hr

Micro1 • 14d ago
Accounting AI Training
Handshake AI fellowship program

Korean AI Evaluation Specialist (South Korea)

$20-30

/hr

Handshake • Bachelor's • 7d ago
Languages English Korean
Handshake AI fellowship program

Japanese AI Evaluation Specialist (Japan)

$20-30

/hr

Handshake • Bachelor's • 7d ago
Languages English Japanese
Handshake AI fellowship program

Spanish AI Evaluation Specialist (Mexico)

$20-25

/hr

Handshake • Bachelor's • 7d ago
Languages English Spanish
Handshake AI fellowship program

AI Evaluation Specialist

$10-15

/hr

Handshake • Associate • 7d ago
Expert AI Training
Handshake AI fellowship program

Bahasa Indonesian AI Evaluation Specialist (Indonesia)

$10-15

/hr

Handshake • Bachelor's • 7d ago
Languages English Indonesian
Terac paid research studies platform

Transaction Finance Professionals: AI Output Annotation and Ranking

$90-120

Terac • 5d ago
Terac paid research studies platform

Architecture Operations Specialists: Structured Data for Post-Training Evaluation

$90-120

/hr

Terac • 5d ago
AI Training Review & QA
High demand
Micro1 AI training platform

Research Engineer - Code Generation & Model Evaluation

$50-100

Micro1 • Master's • 18d ago
Terac paid research studies platform

Management Consultants: Paid AI Output Evaluation

$60-80

Terac • 5d ago
Outlier AI freelance platform

Voice Conversation Quality Evaluator - Chinese (zh-CN)

$15-25

/hr

Outlier • 3d ago
Languages Chinese
Page 1 of 10 Next →

Get new RLHF & AI Evaluation jobs as they open

One browser alert a day when new roles land, RLHF & AI Evaluation first. No email needed.

Browse by category or skill

Similar to: Software Engineers: Paid Code Review for AI Agent Evaluation

Similar to: Photography Evaluators: Paid Task on Image Quality and Brand Guidelines

Similar to: LATAM Software Engineers: Coding Tasks for AI Evaluation

Frequently Asked Questions

What is an AI rater job?

An AI rater job pays you to read a model's responses and score them against written guidelines, usually picking the better of two answers and explaining why. The title varies by platform (evaluator, reviewer, RLHF annotator), the work is the same, and most roles are remote hourly contracts with no fixed schedule.

Do I need to know machine learning to do RLHF work?

No. Your scores are the training signal; the learning algorithm is the lab's problem. Platforms screen for calibration (do you grade the way the guidelines do), written clarity, and, for expert tiers, a credential in the field being graded.

How much do AI raters and evaluators earn per hour?

The listings on this page currently state $35 to $100/hr for the middle half of roles; the median and the number of rates behind it are in the pay block at the top of the page. Generalist rating sits at the low end, credentialed expert rating at the high end, and a few roles pay per task instead.

Is RLHF evaluation the same as data annotation?

It is one kind of data annotation. Classic annotation labels inputs (boxes on images, tags on text); RLHF evaluation judges outputs (which of two answers is better). Listings and platforms use the words interchangeably, so search both when you look for work.

What is the qualification test for AI rater jobs like?

You get sample prompts, responses and the platform's guidelines, then score the samples. Your scores are compared with a gold standard and with other raters; agreement above a threshold passes. Platforms do not publish the threshold. Treat it as an open-book exam on the guidelines, not a knowledge quiz.

Can I do AI evaluation work part-time or alongside a job?

Yes. Almost every listing here is an hourly contract with hours you choose, and many state a minimum of 10 to 20 hours a week during a project. Steady hours are not guaranteed: projects end, queues empty, and most raters hold accounts on two or three platforms.

Why do AI rater listings disappear and come back?

Labs commission rating in projects of a few weeks, and platforms open and close recruiting as each project starts and fills. The same title can reappear a month later. The browser alert on this page tells you when a role reopens.

How many RLHF & AI Evaluation AI training jobs are open right now?

This page is showing 289 listings, sourced from 15 platforms: Mercor (132), Turing (40), Mindrift (30), Terac (17) and rws (17). The most recent one was picked up today. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.

Explore more opportunities

Browse all job categories and find your next AI training opportunity.

View All Categories