AI Prompt Quality Rater
Claru • Remote
Education
Not stated
Type
Hourly
Pay Rate
$35–$55/hr
Listed
13d ago
About this role
From the Claru listing
You'll evaluate prompt-response pairs generated by large language models, scoring them across dimensions like factual accuracy, instruction-following, coherence, and appropriate refusal behavior. Each session involves reviewing batches of 20–50 pairs using a structured rubric inside a web-based annotation platform, flagging edge cases, and writing brief justifications for non-obvious scores.
Beyond surface-level ratings, you'll identify subtle failure modes — responses that are technically accurate but misleading, answers that follow the letter of a prompt while violating its intent, and outputs that pass a quick read but contain embedded errors. You'll work asynchronously, with a typical batch taking 60–90 minutes, and you're expected to maintain inter-annotator agreement scores above 0.75 kappa.
Feedback you submit feeds directly into preference datasets used to fine-tune and align production models. Calibration sessions run bi-weekly, where you'll review disagreements with a senior evaluator and update your scoring approach. Strong performance can lead to specialized tracks covering domain-specific evaluation (legal, medical, code).
Requirements
- Must be eligible to work in Remote
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
Why this role
This AI Prompt Quality Rater role pays $35–$55/hr fully remote, so the Generalist work here isn't limited to whoever happens to live near an AI lab's office.
Skills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Claru
Discover more opportunities on Claru that match your skills and interests.
View All Claru Jobs →Verified Reviews
Community Reviews
Share your experience with Claru
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
What makes this one of Claru's higher-paying roles?
Claru bands its desk-based work by how much judgement a task needs. This listing sits in the top band at $35-$55/hr, alongside its RLHF evaluation, medical and 3D vision annotation and red-teaming roles, all of which expect graduate-level or professional depth in the subject and the ability to justify a score in writing. Its entry-tier text labeling sits at $20-$35/hr for comparison.
Is Claru desk work a steady queue or does it arrive in batches?
Batches. Claru describes this work as asynchronous with no scheduled hours and no meetings: you claim a batch, complete it and submit. There is no guarantee a new batch is waiting when you finish one, so treat it as recurring project income rather than a fixed weekly wage.
How does Claru measure quality on this kind of work?
Through agreement scoring and review. Its evaluation roles set a target for how closely your scores track other annotators, and disagreements with reviewers are resolved in an async comment thread you are expected to answer within a day. Consistently strong scores are what unlock the senior and lead tracks Claru mentions in its listings.
Can I work on Claru from outside the United States?
For the desk-based annotation and evaluation roles, yes: they are listed as remote with no country gate. The country restrictions on Claru apply to its video capture listings, which are limited to the specific markets named on each one.
Do task-based AI training roles require an interview?
Most skip a live interview entirely and gate access through an automated assessment or qualification task instead. Where we've confirmed the specific requirement for this listing, it's called out above in What We Know About This Role.
What does task-based AI training work look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does Generalist work look like for an AI Prompt Quality Rater?
Tasks here are scoped to Generalist, not generic labeling. As an AI Prompt Quality Rater, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Generalist) rather than following a one-size-fits-all rubric. If you don't have hands-on Generalist background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Data Annotation and Entry Level are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
This listing is posted at $35–$55/hr, an hourly rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.
What happens when I click Apply on this listing?
You'll be taken to Claru's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Who is behind Claru?
Claru is the contributor-facing brand of Reka AI, which collects human data for frontier AI labs across text, vision, video and robotics. Its board splits into regional video capture work and desk-based annotation and evaluation, and the apply link on this page is an affiliate link that does not change your pay or what you are offered.