Skip to content
aitrainer.work - AI Training Jobs Platform
Generalist mercor

Generalist Annotator — Health AI Conversation Quality Evaluation

Mercor • Remote, United States

Education

Bachelor's

Type

Hourly

Pay Rate

$20–$160/hr

Listed

15d ago

Apply opens Mercor in a new tab.

Apply Now →

About this role

From the Mercor listing

About the Role

Mercor is hiring generalist annotators to evaluate conversations between users and a Health & Fitness Assistant AI. You will judge how well the AI communicates: whether it answered the actual question, whether a non-expert could follow it, and whether it was honest even when honesty was the harder answer.

This is a non-clinical role. You do not need medical training, and you should not try to supply it — clinicians assess accuracy and safety on the same conversations, as a separate track. Your value is precisely that you are not one of them: you can say how a real person would experience the response, which a clinician reading the same text cannot.

What you will do

  • Relevance and clarity — did the AI answer what was asked, in language a family member with no medical background could act on?
  • Sycophancy detection — the core of this work. Spotting where the AI told someone what they wanted to hear instead of what they needed to hear: backing down when a user pushed back, hiding behind "you know your body best", or framing a risky decision as empowering rather than naming the risk. Automated evaluators reliably score this as helpful, which is why it needs a person.
  • Tone and outcome — whether the conversation left the user informed and calmer, or anxious and no better off, and whether they walked away with anything concrete to do.
  • Written justification — every rating below the top option requires a specific, quotable explanation of what went wrong and what a better response would have said. The ratings tell us something is wrong; only your writing tells us what to fix.

Conversations run from 1 to 7 turns; many are a single turn. Expect roughly 2–13 minutes each depending on length. Full written instructions and worked examples are provided before you start.

Required qualifications

  • Strong written English, and the ability to explain a judgment in two to four specific sentences rather than a one-line verdict
  • Careful reading — the work rewards noticing what a response quietly left out, not just what it got wrong
  • Comfort applying a detailed rubric consistently across many items
  • Reliable availability, with the ability to concentrate hours when a batch is time-boxed

Preferred qualifications

  • Prior annotation, human-feedback, evaluation, or content-review experience
  • A background where you have had to explain something technical to a non-expert audience — teaching, editing, writing, customer support, patient advocacy
  • Experience with rubric-based grading or quality assurance work

Important

Using an AI tool to write your comments is prohibited and will end your work on this project. This project exists to capture human judgment that AI systems lack; AI-written feedback corrupts the dataset we use to check those systems. It is checked for.

Why this work

Most health-AI failures are not exotic. They are ordinary answers that sound responsible, read as warm and careful, and still leave someone worse off than if they had never asked. Automated evaluators are good at spotting responses that look careful, which is exactly why they miss these. Catching them takes a person who read the whole conversation and noticed where it quietly went wrong.

Requirements

  • Must be eligible to work in one of: Remote, United States
  • Fluent proficiency in English (Written & Verbal)
  • Reliable high-speed internet connection
  • Bachelor's's degree or equivalent professional experience
  • Demonstrated expertise in Generalist

How long hiring takes

Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.

Within 2 weeks
~25%
Within 6 weeks
~60%
Within 3 months
~80%

Talent Pool members

Apply through this link and we can put you forward to Mercor when your profile is a strong match. Not every applicant is submitted. If you're not in the pool yet, set up your profile first.

Set up your profile →

What to Expect

Looking at Mercor Generalist listings we've tracked, contracts in this domain typically run about 6.8 weeks. Actual length varies by project, but this gives you a realistic baseline going in.

Based on 10 extracted Mercor Generalist listings.

Why this role

$20–$160/hr for Generalist Annotator — Health AI Conversation Quality Evaluation work covers both income and flexibility. You set your own hours, and the work draws on Generalist knowledge you already have instead of asking you to learn a new field.

Skills and categories

Explore other opportunities in related specializations:

Related jobs

Mercor

Browse All Jobs from Mercor

Discover more opportunities on Mercor that match your skills and interests.

View All Mercor Jobs →

Verified Reviews

Loading reviews…

Community Reviews

Loading reviews…
💬

Share your experience with Mercor

Help other candidates make better decisions by leaving a review.

Sign in to leave a review

Common questions

Does Mercor's application require an on-camera interview?

Yes, every applicant records a video interview with an AI interviewer that asks questions about your resume. Clients review that recording to judge communication skills before matching, so there's no way to apply without going on camera.

Does it cost money to apply to Mercor?

No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.

What does task-based AI training work look like?

Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.

What does Generalist work look like for a Generalist Annotator — Health AI Conversation Quality Evaluation?

Tasks here are scoped to Generalist, not generic labeling. As a Generalist Annotator — Health AI Conversation Quality Evaluation, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Generalist) rather than following a one-size-fits-all rubric. If you don't have hands-on Generalist background, this is likely not the right listing to start with.

What specific skills does this listing call for?

Data Annotation, English, and Entry Level are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.

How much does this specific role pay?

This listing is posted at $20–$160/hr, an hourly rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.

What happens when I click Apply on this listing?

You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.

Can I apply from outside the United States?

This specific role is open only to people based in the United States. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply.

How soon will I start working after applying to Mercor?

Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.