AI Evaluation Analyst
Micro1 • Remote
Education
Any
Type
per_task
Pay Rate (by country)
$20–$30/task
Actively Hiring
100 openings
✅ Applying through this link supports our platform at no cost to you.
This position is hosted on an external talent platform. Please only apply for this position if it fits your skills and interests.
Micro1: our referral track record
We've referred 735 candidates to Micro1 roles. 1% (8) were placed.
What We Know About This Role
- Start timeline
- Starts within 24–48 hrs of onboarding
About this Role
From the Micro1 listing
Role Title: AI Evaluation Analyst Role Type: Contractor Location: Remote micro1 is engaging AI Evaluation Analysts to contribute to a customer’s project focused on advancing frontier language model capabilities. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input. No prior experience in AI is required — your domain knowledge is what matters. You will play a crucial role in producing evaluation and training data that directly influence how powerful AI models understand and interact. This is an exciting opportunity to operate at the forefront of AI, with a focus on written and verbal clarity, multi-turn conversation design, and deep analysis of model behavior.
Scope of Work
- Author detailed, task-based multi-turn conversations and rubrics aligned with project specifications.
- Test and refine conversation drafts against frontier large language models, iterating to meet quality and difficulty requirements.
- Deliver comprehensive evaluation assets including transcripts, target behaviors, binary rubrics, and supporting evidence.
- Ensure strict fidelity to evolving project specs while maintaining high throughput and attention to detail.
- Validate and calibrate outputs with team leads and quality control as guidelines change.
- Work independently and consistently, meeting expected output rates for deliverable completion.
Preferred Qualifications
- Native-level written English with exceptional clarity, structure, and attention to detail.
- Prior experience in data annotation, RLHF, SFT, evaluation, or prompt engineering for AI systems.
- Working knowledge of frontier LLM behaviors and common model failure patterns.
- Demonstrated ability to interpret and apply highly detailed specifications without supervision.
- Strong critical thinking and analytical skills in writing-heavy or analysis-heavy domains.
- Experience authoring evaluation items, rubrics, or conducting deep analysis of technology outputs.
- Background in research, editorial, technical writing, or quality assurance is a plus.
Compensation Structure
Compensation is output-based; experts are paid per task that meets the project specifications. The time required to complete work may vary depending on the expert’s experience and workflow. Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
Start Timeline & Availability
We typically fill roles within 48 hours and are looking for experts ready to jump in right away. If selected, we expect you to start your first tasks within 24–48 hours of completing onboarding.
Requirements
- Working knowledge of frontier LLM behavior
- Data Annotation
- Written English clarity and structure
- Spec fidelity at volume
- Self-direction to a detailed spec
- Must be eligible to work in Remote
Eligible Languages
Fluent proficiency in English
Why This Role
Compensation is output-based; experts are paid per task that meets the project specifications. The time required to complete work may vary depending on the expert’s experience and workflow. Minimum submission requirements apply. Experts must submit a minimum of tasks per week.
Skills & Categories
Explore other opportunities in related specializations:
Related Jobs
Browse All Jobs from Micro1
Discover more opportunities on Micro1 that match your skills and interests.
View All Micro1 Jobs →Verified Reviews
Community Reviews
Share your experience with Micro1
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Frequently Asked Questions
Can I apply to Micro1 from outside the US or UK?
Yes. Micro1 accepts applicants from most countries worldwide, unlike DataAnnotation or Outlier, both of which cap access to a handful of English-speaking countries. Most roles still require strong written and spoken English.
Does Micro1 monitor your computer while you work?
On many projects, yes, through time-tracking tools that take periodic screenshots to verify active hours. Check the specific project's requirements before accepting if desktop monitoring is a dealbreaker.
What does the Micro1 application process look like?
Expect a screening interview with Zara, Micro1's AI recruiter. Prepare for it like a real video call: good lighting, clear audio, verbal answers to technical questions. A human manager reviews the recording afterward, and some roles add a short skills assessment on top.
What does task-based AI training work actually look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does Generalist work look like for a AI Evaluation Analyst?
Tasks here are scoped to Generalist, not generic labeling. As a AI Evaluation Analyst, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Generalist) rather than following a one-size-fits-all rubric. If you don't have hands-on Generalist background, this is likely not the right listing to start with.
What does the 'openings' count on this listing mean?
It's Micro1's live count of how many candidates it's still trying to place for this specific role, 100 right now, not a countdown on your individual application. A higher number signals more active demand; it doesn't lower the bar for acceptance.
Do I need to be fluent in English?
Yes. This role specifically requires English proficiency. You will likely be evaluated on written fluency during the assessment, not just conversational level. If English is not your first language or you are not professionally fluent, this is not the right role. Filter for your native language to find better-matched listings.
How does per-task pay translate to an hourly rate?
It doesn't directly, that's the catch. Each task has a listed price and an estimated completion time, but your actual effective rate depends on how fast you work. If a task is estimated at 20 minutes and pays $8, that is $24/hr for a fast worker and $12/hr for a slow one. Track your actual time on the first few tasks before treating the listed rate as your real hourly income.
What happens when I click Apply on this listing?
You'll be taken to Micro1's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.