Human Baseliner for Open-Ended ML Research Tasks
Mercor • Remote
Education
PhD
Type
Per task
Pay Rate
$75–$90/task
Listed
105d ago
Apply opens Mercor in a new tab.
Apply Now → ⚡ Boost your chances - Optimize your resume with Rezi.aiWhat We Know About This Role
- Weekly hours
- 20 hrs/week
About this role
From the Mercor listing
Overview
We are hiring experienced machine learning engineers and researchers to serve as human baseliners for evaluations of open-ended machine learning research tasks. These evaluations measure how well AI agents perform on realistic AI R&D problems. To interpret agent performance, we also need strong human reference points: skilled practitioners attempting the same tasks under the same time and compute constraints. As a baseliner, you will complete self-contained ML research tasks in a sandboxed environment, working independently with your preferred tools and workflow. Your performance will be used as a benchmark against which frontier-model agents are evaluated.
What You’ll Do
Attempt open-ended machine learning research tasks under a fixed time and compute budget (work trial)
Work independently in a sandboxed Linux environment with internet access
Use your preferred tooling, including IDEs and AI coding assistants such as Cursor, Claude Code, and ChatGPT
Record your full working session via screen recording
Complete a short pre-task and post-task questionnaire
Submit your final work product, screen recording, and completed questionnaires: Post this you will be hired for a longer commitment
Commitment
Minimum 20 hours per week if selected
More availability is strongly preferred
Requirements
Candidates must meet all of the following:
3+ years of machine learning experience
Time spent in a PhD program counts toward this requirement
Undergraduate and master’s experience does not count
Attended a top-100 university or worked at FAANG or a comparable company
Experience with at least one major ML framework such as PyTorch, JAX, or TensorFlow
Deep, hands-on expertise in at least one of the focus areas below:
Pretraining under tight data and compute budgets
PPO, reward shaping, custom
gym/gymnasiumenvironments, and throughput tuningFull fine-tuning, LoRA, QLoRA, DPO, RLHF, RLAIF, and distillation
Large-scale corpus filtering, deduplication, subsampling, and benchmark contamination avoidance
Architecture design under strict parameter-count or size constraints
Modifying pretrained architectures, including attention patterns, pooling heads, or training objectives
Contrastive training for embedding or retrieval models
Generative vision or video modeling
Multilingual or low-resource language experience
Image or video data pipelines at scale
Experience balancing competing model objectives such as safety and capability
Prior work as an ML evaluator, red-teamer, or baseliner
Required Domain Expertise
Candidates must have strong practical experience in at least one of the following:
Pretraining: training transformer language models from scratch
Reinforcement learning: training agents in custom or existing environments
Post-training: fine-tuning and aligning LLMs
Dataset curation: building and cleaning large text corpora for LLM training
Model architecture: designing and modifying neural network architectures
Logistics (work trial requirements)
One baseline attempt per contractor per task
Each task may only be attempted once by a given contractor
All work is confidential and covered by NDA
Compute and environment are provided; no personal GPU is required
How long hiring takes
Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.
Talent Pool members
Apply through this link and we can put you forward to Mercor when your profile is a strong match. Not every applicant is submitted. If you're not in the pool yet, set up your profile first.
Set up your profile →What to Expect
Looking at Mercor Data Science listings we've tracked, contracts in this domain typically run about 3.6 weeks. Actual length varies by project, but this gives you a realistic baseline going in.
Based on 16 extracted Mercor Data Science listings.
Why this role
At $75–$90/task, this remote Human Baseliner for Open-Ended ML Research Tasks role pays competitively without asking you to relocate, interview in person, or keep fixed hours in Data Science.
Talent pool
We're light on Data Science candidates
We've matched 9 people with a Data Science background against 293 Data Science listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Mercor
Discover more opportunities on Mercor that match your skills and interests.
View All Mercor Jobs →Verified Reviews
Community Reviews
Share your experience with Mercor
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
Does it cost money to apply to Mercor?
No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.
Is Mercor for freelancers or full-time contractors?
Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.
What does task-based AI training work look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does Data Science work look like for a Human Baseliner for Open-Ended ML Research Tasks?
Tasks here are scoped to Data Science, not generic labeling. As a Human Baseliner for Open-Ended ML Research Tasks, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Data Science) rather than following a one-size-fits-all rubric. If you don't have hands-on Data Science background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding and Multilingual are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How many hours per week does this role require?
Based on the listing, this role is scoped at about 20 hours per week. Treat this as a real commitment expectation, not a loose estimate.
How much does this specific role pay?
This listing is posted at $75–$90/task, a per-task rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.
How does per-task pay translate to an hourly rate?
It doesn't directly, that's the catch. Each task has a listed price and an estimated completion time, but your actual effective rate depends on how fast you work. If a task is estimated at 20 minutes and pays $8, that is $24/hr for a fast worker and $12/hr for a slow one. Track your actual time on the first few tasks before treating the listed rate as your real hourly income.
What happens when I click Apply on this listing?
You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Is a PhD required?
For this specific role, yes, or near-equivalent professional depth. The credential gate is enforced at the assessment stage, not just on paper. That said, active PhD candidates and people with equivalent published research have qualified without a formal degree. The assessment is the real filter.
How soon will I start working after applying to Mercor?
Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.