Skip to content
aitrainer.work - AI Training Jobs Platform
Data Science mercor

Human Baseliner for Open-Ended ML Research Tasks

Mercor • Remote

Education

PhD

Type

Per task

Pay Rate

$75–$90/task

Listed

105d ago

Apply opens Mercor in a new tab.

Apply Now →

What We Know About This Role

Weekly hours
20 hrs/week

About this role

From the Mercor listing

Overview

We are hiring experienced machine learning engineers and researchers to serve as human baseliners for evaluations of open-ended machine learning research tasks. These evaluations measure how well AI agents perform on realistic AI R&D problems. To interpret agent performance, we also need strong human reference points: skilled practitioners attempting the same tasks under the same time and compute constraints. As a baseliner, you will complete self-contained ML research tasks in a sandboxed environment, working independently with your preferred tools and workflow. Your performance will be used as a benchmark against which frontier-model agents are evaluated.

What You’ll Do

  • Attempt open-ended machine learning research tasks under a fixed time and compute budget (work trial)

  • Work independently in a sandboxed Linux environment with internet access

  • Use your preferred tooling, including IDEs and AI coding assistants such as Cursor, Claude Code, and ChatGPT

  • Record your full working session via screen recording

  • Complete a short pre-task and post-task questionnaire

  • Submit your final work product, screen recording, and completed questionnaires: Post this you will be hired for a longer commitment

Commitment

  • Minimum 20 hours per week if selected

  • More availability is strongly preferred

Requirements

Candidates must meet all of the following:

  • 3+ years of machine learning experience

    • Time spent in a PhD program counts toward this requirement

    • Undergraduate and master’s experience does not count

  • Attended a top-100 university or worked at FAANG or a comparable company

  • Experience with at least one major ML framework such as PyTorch, JAX, or TensorFlow

  • Deep, hands-on expertise in at least one of the focus areas below:

    • Pretraining under tight data and compute budgets

    • PPO, reward shaping, custom gym / gymnasium environments, and throughput tuning

    • Full fine-tuning, LoRA, QLoRA, DPO, RLHF, RLAIF, and distillation

    • Large-scale corpus filtering, deduplication, subsampling, and benchmark contamination avoidance

    • Architecture design under strict parameter-count or size constraints

    • Modifying pretrained architectures, including attention patterns, pooling heads, or training objectives

    • Contrastive training for embedding or retrieval models

    • Generative vision or video modeling

    • Multilingual or low-resource language experience

    • Image or video data pipelines at scale

    • Experience balancing competing model objectives such as safety and capability

    • Prior work as an ML evaluator, red-teamer, or baseliner

Required Domain Expertise

Candidates must have strong practical experience in at least one of the following:

  • Pretraining: training transformer language models from scratch

  • Reinforcement learning: training agents in custom or existing environments

  • Post-training: fine-tuning and aligning LLMs

  • Dataset curation: building and cleaning large text corpora for LLM training

  • Model architecture: designing and modifying neural network architectures

Logistics (work trial requirements)

  • One baseline attempt per contractor per task

  • Each task may only be attempted once by a given contractor

  • All work is confidential and covered by NDA

  • Compute and environment are provided; no personal GPU is required

How long hiring takes

Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.

Within 2 weeks
~25%
Within 6 weeks
~60%
Within 3 months
~80%

Talent Pool members

Apply through this link and we can put you forward to Mercor when your profile is a strong match. Not every applicant is submitted. If you're not in the pool yet, set up your profile first.

Set up your profile →

What to Expect

Looking at Mercor Data Science listings we've tracked, contracts in this domain typically run about 3.6 weeks. Actual length varies by project, but this gives you a realistic baseline going in.

Based on 16 extracted Mercor Data Science listings.

Why this role

At $75–$90/task, this remote Human Baseliner for Open-Ended ML Research Tasks role pays competitively without asking you to relocate, interview in person, or keep fixed hours in Data Science.

Talent pool

We're light on Data Science candidates

We've matched 9 people with a Data Science background against 293 Data Science listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.

Set up your profile

Skills and categories

Explore other opportunities in related specializations:

Related jobs

Mercor

Browse All Jobs from Mercor

Discover more opportunities on Mercor that match your skills and interests.

View All Mercor Jobs →

Verified Reviews

Loading reviews…

Community Reviews

Loading reviews…
💬

Share your experience with Mercor

Help other candidates make better decisions by leaving a review.

Sign in to leave a review

Common questions

Does it cost money to apply to Mercor?

No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.

Is Mercor for freelancers or full-time contractors?

Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.

What does task-based AI training work look like?

Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.

What does asynchronous AI training work mean in practice?

No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.

What does Data Science work look like for a Human Baseliner for Open-Ended ML Research Tasks?

Tasks here are scoped to Data Science, not generic labeling. As a Human Baseliner for Open-Ended ML Research Tasks, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Data Science) rather than following a one-size-fits-all rubric. If you don't have hands-on Data Science background, this is likely not the right listing to start with.

What specific skills does this listing call for?

Coding and Multilingual are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.

How many hours per week does this role require?

Based on the listing, this role is scoped at about 20 hours per week. Treat this as a real commitment expectation, not a loose estimate.

How much does this specific role pay?

This listing is posted at $75–$90/task, a per-task rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.

How does per-task pay translate to an hourly rate?

It doesn't directly, that's the catch. Each task has a listed price and an estimated completion time, but your actual effective rate depends on how fast you work. If a task is estimated at 20 minutes and pays $8, that is $24/hr for a fast worker and $12/hr for a slow one. Track your actual time on the first few tasks before treating the listed rate as your real hourly income.

What happens when I click Apply on this listing?

You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.

Is a PhD required?

For this specific role, yes, or near-equivalent professional depth. The credential gate is enforced at the assessment stage, not just on paper. That said, active PhD candidates and people with equivalent published research have qualified without a formal degree. The assessment is the real filter.

How soon will I start working after applying to Mercor?

Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.