Senior Machine Learning Engineer - LLM Evaluation / Task Creations (India Based)
Mercor • Remote, India
Education
Any
Type
hourly
Pay Rate
$35/hr
Listed
211d ago
✅ Applying through this link supports our platform at no cost to you.
This position is hosted on an external talent platform. Please only apply for this position if it fits your skills and interests.
In our Talent Pool?
Apply through this link and we can vouch for you to Mercor. ? We vouch for Talent Pool members who apply through our referral link, when we believe they're a strong match. Not every applicant gets a vouch. Not in the pool yet? Set up your profile first.
Set up your profile →Mercor: our referral track record
We've referred 556 candidates to Mercor roles. 5% (29) were placed.
What We Know About This Role
- Weekly hours
- 30–40 hrs/week
About this Role
From the Mercor listing
Role Description
- Mercor is hiring on behalf of a leading AI research lab to bring on highly skilled Machine Learning Engineers with a proven record of building, training, and evaluating high-performance ML systems in real-world environments. In this role, you will design, implement, and curate high-quality machine learning datasets, tasks, and evaluation workflows that power the training and benchmarking of advanced AI systems.
- This position is ideal for engineers who have excelled in competitive machine learning settings such as Kaggle, possess deep modelling intuition, and can translate complex real-world problem statements into robust, well-structured ML pipelines and datasets. You will work closely with researchers and engineers to develop realistic ML problems, ensure dataset quality, and drive reproducible, high-impact experimentation. Candidates should have 2+ years of applied ML experience or a strong record in competitive ML, and must be based in India. Ideal applicants are proficient in Python, experienced in building reproducible pipelines, and familiar with benchmarking frameworks, scoring methodologies, and ML evaluation best practices.
Responsibilities
- Frame unique ML problems for enhancing ML capabilities of LLMs.
- Design, build, and optimise machine learning models for classification, prediction, NLP, recommendation, or generative tasks.
- Run rapid experimentation cycles, evaluate model performance, and iterate continuously.
- Conduct advanced feature engineering and data preprocessing.
- Implement adversarial testing, model robustness checks, and bias evaluations.
- Fine-tune, evaluate, and deploy transformer-based models where necessary.
- Maintain clear documentation of datasets, experiments, and model decisions. Stay updated on the latest ML research, tools, and techniques to push modelling capabilities forward.
Required Qualifications
- At least 2 years of full-time experience in machine learning model development
- Technical degree in Computer Science, Electrical Engineering, Statistics, Mathematics, or a related field
- Demonstrated competitive machine learning experience (Kaggle, DrivenData, or equivalent)
- Evidence of top-tier performance in ML competitions (Kaggle medals, finalist placements, leaderboard rankings)
- Strong proficiency in Python, PyTorch/TensorFlow, and modern ML/NLP frameworks
- Solid understanding of ML fundamentals: statistics, optimisation, model evaluation, architectures
- Experience with distributed training, ML pipelines, and experiment tracking
- Strong problem-solving skills and algorithmic thinking
- Experience working with cloud environments (AWS/GCP/Azure)
- Exceptional analytical, communication, and interpersonal skills
- Ability to clearly explain modelling decisions, tradeoffs, and evaluation results
- Fluency in English
Preferred / Nice to Have
- Kaggle Grandmaster, Master, or multiple Gold Medals
- Experience creating benchmarks, evaluations, or ML challenge problems
- Background in generative models, LLMs, or multimodal learning
- Experience with large-scale distributed training
- Prior experience in AI research, ML platforms, or infrastructure teams
- Contributions to technical blogs, open-source projects, or research publications
- Prior mentorship or technical leadership experience
- Published research papers (conference or journal)
- Experience with LLM fine-tuning, vector databases, or generative AI workflows
- Familiarity with MLOps tools: Weights & Biases, MLflow, Airflow, Docker, etc.
- Experience optimising inference performance and deploying models at scale
Why Join
- Gain exposure to cutting-edge AI research workflows, collaborating closely with data scientists, ML engineers, and research leaders shaping next-generation AI systems.
- Work on high-impact machine learning challenges while experimenting with advanced modelling strategies, new analytical methods, and competition-grade validation techniques.
- Collaborate with world-class AI labs and technical teams operating at the frontier of forecasting, experimentation, tabular ML, and multimodal analytics.
- Flexible engagement options (30–40 hrs/week or full-time) — ideal for ML engineers eager to apply Kaggle-level problem solving to real-world, production-grade AI systems. We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
- You will be engaged as an independent contractor.
- This is a fully remote role that can be completed on your own schedule.
- Projects can be extended, shortened, or concluded early depending on needs and performance.
- Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.
- Payments are weekly on Stripe or Wise based on services rendered.
- Please note: We are unable to support H1-B or STEM OPT candidates at this time.
Requirements
- Must be eligible to work in one of: Remote, India
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
Eligible Languages
Fluent proficiency in English
Key Responsibilities
- Frame unique ML problems for enhancing ML capabilities of LLMs.
- Design, build, and optimise machine learning models for classification, prediction, NLP, recommendation, or generative tasks.
- Run rapid experimentation cycles, evaluate model performance, and iterate continuously.
- Conduct advanced feature engineering and data preprocessing.
Why This Role
Work from anywhere, at any time. This fully remote Senior Machine Learning Engineer - LLM Evaluation / Task Creations (India Based) position ($35/hr) breaks down geographic barriers, allowing you to earn US-competitive rates regardless of your local market. It is a perfect stepping stone for building a career in the Software Engineering ecosystem.
Skills & Categories
Explore other opportunities in related specializations:
Related Jobs
Software Engineering - Research & Evaluation Studies
terac • Software Engineering
$250 /task
Engineering & Data tools Specialist
micro1 • Software Engineering
$80 /hr
QA / Software Engineering Reviewer – Browser Test Validation
mercor • Software Engineering
$60 /hr
Head of AI & Engineering Expert
ethos • Software Engineering
$150 /hr
Browse All Jobs from Mercor
Discover more opportunities on Mercor that match your skills and interests.
View All Mercor Jobs →Verified Reviews
Community Reviews
Share your experience with Mercor
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Frequently Asked Questions
Is Mercor for freelancers or full-time contractors?
Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.
Does Mercor's application require an on-camera interview?
Yes, every applicant records a video interview with an AI interviewer that asks questions about your resume. Clients review that recording to judge communication skills before matching, so there's no way to apply without going on camera.
Does it cost money to apply to Mercor?
No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.
What does task-based AI training work actually look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does Software Engineering work look like for a Senior Machine Learning Engineer - LLM Evaluation / Task Creations (India Based)?
Tasks here are scoped to Software Engineering, not generic labeling. As a Senior Machine Learning Engineer - LLM Evaluation / Task Creations (India Based), expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Software Engineering) rather than following a one-size-fits-all rubric. If you don't have hands-on Software Engineering background, this is likely not the right listing to start with.
How many hours per week does this role require?
Based on the listing, this role is scoped at 30–40 hours per week. Treat this as a real commitment expectation, not a loose estimate.
Do I need to be fluent in English?
Yes. This role specifically requires English proficiency. You will likely be evaluated on written fluency during the assessment, not just conversational level. If English is not your first language or you are not professionally fluent, this is not the right role. Filter for your native language to find better-matched listings.
What happens when I click Apply on this listing?
You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Can I apply from outside India?
This specific role is open only to people based in India. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply.
How soon will I start working after applying to Mercor?
Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.