Applied Engineer, Evaluations Role
Mercor • Remote
Education
Any
Type
Hourly
Pay Rate
$100–$120/hr
Listed
114d ago
Apply opens Mercor in a new tab.
Apply Now → ⚡ Boost your chances - Optimize your resume with Rezi.aiAbout this role
From the Mercor listing
Mercor is seeking software engineers to build and refine evaluations. You’ll turn completed pull requests into engineering tasks and use our in-house evaluation framework to multiple frontier coding agents, including Claude Code, Codex and more.
We work alongside an in-house research team and have produced industry leading benchmarks to compare the performance of frontier large language models.
In this role, you will:
Identify repositories with enough substantive work to support challenging evaluations, and select suitable completed PRs.
Investigate each problem, its reference solution, and the repository’s architecture, tests, and conventions.
Write task prompts and build evaluations using automated tests, shell commands, and LLM grading prompts.
Validate evaluations against reference solutions and deliberately flawed implementations. Find missing checks, incorrect grades, and criteria that unnecessarily constrain how a problem can be solved.
Run evaluations repeatedly across multiple models and harnesses. Investigate whether failures come from the agent’s solution, the environment, or the grading, and establish that tasks expose meaningful weaknesses in agent performance.
Refine evaluations through repeated testing and review. Document findings and grading decisions, and work through feedback.
You’ll need:
Strong programming fundamentals and practical experience working in substantial and complex codebases. Languages we create evals for include TypeScript/JavaScript, Python, Java, Kotlin, Go, Ruby, PHP, C++ and Rust
The ability to understand unfamiliar code, investigate subtle behavior, and assess whether different implementations solve the same problem correctly.
Experience writing meaningful tests, including edge cases and regression coverage.
Attention to detail and patience for repeated investigation and refinement.
Practical experience using AI coding tools, with the judgment to verify their output and catch mistakes.
Confidence using Git, test runners, and CI tooling.
Clear and fluent written and verbal communication with the ability to own a task independently while raising questions when requirements are ambiguous.
Relevant experience can come from open-source, private, or enterprise repositories. Familiarity with a particular language or ecosystem is helpful; the ability to learn the repository and make sound engineering judgments matters more than its popularity or your public contribution history.
We provide onboarding to the evaluation framework and ongoing review feedback. After onboarding, you’ll be expected to own task selection, evaluation development, and iteration without step-by-step direction.
Requirements
- Must be eligible to work in Remote
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
- Any's degree or equivalent professional experience
- Demonstrated expertise in STEM
How long hiring takes
Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.
Talent Pool members
Apply through this link and we can put you forward to Mercor when your profile is a strong match. Not every applicant is submitted. If you're not in the pool yet, set up your profile first.
Set up your profile →Interview Prep
This listing calls for this tool directly. Prep for the technical screen:
Why this role
This Applied Engineer, Evaluations Role role pays $100–$120/hr like specialized consulting, minus the client overhead. You set the hours, and the assignments draw directly on your STEM background.
Talent pool
We're light on STEM candidates
We've matched 65 people with a STEM background against 726 STEM listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Mercor
Discover more opportunities on Mercor that match your skills and interests.
View All Mercor Jobs →Verified Reviews
Community Reviews
Share your experience with Mercor
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
Is Mercor for freelancers or full-time contractors?
Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.
Does Mercor's application require an on-camera interview?
Yes, every applicant records a video interview with an AI interviewer that asks questions about your resume. Clients review that recording to judge communication skills before matching, so there's no way to apply without going on camera.
Why do these AI training roles pay so much?
Because general knowledge isn't what's being tested. The model already knows the basics; what it needs is expertise on edge cases, the rare, difficult, highly technical judgment calls only a senior professional in the field would make correctly.
What does the day-to-day workload look like for elite-expert AI training roles?
Slow and deep, not fast and repetitive. A single task can take 45-60 minutes of researching citations or verifying complex calculations. Quality is what's being measured here, not throughput.
What does STEM work look like for an Applied Engineer, Evaluations Role?
Tasks here are scoped to STEM, not generic labeling. As an Applied Engineer, Evaluations Role, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to STEM) rather than following a one-size-fits-all rubric. If you don't have hands-on STEM background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding, Kotlin, and Java are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
This listing is posted at $100–$120/hr, an hourly rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.
What happens when I click Apply on this listing?
You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
How soon will I start working after applying to Mercor?
Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.