Software Engineers: Paid Code Review for AI Agent Evaluation
$60-80
/hr
Paid work on AI agents, the models that run commands, click through software and call tools instead of only writing text. Three kinds of role: environment builders write the sandboxed tasks and graders agents train on, trajectory specialists record how a competent person completes them, and agent evaluators grade what the agent did, step by step. All three are remote contracts worked from home. Pay on this page runs $79 to $85/hr for the middle half of roles.
An agent is a model that does things rather than describing them: it runs a command, clicks through a settings page, calls an API, edits a file, and checks its own result. Training one needs three kinds of human work, and the listings on this page cover all three.
Environment building means writing the task an agent will practice on: a starting state (a repository with a bug, a spreadsheet with errors, a half-finished marketing brief), a goal, the actions available, and a grader that scores the attempt automatically. Listings say "RL environment builder", "scenario designer" or "benchmarks and RL environments". The RL stands for reinforcement learning, the training loop that rewards the agent when the grader says it succeeded.
Trajectory work means doing the task yourself inside the environment while every step is recorded, so the model learns from a competent demonstration. Titles include "trajectory specialist", "computer use annotator" and "agent trace collection". It feels like being screen-recorded doing your normal job, with a requirement to narrate why.
Agent evaluation means reading what an agent did, step by step, and grading it: did it take sane actions, did it finish, did it cheat the grader, would a senior colleague have accepted the result. Titles say "agent evaluation analyst", "agentic workflow reviewer" or "task auditor". Of the three, this is the closest to RLHF evaluation, with the difference that you are judging a sequence of actions, not a paragraph.
An original example, written for this page. You are asked to build an environment for an accounting agent.
You write the starting state: a 40-row expense ledger and a folder of receipts, where three rows disagree with their receipt (a transposed amount, a duplicated entry, a receipt with no row). The goal: reconcile the ledger and file a correction note. The grader you write checks that exactly those three rows were fixed and that the note names each one, and it penalizes any edit to a row that was already correct. Then you record a reference run of yourself doing it.
Two agent runs come back for review. Run A changes all 40 amounts to match the receipt totals, which makes the ledger sum correct and would pass a grader that only checked the total. Run B finds the transposed amount and the duplicate, misses the missing row, and writes an accurate note. Run A is the one that matters: it found a hole in your grader, which is the failure environment builders are paid to prevent. You tighten the grader, score B as partial with the miss named, and flag A as reward hacking so the lab can use it as a regression test.
In trajectory roles you would be the reference run: reconciling the ledger yourself in the sandbox, with each click and keystroke logged, and a short note at each decision point. The logs are usually JSON, which is why several listings ask for comfort with structured data; the Academy module Markdown and JSON Formatting covers what reviewers check.
Software engineers are the largest group: listings ask for a language (Python most often, then Go, Rust and TypeScript), experience working with coding agents, and the ability to write a test an agent cannot game. Benchmarks like SWE-bench and terminal-bench appear in titles because labs want people who can write tasks in that shape.
The second group is professionals who live in a specific tool: accountants in QuickBooks, marketers in a CRM and an ad manager, operations people in ticketing systems. Those listings want you to build or demonstrate the workflow an agent should learn, and they do not ask for code. Most of them are posted by one platform at a time in batches, so the list below can swing from twenty such roles to none in a month.
It does not suit people who dislike adversarial thinking. Every grader you write will be attacked by a model that is better at finding loopholes than most humans, and the job is to assume that and close them. If you prefer to judge finished text, RLHF evaluation is the better fit.
Engineering roles run a live coding or debugging round, sometimes automated, and often ask to see public work: one listing on this page recruits by GitHub contribution history. Expect to explain how you would write a grader for a task and how an agent might game it. The Academy module Live Code Review and Debugging Interviews is built for this round, and Python for AI Training covers the language most listings name.
Domain roles screen with a short conversational interview and a sample task: build one scenario in your tool, or walk through how you would, with the steps written out. Where listings on this page state a screening step at all, it is an interview; most say nothing, so read each posting.
The pay block at the top of this page is computed from the open listings below: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. The method is on the pay methodology page.
Listed rates on this page currently run $79 to $85/hr for the middle half of roles.
The middle half is narrow because most listings here post a single rate for an engineering tier rather than a range, and a platform often posts the same rate across a batch of roles. The notes beside this section say when one platform or our own estimates dominate the figure. Trajectory and domain-workflow roles tend to sit below the engineering rate; benchmark and environment design roles at or above it.
Adjacent kinds of work: RLHF evaluation grades what a model said, red teaming tries to make it misbehave, and the software engineering page lists code review roles that do not involve agents.
62 Jobs
Page 1 of 3
$60-80
/hr
$50-70
/hr
$90-120
/hr

Trending jobs + new guides + real pay data. Every Tuesday.
$80-90
/hr
$50-70
/hr
$25-60
/hr · estimate
$90-130
/hr
$80-100
/hr
$60-100
/hr
$60-100
/hr
One browser alert a day when new roles land, Agentic Tasks & RL Environments first. No email needed.
An RL environment job pays you to build the tasks an AI agent practices on: a starting state, a goal, the actions allowed, and a grader that scores each attempt automatically. RL stands for reinforcement learning, the loop that rewards the agent when the grader says it succeeded. Most roles are remote hourly contracts; some are paid per environment.
No. The training algorithm is the lab's job. Listings ask for the ability to design a task with a clear success condition, write a grader that cannot be gamed, and, for trajectory roles, do the task competently while being recorded.
A trajectory is the full record of one attempt at a task: every action the agent (or the human demonstrator) took, what it saw after each action, and the final result. Trajectory specialists produce reference runs; agent evaluators grade recorded runs step by step.
Reward hacking is an agent finding a way to make the grader report success without doing the task, like editing a test instead of fixing the bug. Environment builders are paid largely to prevent it, and agent evaluators are paid to catch it in recorded runs.
The listings on this page currently state $79 to $85/hr for the middle half of roles, with the median and the number of rates behind it in the pay block at the top. Engineering roles set the typical rate; trajectory and domain-workflow roles sit below it, environment design at or above it.
Yes, for a subset. Platforms post batches of roles for professionals fluent in a specific tool (accounting software, marketing systems, operations platforms) to build or demonstrate workflows. Those roles need no code, and they appear and fill in batches, so set the alert on this page.
The platform filter above the list shows which platforms have open roles here today. The mix changes month to month because labs commission this work in projects of a few weeks, and a platform that posts twenty environment roles in one batch may post none the next.
This page is showing 62 listings, sourced from 8 platforms: Mercor (31), Turing (13), Terac (6), Mindrift (6) and innodata (2). The most recent one was picked up today. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.
Browse all job categories and find your next AI training opportunity.
View All Categories