Skip to content
aitrainer.work - AI Training Jobs Platform

Agentic AI and RL Environment Jobs: Build and Grade the Tasks Agents Learn From

Paid work on AI agents, the models that run commands, click through software and call tools instead of only writing text. Three kinds of role: environment builders write the sandboxed tasks and graders agents train on, trajectory specialists record how a competent person completes them, and agent evaluators grade what the agent did, step by step. All three are remote contracts worked from home. Pay on this page runs $79 to $85/hr for the middle half of roles.

63 open jobs
Typical pay
$80/hr
Most roles pay
$79 to $85/hr
Based on
48 posted rates
Updated
October 4, 2026

About Agentic Tasks & RL Environments work

Agentic work trains models to act, so the data is tasks, demonstrations and graded runs

An agent is a model that does things rather than describing them: it runs a command, clicks through a settings page, calls an API, edits a file, and checks its own result. Training one needs three kinds of human work, and the listings on this page cover all three.

Environment building means writing the task an agent will practice on: a starting state (a repository with a bug, a spreadsheet with errors, a half-finished marketing brief), a goal, the actions available, and a grader that scores the attempt automatically. Listings say "RL environment builder", "scenario designer" or "benchmarks and RL environments". The RL stands for reinforcement learning, the training loop that rewards the agent when the grader says it succeeded.

Trajectory work means doing the task yourself inside the environment while every step is recorded, so the model learns from a competent demonstration. Titles include "trajectory specialist", "computer use annotator" and "agent trace collection". It feels like being screen-recorded doing your normal job, with a requirement to narrate why.

Agent evaluation means reading what an agent did, step by step, and grading it: did it take sane actions, did it finish, did it cheat the grader, would a senior colleague have accepted the result. Titles say "agent evaluation analyst", "agentic workflow reviewer" or "task auditor". Of the three, this is the closest to RLHF evaluation, with the difference that you are judging a sequence of actions, not a paragraph.

One example task, start to finish

An original example, written for this page. You are asked to build an environment for an accounting agent.

You write the starting state: a 40-row expense ledger and a folder of receipts, where three rows disagree with their receipt (a transposed amount, a duplicated entry, a receipt with no row). The goal: reconcile the ledger and file a correction note. The grader you write checks that exactly those three rows were fixed and that the note names each one, and it penalizes any edit to a row that was already correct. Then you record a reference run of yourself doing it.

Two agent runs come back for review. Run A changes all 40 amounts to match the receipt totals, which makes the ledger sum correct and would pass a grader that only checked the total. Run B finds the transposed amount and the duplicate, misses the missing row, and writes an accurate note. Run A is the one that matters: it found a hole in your grader, which is the failure environment builders are paid to prevent. You tighten the grader, score B as partial with the miss named, and flag A as reward hacking so the lab can use it as a regression test.

In trajectory roles you would be the reference run: reconciling the ledger yourself in the sandbox, with each click and keystroke logged, and a short note at each decision point. The logs are usually JSON, which is why several listings ask for comfort with structured data; the Academy module Markdown and JSON Formatting covers what reviewers check.

It suits engineers and people who are fluent in the software of their own trade

Software engineers are the largest group: listings ask for a language (Python most often, then Go, Rust and TypeScript), experience working with coding agents, and the ability to write a test an agent cannot game. Benchmarks like SWE-bench and terminal-bench appear in titles because labs want people who can write tasks in that shape.

The second group is professionals who live in a specific tool: accountants in QuickBooks, marketers in a CRM and an ad manager, operations people in ticketing systems. Those listings want you to build or demonstrate the workflow an agent should learn, and they do not ask for code. Most of them are posted by one platform at a time in batches, so the list below can swing from twenty such roles to none in a month.

It does not suit people who dislike adversarial thinking. Every grader you write will be attacked by a model that is better at finding loopholes than most humans, and the job is to assume that and close them. If you prefer to judge finished text, RLHF evaluation is the better fit.

Screening is a technical interview for engineers and a workflow walkthrough for everyone else

Engineering roles run a live coding or debugging round, sometimes automated, and often ask to see public work: one listing on this page recruits by GitHub contribution history. Expect to explain how you would write a grader for a task and how an agent might game it. The Academy module Live Code Review and Debugging Interviews is built for this round, and Python for AI Training covers the language most listings name.

Domain roles screen with a short conversational interview and a sample task: build one scenario in your tool, or walk through how you would, with the steps written out. Where listings on this page state a screening step at all, it is an interview; most say nothing, so read each posting.

Pay comes from the listings above, and most of them post one engineering rate

The pay block at the top of this page is computed from the open listings below: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. The method is on the pay methodology page.

Listed rates on this page currently run $79 to $85/hr for the middle half of roles.

The middle half is narrow because most listings here post a single rate for an engineering tier rather than a range, and a platform often posts the same rate across a batch of roles. The notes beside this section say when one platform or our own estimates dominate the figure. Trajectory and domain-workflow roles tend to sit below the engineering rate; benchmark and environment design roles at or above it.

How to start this week

  1. Pick the track that matches what you already do. Engineers: environment building and agentic coding. Tool-fluent professionals: trajectory and workflow roles. Careful readers with some technical comfort: agent evaluation.
  2. Build one environment on your own before applying. A task with a starting state, a goal, a grader and a reference run, in a public repository, is the strongest application material the listings on this page respond to. Keep it small and make the grader hard to game.
  3. Prepare for the adversarial question. "How would an agent cheat this?" comes up in interviews and in the work. Practice answering it for your own task.
  4. Apply to the listings sorted by newest. Batches here open and close quickly; a role that is three weeks old is often filled.
  5. Keep your JSON and Markdown clean. Trajectory logs and task specs are reviewed as structured data, and formatting errors are the fastest way to fail a trial batch.

Adjacent kinds of work: RLHF evaluation grades what a model said, red teaming tries to make it misbehave, and the software engineering page lists code review roles that do not involve agents.

Type:
Popular:

62 Jobs

Page 1 of 3

New job posted today
Terac paid research studies platform

Software Engineers: Paid Code Review for AI Agent Evaluation

$60-80

/hr

Terac • 🆕 Today
New job posted today
Mercor AI hiring platform

Task Writer: Project Agentic-MME

$50-70

/hr

Mercor • 1d ago
Writing AI Training
Ethos expert network platform

RL Engineer Expert

$90-120

/hr

Ethos • 128d ago
Terac paid research studies platform

Architecture Operations Specialists: Structured Data for Post-Training Evaluation

$90-120

/hr

Terac • 5d ago
AI Training Review & QA
Terac paid research studies platform

SWE Coding Benchmarks & RL Environments

$70-90

/hr

Terac • 8d ago
Newsletter subscription
AI Training Jobs

Weekly AI Training Intelligence

Trending jobs + new guides + real pay data. Every Tuesday.

✅ 500+🔒 No spam

Unsubscribe anytime.

Engagement Manager, Agentic AI Workflow Evaluations

$80-90

/hr

innodata • 14d ago
Business AI Training

AI Agentic Workflow Reviewer

$50-60

/hr

innodata • 14d ago
Meridial expert contractor network, by Invisible Technologies

SWE-Bench AI Task Auditor - Freelance AI Trainer Project

$50-70

/hr

Meridial • 19d ago
STEM AI Training
Turing remote developer platform

Tech Lead - Agent Development

$15-35

/hr · estimate

Turing • 24d ago
Turing remote developer platform

LLM Trainer – Terminal-Bench (Python & Linux Systems)

$10-30

/hr · estimate

Turing • 24d ago
Turing remote developer platform

OSWorld GUI Data Annotator (Ubuntu Desktop)

$15-30

/hr · estimate

Turing • 24d ago
Turing remote developer platform

Agentic Tasker (Frontier STEM)

$15-30

/hr · estimate

Turing • Bachelor's • 25d ago
Turing remote developer platform

Software / Platform / Data Engineer — Agent Trace Collection

$25-60

/hr · estimate

Turing • Bachelor's • 26d ago
Terac paid research studies platform

RL Environment Builder: Experienced QuickBooks User (Accounting and Sales Operations)

$90-130

/hr

Terac • 18d ago
Terac paid research studies platform

Machine Learning Engineers: Scenario Building for Reinforcement Learning

$80-100

/hr

Terac • 18d ago
Data Science AI Training
Mercor AI hiring platform

SWE-Bench Task Auditor

$70-90

/hr

Mercor • 33d ago
Accounting Coding TypeScript Django
Micro1 AI training platform

MCP Expert

$60-120

Micro1 • Master's • 39d ago
Mercor AI hiring platform

Marketing Expert — Brand Marketing (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Field Marketing (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Marketing Operations (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Lifecycle (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — PR & Communications (AI Agent Environments)

$60-100

/hr

Mercor • Bachelor's • 39d ago
Mercor AI hiring platform

Marketing Expert — Creator & Influencer (AI Agent Environments)

$60-100

/hr

Mercor • Master's • 39d ago
Mercor AI hiring platform

Marketing Expert — Social Media (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Content Marketing (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Product Marketing (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Paid Growth (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Marketing Analytics (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Mercor AI hiring platform

Marketing Expert — Organic Growth (AI Agent Environments)

$60-100

/hr

Mercor • 39d ago
Terac paid research studies platform

RL Environment Creation - World Building [Various Domains]

$80-100

/hr

Terac • 38d ago
AI Training Content Review
Page 1 of 3 Next →

Get new Agentic Tasks & RL Environments jobs as they open

One browser alert a day when new roles land, Agentic Tasks & RL Environments first. No email needed.

Browse by category or skill

Similar to: Software Engineers: Paid Code Review for AI Agent Evaluation

Similar to: Architecture Operations Specialists: Structured Data for Post-Training Evaluation

Frequently Asked Questions

What is an RL environment job?

An RL environment job pays you to build the tasks an AI agent practices on: a starting state, a goal, the actions allowed, and a grader that scores each attempt automatically. RL stands for reinforcement learning, the loop that rewards the agent when the grader says it succeeded. Most roles are remote hourly contracts; some are paid per environment.

Do I need to know reinforcement learning to do agentic AI work?

No. The training algorithm is the lab's job. Listings ask for the ability to design a task with a clear success condition, write a grader that cannot be gamed, and, for trajectory roles, do the task competently while being recorded.

What is an agent trajectory or trace?

A trajectory is the full record of one attempt at a task: every action the agent (or the human demonstrator) took, what it saw after each action, and the final result. Trajectory specialists produce reference runs; agent evaluators grade recorded runs step by step.

What is reward hacking and why do listings mention it?

Reward hacking is an agent finding a way to make the grader report success without doing the task, like editing a test instead of fixing the bug. Environment builders are paid largely to prevent it, and agent evaluators are paid to catch it in recorded runs.

How much do agentic AI and RL environment jobs pay?

The listings on this page currently state $79 to $85/hr for the middle half of roles, with the median and the number of rates behind it in the pay block at the top. Engineering roles set the typical rate; trajectory and domain-workflow roles sit below it, environment design at or above it.

Can I do RL environment work without being a software engineer?

Yes, for a subset. Platforms post batches of roles for professionals fluent in a specific tool (accounting software, marketing systems, operations platforms) to build or demonstrate workflows. Those roles need no code, and they appear and fill in batches, so set the alert on this page.

Which platforms post agentic and RL environment work?

The platform filter above the list shows which platforms have open roles here today. The mix changes month to month because labs commission this work in projects of a few weeks, and a platform that posts twenty environment roles in one batch may post none the next.

How many Agentic Tasks & RL Environments AI training jobs are open right now?

This page is showing 62 listings, sourced from 8 platforms: Mercor (31), Turing (13), Terac (6), Mindrift (6) and innodata (2). The most recent one was picked up today. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.

Explore more opportunities

Browse all job categories and find your next AI training opportunity.

View All Categories