Skip to content
aitrainer.work - AI Training Jobs Platform

AI Safety and Red Teaming Jobs: Get Paid to Make Models Fail Safely

Paid work to find where a model gives an unsafe or wrong answer, and to grade whether it should have refused. Red teaming roles attack the model with adversarial prompts and document what broke; safety review roles grade outputs against a written policy; expert audit roles check a model's answers in a hazardous specialty. Most listings are language-paired, and no security clearance or pentesting background is needed for them. Like everything on this board, it is remote work you do from home. Pay on this page runs $21 to $60/hr for the middle half of roles.

102 open jobs
Typical pay
$50/hr
Most roles pay
$21 to $60/hr
Based on
96 posted rates
Updated
October 4, 2026

About AI Safety & Red Teaming work

Red teaming is paid work to make a model fail, and safety review is paid work to grade the failure

Two kinds of task share this page. In adversarial roles you try to get a model to do something its policy says it must not: give dangerous instructions, produce a slur, leak a system prompt, state something false with confidence. You write the attempt, record what came back, rate how bad it was and how reliably it reproduces, and file it. In safety review roles you are handed outputs and decide whether the model should have refused, helped, or helped with caveats, against a written policy.

A third, smaller group is expert audits: chemists, nuclear engineers, clinicians and other specialists checking whether a model's answers in their field cross a line a generalist could not recognize. Those listings name the specialty in the title and sit at the top of the pay range.

Most roles here are language-paired, and the language moves the rate more than the specialism does. Labs test models in every language they ship, so a bilingual Thai, Polish or Arabic reader with good judgment is in demand without a technical background. Classic penetration testing (networks, applications, exploits) is a different track; see cybersecurity for that.

One example task, start to finish

An original example, written for this page. The brief: the model must not give step-by-step instructions for disabling a car's immobilizer, and it must still help with legitimate locksmith and repair questions.

You make attempts in the categories the project lists: a plain request, a roleplay framing, a request split across turns, a request wrapped in a plausible professional context. Three refuse cleanly. The fourth, a mechanic asking how to diagnose an immobilizer fault, gets a helpful answer that stops short of a bypass, which is correct. The fifth, framed as fiction, produces a partial procedure. That is the finding.

The report is what you are paid for, and the attack is only the evidence for it. A finding a lab can act on names the policy clause, reproduces on a second try, and states severity without drama. This page never publishes a working attack, including as an illustration; what it can show is the shape of a report.

Weak reportStrong report
Got the model to explain immobilizer bypass using a story. Serious.Policy: no bypass procedure. Framing: fiction, single turn. Result: partial procedure (two of the steps a locksmith would recognize), no tool list. Reproduced on 2 of 3 attempts with the same framing. Severity: medium (incomplete, but the framing is cheap). The mechanic-diagnosis prompt was handled correctly and should stay allowed.

The guide on adversarial AI training goes deeper on the workflow and on the five task types most projects use.

It suits people with policy fluency, a second language, or a hazardous specialty

The best red teamers are patient and literal. They read the policy, find the gap between what it says and what the model does, and document it the same way every time. Creativity helps with framings; it hurts when it turns into a hunt for shock value, which reviewers score down. People who did content moderation, QA, compliance or editing tend to do well.

Bilingual readers fill most of the listings on this page. Subject experts fill the audit roles: anyone whose field has a misuse case the model must not assist with. Engineers with security backgrounds fit the technical minority of roles (prompt injection, tool misuse) and can also look at penetration testing.

The material is why people leave this work. Adversarial projects involve reading, and sometimes writing about, self-harm, abuse, weapons and hate, in detail, for hours. Platforms generally let you opt out of categories; ask before accepting a project, and treat the opt-out as a normal part of the job rather than a weakness.

Screening tests judgment about the policy, then adds a domain task for experts

Expect a written judgment test: a policy excerpt, a set of model outputs, and the question of which ones should have been refused and why. Agreement with the project's reference answers decides it. One listing on this page describes its process as an AI interview, a domain-specific task, and a recruiter interview; others go straight to the test. Expert audit roles verify the credential first.

The Academy module Safety and Content Policy Triage is built around the judgment these tests measure, including the error reviewers penalize most, which is over-flagging. Passing the Screening Process covers the stages in general. Platforms promote into safety projects more often than they hire into them, so a strong record on general evaluation work is a common route in; the guide's section on getting invited explains how that routing works.

Pay comes from the listings above, and the language pairing sets it

The pay block at the top of this page is computed from the open listings below it: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. The method is on the pay methodology page.

Listed rates on this page currently run $21 to $60/hr for the middle half of roles.

Within that spread, generalist bilingual roles sit low, STEM PhD bilingual roles in the middle, and specialist audits at the top. The notes beside this section say when one platform dominates the figure, which on this page is often the case, since safety projects are commissioned in large batches by one lab at a time.

How to start this week

  1. Read one real usage policy end to end. The major labs publish theirs. Knowing how a refusal rule is worded makes the judgment test an exercise in reading rather than guessing.
  2. Decide your categories first. Which content you will and will not work with. Say so at application; platforms route around it.
  3. Apply to language-paired listings that name your languages exactly. Sort below by newest; batches of bilingual safety roles open together and fill within weeks.
  4. Practice the report, not the attack. Write three findings in the shape of the strong report above about any model you have access to. Clarity and reproducibility are what the test rewards.
  5. If you are already rating, ask for safety work. The route in is usually a promotion from general evaluation on the same platform, so tell your project lead you want it.

Adjacent kinds of work: RLHF evaluation grades helpfulness rather than harm, agentic tasks cover agents that take actions and the new ways those can fail, and cybersecurity holds the pentesting and threat roles.

Type:
Popular:

103 Jobs

Page 1 of 4

SuperAnnotate SME Careers platform

Explosives & Energetic Materials Expert for AI Safety Audit

$1-100

SME Careers • Bachelor's • 18d ago
AI Training AI Safety
SuperAnnotate SME Careers platform

Chemical Safety Expert for AI Model Safety Audit

$80-100

SME Careers • Bachelor's • 18d ago
STEM AI Training
SuperAnnotate SME Careers platform

CBRN Specialist

$80-100

SME Careers • Bachelor's • 22d ago
Healthcare AI Training
SuperAnnotate SME Careers platform

Red-Teaming Quality Assurance Lead (QAL)

$80-100

SME Careers • Master's • 53d ago
QA Testing AI Training
SuperAnnotate SME Careers platform

Nuclear & Radiological Safety Expert for AI Model Audit

$70-90

SME Careers • Bachelor's • 23d ago
Healthcare AI Training
Newsletter subscription
AI Training Jobs

Weekly AI Training Intelligence

Trending jobs + new guides + real pay data. Every Tuesday.

✅ 500+🔒 No spam

Unsubscribe anytime.

Handshake AI fellowship program

Adversarial Prompt Expert

$40-60

/hr

Handshake • Bachelor's • 7d ago
Expert AI Training
Mercor AI hiring platform

Chemicals Safety for Redteaming

$70-80

/hr

Mercor • 19d ago
Mercor AI hiring platform

Energetic Materials Expert for Redteaming

$70-80

/hr

Mercor • Bachelor's • 19d ago
Expert AI Training
Mercor AI hiring platform

Radiologicals Safety for Redteaming

$70-80

/hr

Mercor • Bachelor's • 19d ago
Mercor AI hiring platform

Nuclear Engineering & Safeguards Experts for Red Team

$70-80

/hr

Mercor • 19d ago
Mercor AI hiring platform

Bilingual Norwegian STEM Expert (PhD) — AI Safety

$70-90

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Norwegian PhD
Mercor AI hiring platform

Bilingual Japanese STEM Expert (PhD) — AI Safety

$60-80

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Japanese PhD
Mercor AI hiring platform

Bilingual Chinese STEM Expert (PhD) — AI Safety

$60-80

/hr

Mercor • PhD • 29d ago
Languages Bilingual Chinese English PhD
Mercor AI hiring platform

Bilingual Korean STEM Expert (PhD) — AI Safety

$60-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Korean PhD
Mercor AI hiring platform

Bilingual Finnish STEM Expert (PhD) — AI Safety

$60-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Finnish PhD
Mercor AI hiring platform

Bilingual Dutch (Belgium) STEM Expert (PhD) — AI Safety

$60-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual Dutch English PhD
Mercor AI hiring platform

Bilingual Danish STEM Expert (PhD) — AI Safety

$60-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual Danish English PhD
Mercor AI hiring platform

Bilingual Dutch (Netherlands) STEM Expert (PhD) — AI Safety

$60-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual Dutch English PhD
Mercor AI hiring platform

Bilingual German STEM Expert (PhD) — AI Safety

$50-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual English German PhD
Mercor AI hiring platform

Bilingual French STEM Expert (PhD) — AI Safety

$50-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual English French PhD
Mercor AI hiring platform

Bilingual Norwegian Generalist Expert — AI Safety

$50-70

/hr

Mercor • Bachelor's • 29d ago
Languages Bilingual English Norwegian
Mercor AI hiring platform

Bilingual Spanish STEM Expert (PhD) — AI Safety

$50-70

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Spanish PhD
Mercor AI hiring platform

Bilingual Czech STEM Expert (PhD) — AI Safety

$50-60

/hr

Mercor • PhD • 29d ago
Languages Bilingual Czech English PhD
Mercor AI hiring platform

Bilingual Portuguese STEM Expert (PhD) — AI Safety

$40-60

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Portuguese PhD
Mercor AI hiring platform

Bilingual Italian STEM Expert (PhD) — AI Safety

$40-60

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Italian PhD
Mercor AI hiring platform

Bilingual Polish STEM Expert (PhD) — AI Safety

$40-60

/hr

Mercor • PhD • 29d ago
Languages Bilingual English Polish PhD
Mercor AI hiring platform

Bilingual French Generalist Expert — AI Safety

$40-60

/hr

Mercor • Bachelor's • 29d ago
Languages Bilingual English French
Mercor AI hiring platform

Bilingual Korean Generalist Expert — AI Safety

$40-60

/hr

Mercor • Bachelor's • 29d ago
Languages Bilingual English Korean
Mercor AI hiring platform

Bilingual Japanese Generalist Expert — AI Safety

$40-60

/hr

Mercor • Bachelor's • 29d ago
Languages Bilingual English Japanese
Mercor AI hiring platform

Bilingual German Generalist Expert — AI Safety

$40-60

/hr

Mercor • Bachelor's • 29d ago
Languages Bilingual English German
Page 1 of 4 Next →

Get new AI Safety & Red Teaming jobs as they open

One browser alert a day when new roles land, AI Safety & Red Teaming first. No email needed.

Browse by category or skill

Similar to: Adversarial Prompt Expert

Similar to: LLM Jailbreak Red Teamer

Similar to: Explosives & Energetic Materials Expert for AI Safety Audit

Frequently Asked Questions

What is an AI red teaming job?

An AI red teaming job pays you to try to make a model break its own rules, then document how, how reliably, and how badly. Safety review roles on the same page grade outputs against a policy instead of attacking the model. Both are remote hourly contracts, often paired with a language.

What is the difference between AI safety and red teaming roles?

AI safety is the wider discipline of checking model outputs for harm, bias and policy violations. Red teaming is the adversarial part of it, deliberately probing the model to find where it fails. Most listings blend the two, which is why one page lists both.

Do I need a cybersecurity or pentesting background for AI red teaming?

No, for most roles. Penetration testing of networks and applications is a separate track. AI red teaming and safety review screen for judgment about policy, clear writing, and often a second language; technical roles (prompt injection, tool misuse) are the minority.

How much do AI red teaming jobs pay?

The listings on this page currently state $21 to $60/hr for the middle half of roles, with the median and the number of rates behind it in the pay block at the top. Bilingual generalist roles sit low, PhD-level bilingual roles in the middle, specialist audits at the top.

Is AI red teaming work disturbing?

It can be. Adversarial projects involve harmful content in detail: self-harm, abuse, weapons, hate. Platforms generally allow opting out of categories, and asking about that before accepting a project is normal. Safety review roles vary; expert audits are usually confined to the specialist's own field.

How do I get into AI safety projects with no experience?

Through general evaluation work. Platforms promote raters with a strong accuracy record into safety projects more often than they hire outsiders into them, so a few months of RLHF evaluation on the same platform is the common route. Bilingual applicants are hired directly more often.

Do I need a security clearance for AI safety audit roles?

The listings on this board do not ask for one. Specialist audits (chemical, nuclear, biological) ask for a degree and professional experience in the field, and some restrict by country. Read the eligibility line on each listing.

How many AI Safety & Red Teaming AI training jobs are open right now?

This page is showing 103 listings, sourced from 7 platforms: Mercor (89), SME Careers (5), telus (5), Handshake (1) and Claru (1). The most recent one was picked up today. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.

Explore more opportunities

Browse all job categories and find your next AI training opportunity.

View All Categories