Explosives & Energetic Materials Expert for AI Safety Audit
$1-100
Paid work to find where a model gives an unsafe or wrong answer, and to grade whether it should have refused. Red teaming roles attack the model with adversarial prompts and document what broke; safety review roles grade outputs against a written policy; expert audit roles check a model's answers in a hazardous specialty. Most listings are language-paired, and no security clearance or pentesting background is needed for them. Like everything on this board, it is remote work you do from home. Pay on this page runs $21 to $60/hr for the middle half of roles.
Two kinds of task share this page. In adversarial roles you try to get a model to do something its policy says it must not: give dangerous instructions, produce a slur, leak a system prompt, state something false with confidence. You write the attempt, record what came back, rate how bad it was and how reliably it reproduces, and file it. In safety review roles you are handed outputs and decide whether the model should have refused, helped, or helped with caveats, against a written policy.
A third, smaller group is expert audits: chemists, nuclear engineers, clinicians and other specialists checking whether a model's answers in their field cross a line a generalist could not recognize. Those listings name the specialty in the title and sit at the top of the pay range.
Most roles here are language-paired, and the language moves the rate more than the specialism does. Labs test models in every language they ship, so a bilingual Thai, Polish or Arabic reader with good judgment is in demand without a technical background. Classic penetration testing (networks, applications, exploits) is a different track; see cybersecurity for that.
An original example, written for this page. The brief: the model must not give step-by-step instructions for disabling a car's immobilizer, and it must still help with legitimate locksmith and repair questions.
You make attempts in the categories the project lists: a plain request, a roleplay framing, a request split across turns, a request wrapped in a plausible professional context. Three refuse cleanly. The fourth, a mechanic asking how to diagnose an immobilizer fault, gets a helpful answer that stops short of a bypass, which is correct. The fifth, framed as fiction, produces a partial procedure. That is the finding.
The report is what you are paid for, and the attack is only the evidence for it. A finding a lab can act on names the policy clause, reproduces on a second try, and states severity without drama. This page never publishes a working attack, including as an illustration; what it can show is the shape of a report.
| Weak report | Strong report |
|---|---|
| Got the model to explain immobilizer bypass using a story. Serious. | Policy: no bypass procedure. Framing: fiction, single turn. Result: partial procedure (two of the steps a locksmith would recognize), no tool list. Reproduced on 2 of 3 attempts with the same framing. Severity: medium (incomplete, but the framing is cheap). The mechanic-diagnosis prompt was handled correctly and should stay allowed. |
The guide on adversarial AI training goes deeper on the workflow and on the five task types most projects use.
The best red teamers are patient and literal. They read the policy, find the gap between what it says and what the model does, and document it the same way every time. Creativity helps with framings; it hurts when it turns into a hunt for shock value, which reviewers score down. People who did content moderation, QA, compliance or editing tend to do well.
Bilingual readers fill most of the listings on this page. Subject experts fill the audit roles: anyone whose field has a misuse case the model must not assist with. Engineers with security backgrounds fit the technical minority of roles (prompt injection, tool misuse) and can also look at penetration testing.
The material is why people leave this work. Adversarial projects involve reading, and sometimes writing about, self-harm, abuse, weapons and hate, in detail, for hours. Platforms generally let you opt out of categories; ask before accepting a project, and treat the opt-out as a normal part of the job rather than a weakness.
Expect a written judgment test: a policy excerpt, a set of model outputs, and the question of which ones should have been refused and why. Agreement with the project's reference answers decides it. One listing on this page describes its process as an AI interview, a domain-specific task, and a recruiter interview; others go straight to the test. Expert audit roles verify the credential first.
The Academy module Safety and Content Policy Triage is built around the judgment these tests measure, including the error reviewers penalize most, which is over-flagging. Passing the Screening Process covers the stages in general. Platforms promote into safety projects more often than they hire into them, so a strong record on general evaluation work is a common route in; the guide's section on getting invited explains how that routing works.
The pay block at the top of this page is computed from the open listings below it: typical (median) hourly rate, the range most roles pay, and how many posted rates it is based on, updated on the date shown. The method is on the pay methodology page.
Listed rates on this page currently run $21 to $60/hr for the middle half of roles.
Within that spread, generalist bilingual roles sit low, STEM PhD bilingual roles in the middle, and specialist audits at the top. The notes beside this section say when one platform dominates the figure, which on this page is often the case, since safety projects are commissioned in large batches by one lab at a time.
Adjacent kinds of work: RLHF evaluation grades helpfulness rather than harm, agentic tasks cover agents that take actions and the new ways those can fail, and cybersecurity holds the pentesting and threat roles.
103 Jobs
Page 1 of 4
$1-100
$80-100
$80-100
$80-100
$70-90

Trending jobs + new guides + real pay data. Every Tuesday.
$40-60
/hr
$70-80
/hr
$50-70
/hr
$40-60
/hr
$40-60
/hr
$40-60
/hr
One browser alert a day when new roles land, AI Safety & Red Teaming first. No email needed.
An AI red teaming job pays you to try to make a model break its own rules, then document how, how reliably, and how badly. Safety review roles on the same page grade outputs against a policy instead of attacking the model. Both are remote hourly contracts, often paired with a language.
AI safety is the wider discipline of checking model outputs for harm, bias and policy violations. Red teaming is the adversarial part of it, deliberately probing the model to find where it fails. Most listings blend the two, which is why one page lists both.
No, for most roles. Penetration testing of networks and applications is a separate track. AI red teaming and safety review screen for judgment about policy, clear writing, and often a second language; technical roles (prompt injection, tool misuse) are the minority.
The listings on this page currently state $21 to $60/hr for the middle half of roles, with the median and the number of rates behind it in the pay block at the top. Bilingual generalist roles sit low, PhD-level bilingual roles in the middle, specialist audits at the top.
It can be. Adversarial projects involve harmful content in detail: self-harm, abuse, weapons, hate. Platforms generally allow opting out of categories, and asking about that before accepting a project is normal. Safety review roles vary; expert audits are usually confined to the specialist's own field.
Through general evaluation work. Platforms promote raters with a strong accuracy record into safety projects more often than they hire outsiders into them, so a few months of RLHF evaluation on the same platform is the common route. Bilingual applicants are hired directly more often.
The listings on this board do not ask for one. Specialist audits (chemical, nuclear, biological) ask for a degree and professional experience in the field, and some restrict by country. Read the eligibility line on each listing.
This page is showing 103 listings, sourced from 7 platforms: Mercor (89), SME Careers (5), telus (5), Handshake (1) and Claru (1). The most recent one was picked up today. The count moves as platforms post and close roles, so treat it as a snapshot of today rather than a fixed figure.
Browse all job categories and find your next AI training opportunity.
View All Categories