AI Safety Red Teamer
Mercor • Remote
Education
Any
Type
hourly
Pay Rate (by country)
$70–$84/hr
Listed
13d ago
✅ Applying through this link supports our platform at no cost to you.
This position is hosted on an external talent platform. Please only apply for this position if it fits your skills and interests.
In our Talent Pool?
Apply through this link and we can vouch for you to Mercor. ? We vouch for Talent Pool members who apply through our referral link, when we believe they're a strong match. Not every applicant gets a vouch. Not in the pool yet? Set up your profile first.
Set up your profile →Mercor: our referral track record
We've referred 191 candidates to Mercor roles. 14% (27) were placed.
About this Role
From the Mercor listing
We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics.
Responsibilities
Design adversarial prompts to stress-test frontier AI models.
Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.
Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.
Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.
Collaborate with AI researchers to improve model alignment, robustness, and safety.
Required Qualifications
Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.
5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.
Strong analytical reasoning, prompt design, and written communication skills.
Experience designing adversarial prompts or evaluating frontier AI systems.
Preferred Qualifications
Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.
Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.
Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.
Why Join?
Help secure and strengthen the next generation of frontier AI models.
Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams.
Influence how AI systems respond to complex, real-world safety challenges.
Requirements
- Must be eligible to work in Remote
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
- Bachelor's degree or equivalent professional experience
- Demonstrated expertise in Software Engineering
Why This Role
No office, no fixed hours, no relocation. This AI Safety Red Teamer role pays $77/hr fully remote, giving you access to Software Engineering work that would otherwise be limited to a handful of major cities.
Skills & Categories
Explore other opportunities in related specializations:
Related Jobs
Engineering & Data tools Specialist
micro1 • Software Engineering
$80 /hr
QA / Software Engineering Reviewer – Browser Test Validation
mercor • Software Engineering
$60 /hr
Head of AI & Engineering Expert
ethos • Software Engineering
$150 /hr
Senior Machine Learning Engineer / Model Evaluations Expert
ethos • Software Engineering
$125 /hr
Browse All Jobs from Mercor
Discover more opportunities on Mercor that match your skills and interests.
View All Mercor Jobs →Verified Reviews
Community Reviews
Share your experience with Mercor
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Frequently Asked Questions
Is Mercor for freelancers or full-time contractors?
Mercor places you with one client for a defined engagement, like 'Python Tutor for 3 months', rather than having you grab small tasks from a shared queue. Most roles function as steady contract work, not one-off gigs.
Does Mercor's application require an on-camera interview?
Yes, every applicant records a video interview with an AI interviewer that asks questions about your resume. Clients review that recording to judge communication skills before matching, so there's no way to apply without going on camera.
Does it cost money to apply to Mercor?
No, applying and joining Mercor is free. Mercor's revenue comes from a fee it charges the client on top of your hourly rate, not from applicants. Treat any request for payment to join as a red flag.
What does task-based AI training work actually look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does Software Engineering work look like for a AI Safety Red Teamer?
Tasks here are scoped to Software Engineering, not generic labeling. As a AI Safety Red Teamer, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Software Engineering) rather than following a one-size-fits-all rubric. If you don't have hands-on Software Engineering background, this is likely not the right listing to start with.
What happens when I click Apply on this listing?
You'll be taken to Mercor's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
How soon will I start working after applying to Mercor?
Not immediately. Mercor is a talent marketplace, not a task queue, so applying puts you in a pool of candidates. You start working only once a specific client, like a major AI lab, selects your profile, and that matching process can take weeks.