Skip to content
aitrainer.work - AI Training Jobs Platform
Generalist mindrift

Evaluation Scenario Writer - AI Agent Testing Specialist

Mindrift • United States

Education

Master's

Type

Hourly

Pay Rate

$80/hr

Listed

271d ago

Apply Now →

About this role

From the Mindrift listing

This opportunity is only for candidates currently residing in the specified country. Your location may affect eligibility and rates. Please submit your resume in English and indicate your level of English. At Mindrift , innovation meets opportunity. We believe in using the power of collective human intelligence to ethically shape the future of AI.  What we do The Mindrift platform connects specialists with AI projects from major tech innovators. Our mission is to unlock the potential of Generative AI by tapping into real-world expertise from across the globe. About the Role We’re looking for someone who can design realistic and structured evaluation scenarios for LLM-based agents. You’ll create test cases that simulate human-performed tasks and define gold-standard behavior to compare agent actions against. You’ll work to ensure each scenario is clearly defined, well-scored, and easy to execute and reuse. You’ll need a sharp analytical mindset, attention to detail, and an interest in how AI agents make decisions. Although every project is unique, you might typically: Create structured test cases that simulate complex human workflows. Define gold-standard behavior and scoring logic to evaluate agent actions. Analyze agent logs, failure modes, and decision paths. Work with code repositories and test frameworks to validate your scenarios. Iterate on prompts, instructions, and test cases to improve clarity and difficulty. Ensure that scenarios are production-ready, easy to run, and reusable. How to get started Simply apply to this post, qualify, and get the chance to contribute to projects aligned with your skills, on your own schedule. From creating training prompts to refining model responses, you’ll help shape the future of AI while ensuring technology benefits everyone.

Requirements

  • Must be eligible to work in United States
  • Fluent proficiency in English (Written & Verbal)
  • Reliable high-speed internet connection
  • Master's's degree or equivalent professional experience
  • Demonstrated expertise in Generalist

Why this role

This Evaluation Scenario Writer - AI Agent Testing Specialist opening pays $80/hr and draws on Generalist knowledge you've already built. Onboarding covers the AI training process itself.

Skills and categories

Explore other opportunities in related specializations:

Related jobs

Mindrift

Browse All Jobs from Mindrift

Discover more opportunities on Mindrift that match your skills and interests.

View All Mindrift Jobs →

Verified Reviews

Loading reviews…

Community Reviews

Loading reviews…
💬

Share your experience with Mindrift

Help other candidates make better decisions by leaving a review.

Sign in to leave a review

Common questions

What does a typical Mindrift task involve?

Read a scenario, write a short prompt for the AI (around 100 words), evaluate two AI responses to it, fact-check the outputs, then write a brief explanation (around 50 words) on which response is better against the project rubric. The focus is clarity, safety, and rule-following, not creative writing.

Who is Mindrift good for?

Freelance writers, editors, and generalist AI tutors with strong English fluency and solid research skills, no specialized tech degree required. Built by data-labeling company Toloka, Mindrift also opens higher-paying specialized projects (cybersecurity, medicine, law) to domain experts once verified.

What equipment do I need for creative AI training roles?

For voice or audio roles at this pay level, a professional home studio setup (XLR microphone, acoustically treated room) is typically required. Phone recordings and laptop mics usually get rejected by quality control.

How is my creative work used in AI training?

As ground-truth data: for writers, that means creative generation samples; for voice actors, it often means training text-to-speech models. Check the specific contract details on rights usage for your voice or likeness before signing.

What does Generalist work look like for an Evaluation Scenario Writer - AI Agent Testing Specialist?

Tasks here are scoped to Generalist, not generic labeling. As an Evaluation Scenario Writer - AI Agent Testing Specialist, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Generalist) rather than following a one-size-fits-all rubric. If you don't have hands-on Generalist background, this is likely not the right listing to start with.

What specific skills does this listing call for?

AI Training and Review & QA are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.

How much does this specific role pay?

This listing is posted at $80/hr, an hourly rate. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.

Do I need a Master's to qualify?

This role lists a master's degree as a requirement. In practice, the domain assessment is the real gate. If you can pass it, the degree is usually secondary. However, some platforms verify credentials formally, so list your actual qualifications accurately on your profile.

Can I apply from outside the United States?

This specific role is open only to people based in the United States. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply.