AI Benchmark Engineer — Knowledge / Research
Turing • Remote, Bangladesh, Brazil, Colombia, Egypt, India, Indonesia, Kenya, Nigeria, Turkey, Vietnam, Ghana
Education
Not stated
Type
Hourly
Listed
160d ago
Apply opens Turing in a new tab.
Apply Now → ⚡ Boost your chances - Optimize your resume with Rezi.aiAbout this role
From the Turing listing
About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We are seeking a highly analytical and computationally proficient individual to join our team with a strong research background. You will be instrumental in contributing to this role by either crafting challenging and insightful problems in your respective research domain, devising elegant computational solutions.
Responsibilities
- Build multi-agent benchmark tasks that require reading, analyzing, and synthesizing large document collections
- Curate real-world research corpora — academic papers, case studies, technical reports — and design questions that require comprehensive analysis
- Write structured ground-truth oracles (JSON) with specific, verifiable answers that prove the agent actually read the source material
- Design LLM judge prompts that evaluate agent output field-by-field against the oracle
- Create decomposition guides that split research across multiple parallel sub-agents (one per document, one per domain, then synthesis)
Required Qualifications
- 5+ years of research experience (academic or industry) in any scientific domain
- Strong reading comprehension with ability to extract structured data from unstructured text
- Experience with JSON and data structures, including schema design and output validation
- Proficiency in Python scripting for data processing and evaluation (e.g., judge scripts)
- Familiarity with AI coding benchmarks such as SWE-bench and Terminal-bench
- Hands-on experience with Docker (writing Dockerfiles, building images, debugging containers)
- High attention to detail, especially for creating precise evaluation oracles without approximations
Nice to have
- Experience with systematic reviews, meta-analyses, or large-scale literature surveys
- Familiarity with medical, legal, or scientific document analysis
- Experience with NLP or information extraction tasks
- Knowledge of LLM evaluation and benchmarking (e.g., MMLU, GPQA, SimpleQA)
- Experience curating datasets for AI evaluation
Perks of Freelancing With Turing
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
- Potential for contract extension based on performance and project needs.
Offer Details
- Commitments Required : 40 hours /week with 4 hours of PST Overlap
- Engagement type : Contractor assignment/freelancer (no medical/paid leave)
- Duration of contract : 1 month; [expected start date is next week]
- Location: Bangladesh, Brazil, Colombia, Egypt, Ghana, India, Indonesia, Kenya, Nigeria, Turkey, Vietnam
Requirements
- 5+ years of research experience (academic or industry) in any scientific domain
- Strong reading comprehension with ability to extract structured data from unstructured text
- Experience with JSON and data structures, including schema design and output validation
- Proficiency in Python scripting for data processing and evaluation (e.g., judge scripts)
- Familiarity with AI coding benchmarks such as SWE-bench and Terminal-bench
- Hands-on experience with Docker (writing Dockerfiles, building images, debugging containers)
- Must be eligible to work in one of: Remote, Bangladesh
How long hiring takes
Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.
Interview Prep
This listing calls for this tool directly. Prep for the technical screen:
What to Expect
Looking at Turing STEM listings we've tracked, contracts in this domain typically run about 8.5 weeks. Actual length varies by project, but this gives you a realistic baseline going in.
Based on 11 extracted Turing STEM listings.
Why this role
This AI Benchmark Engineer role covers building knowledge and research benchmark tasks used to test AI models on factual recall and research-style reasoning. The tasks need to be hard enough that a model can't answer correctly by pattern-matching training data alone.
Talent pool
We're light on STEM candidates
We've matched 65 people with a STEM background against 726 STEM listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Turing
Discover more opportunities on Turing that match your skills and interests.
View All Turing Jobs →Verified Reviews
Community Reviews
Share your experience with Turing
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
Do I need to be a software engineer to work for Turing?
No, not anymore. Turing built its name matching senior engineers with Silicon Valley companies, but it has since expanded into AGI infrastructure work and now hires non-engineering domain experts, technical writers, and researchers for post-training data annotation and RLHF. A strong analytical background and excellent English matter more than coding ability.
How does Turing's talent matching work?
Turing calls it the Intelligent Talent Cloud. You build a profile and go through vetting (automated tests, an AI-powered interview, practical skill assessments), and once vetted, Turing's algorithm surfaces your profile directly to partner companies like Fortune 500s and top AI labs. You don't browse listings or bid on work; matches come to you.
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
What does STEM work look like for an AI Benchmark Engineer — Knowledge / Research?
Tasks here are scoped to STEM, not generic labeling. As an AI Benchmark Engineer — Knowledge / Research, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to STEM) rather than following a one-size-fits-all rubric. If you don't have hands-on STEM background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding and Python are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
The listing doesn't state a rate. The $12–$30/hr shown here is our estimate from the role type and location (see /pay-methodology), so treat it as a rough guide and confirm the actual rate with the platform before committing time.
What happens when I click Apply on this listing?
You'll be taken to Turing's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Can I apply from outside Bangladesh, Brazil, Colombia and 8 other countries?
This specific role is open only to people based in Bangladesh, Brazil, Colombia, Egypt, India, Indonesia, Kenya, Nigeria, Turkey, Vietnam, and Ghana. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply, since they can differ by country.