LLM - Applied AI Research Scientists
Turing • Remote
Education
Master
Type
Hourly
Listed
25d ago
Apply opens Turing in a new tab.
Apply Now → ⚡ Boost your chances - Optimize your resume with Rezi.aiAbout this role
From the Turing listing
About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We are seeking highly skilled and motivated Applied AI Research Scientists in Computer Science and Computer Engineering with an MS or Ph.D. in a relevant technical field to join our team at Turing. In this role, you will contribute to the design, validation, and execution of expert-level evaluation tasks that probe the limits of state-of-the-art AI systems. Your work will focus on creating headroom-level, rigorously verifiable questions across hardware, systems, and computing domains to assess and stress-test advanced multimodal and reasoning-capable AI models. This position requires deep domain expertise, strong analytical rigor, and the ability to translate complex technical concepts into precise, evaluable challenges that expose model limitations beyond surface-level reasoning. You will work closely with a collaborative, cross-functional team and are expected to be a reliable team player who is highly detail-oriented and committed to accuracy and quality.
Roles & Responsibilities
- Design headroom-level evaluation questions requiring advanced reasoning and graduate-level domain expertise in CS/CE.
- Ensure all tasks are objectively verifiable with clear, definitive ground-truth answers.Develop high-quality multimodal prompts, including accurate technical diagrams or visuals when appropriate.
- Identify and document model headroom, focusing on SOTA models like Gemini and ChatGPT, and conduct structured side-by-side evaluations.
- Document model failures and reasoning gaps, provide correct solutions, and maintain accurate records of prompts, answers, and evaluation results in shared tracking systems.
Requirements
- MS or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field.
- Strong expertise in two or more Computer Engineering or Computer Science domains, such as hardware, computer architecture, systems, VLSI design, embedded systems and IoT, operating systems, compilers, systems security, or AI/ML
- Proven experience with applied AI research, technical evaluation, or research-driven problem formulation in real-world or production-oriented settings
- Strong programming proficiency, with experience in Python for analysis, verification, and evaluation workflows
- Strong written communication skills and the ability to collaborate effectively as a detail-oriented team player
- Familiarity with modern AI model capabilities, limitations, and benchmarking practices is a strong plus
Perks of Freelancing With Turing
- Work in a fully remote environment
- Opportunity to work on cutting-edge AI projects with leading LLM companies
Offer Details
- Commitment Required: 8 hours per day, with 4 hours of mandatory overlap with PST
- Employment Type: Contractor assignment (no medical/paid leave).
- Contract Duration: 4 months (expected start date: next week).
Evaluation Process
- Round 1: Take home assessmentOffline assessment to completed and submitted for reveiw.
- Round 2: Delivery Interview (60 minutes)A combined technical and cultural discussion with the Delivery Team.
How long hiring takes
Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.
Interview Prep
This listing calls for this tool directly. Prep for the technical screen:
What to Expect
Looking at Turing STEM listings we've tracked, contracts in this domain typically run about 8.5 weeks. Actual length varies by project, but this gives you a realistic baseline going in.
Based on 11 extracted Turing STEM listings.
Why this role
This LLM Applied AI Research Scientist role is for research scientists to design evaluation methods and benchmarks that test frontier model capabilities in reasoning, coding, and STEM domains. The work sits closer to research design than task labeling, building the frameworks other reviewers use.
Talent pool
We're light on STEM candidates
We've matched 65 people with a STEM background against 726 STEM listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Turing
Discover more opportunities on Turing that match your skills and interests.
View All Turing Jobs →Verified Reviews
Community Reviews
Share your experience with Turing
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
How and when does Turing pay contractors?
Monthly, in USD, via Deel, Payoneer, or direct bank transfer. You're engaged as an independent contractor responsible for your own local taxes. Plan your cash flow around a monthly cycle if you're used to weekly payouts elsewhere.
Do I need to be a software engineer to work for Turing?
No, not anymore. Turing built its name matching senior engineers with Silicon Valley companies, but it has since expanded into AGI infrastructure work and now hires non-engineering domain experts, technical writers, and researchers for post-training data annotation and RLHF. A strong analytical background and excellent English matter more than coding ability.
Is academic-niche AI training just data labeling?
No, it's closer to academic research. Expect to write or verify complex proofs, solve advanced equations, or check the logic behind a model's step-by-step reasoning. The goal is teaching AI systems to reason deeply within your specific field.
Do I need a PhD for academic-niche AI training roles?
For the top pay tiers, a PhD or current enrollment is usually expected. But the domain assessment is what decides it: if you can solve the problems, the degree becomes secondary.
What does STEM work look like for a LLM - Applied AI Research Scientists?
Tasks here are scoped to STEM, not generic labeling. As a LLM - Applied AI Research Scientists, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to STEM) rather than following a one-size-fits-all rubric. If you don't have hands-on STEM background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding and Python are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
The listing doesn't state a rate. The $30–$70/hr shown here is our estimate from the role type and location (see /pay-methodology), so treat it as a rough guide and confirm the actual rate with the platform before committing time.
What happens when I click Apply on this listing?
You'll be taken to Turing's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Do I need a Master's to qualify?
This role lists a master's degree as a requirement. In practice, the domain assessment is the real gate. If you can pass it, the degree is usually secondary. However, some platforms verify credentials formally, so list your actual qualifications accurately on your profile.