LLM Trainer – Terminal-Bench (Python & Linux Systems)
Turing • Bangladesh, Egypt, India, Kenya, Mexico, Nigeria, Pakistan, Turkey, Ghana
Education
Not stated
Type
Hourly
Listed
24d ago
Apply opens Turing in a new tab.
Apply Now → ⚡ Boost your chances - Optimize your resume with Rezi.aiAbout this role
From the Turing listing
About Turing
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and second, by applying that expertise to help enterprises transform AI from proof of concept into proprietary intelligence with systems that perform reliably, deliver measurable impact, and drive lasting results on the P&L
Role Overview
We are seeking skilled Terminal-Bench Task 2.0 (Harbor) Authors to design, develop, and validate high-quality benchmark tasks for evaluating large language models (LLMs) in simulated Linux terminal environments.In this role, you will create challenging, deterministic, and reproducible tasks that rigorously test AI capabilities across software engineering, systems, data science, and mathematical domains. Your work will directly contribute to evaluating and stress-testing frontier AI models under real-world terminal constraints.
What does a typical day look like?
- Author original Terminal-Bench 2.0 (Harbor) tasks with precise, unambiguous instructions
- Design realistic Linux terminal workflows involving filesystems, processes, networking, and containers
- Implement golden solutions and pytest-based evaluation scripts with deterministic outcomes
- Build and maintain Dockerized environments with pinned dependencies for reproducibility
- Anticipate edge cases and failure modes to prevent reward hacking
- Validate tasks by running them against frontier LLMs and iterating to achieve target pass/fail rates
- Deliver approximately 5 fully validated benchmark tasks per week
Key Responsibilities
- Design & Author Tasks: Create unique, high-quality Terminal-Bench 2.0 tasks with clear goals, environment setup, and expected outputs
- Write Deterministic Tests: Develop robust pytest-based test suites with no hidden test cases
- Build Docker Environments: Configure reproducible containerized setups with locked dependencies
- Model Validation: Execute tasks against AI models and refine difficulty and clarity
- Quality Assurance: Ensure strong alignment between instructions, tests, and evaluation logic to maintain benchmark integrity
Required Skills & Qualifications
Technical Skils
- Python: Strong proficiency with clean, testable code (3+ years experience)
- Bash / Shell Scripting: Confident with Unix command-line tools and workflows (2+ years experience)
- Linux Systems: Familiarity with filesystems, permissions, processes, and basic networking (1+ year experience)
- Docker: Experience building, configuring, and debugging containerized environments
Domain Expertise (one or more required)
- Software Engineering, System Administration, or Debugging
- Data Science, Machine Learning, or Model Training
- Mathematics, Algorithm Design, or Scientific Computing
- Frontend or Backend Development
- Data Preprocessing and Analysis
Core Competencies
- Strong analytical and problem-solving ability
- Extreme attention to detail in technical writing
- Ability to anticipate corner cases and unintended model behaviors
- Comfort working with deterministic evaluation and strict correctness criteria
Expected Output
- ~5 fully validated Terminal-Bench tasks per week, including instructions, Docker setup, golden solutions, and test suites
Perks of Freelancing With Turing:
- Work in a fully remote environment.
- Opportunity to work on cutting-edge AI projects with leading LLM companies.
Offer Details
- Commitments Required: 8 hours per day with overlap of 4 hours with PST.
- Employment type : Contractor assignment (no medical/paid leave)
- Duration of contract : 1 month; [expected start date is next week]
- Location : India, Pakistan, Nigeria, Kenya, Egypt, Ghana, Bangladesh, Turkey, Mexico
Requirements
- Must be eligible to work in one of: Bangladesh, Egypt, India, etc.
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
How long hiring takes
Across the AI training platforms we refer candidates to, the median gap between referral and hire is about 30 days. It varies by platform and role, so treat it as a rough guide for this one.
Interview Prep
This listing calls for this tool directly. Prep for the technical screen:
Why this role
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers in two ways: first, by accelerating frontier research with high-quality data, advanced training pipelines, plus top AI researchers who speci
Talent pool
We're light on Data Science candidates
We've matched 9 people with a Data Science background against 293 Data Science listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from Turing
Discover more opportunities on Turing that match your skills and interests.
View All Turing Jobs →Verified Reviews
Community Reviews
Share your experience with Turing
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
How does Turing's talent matching work?
Turing calls it the Intelligent Talent Cloud. You build a profile and go through vetting (automated tests, an AI-powered interview, practical skill assessments), and once vetted, Turing's algorithm surfaces your profile directly to partner companies like Fortune 500s and top AI labs. You don't browse listings or bid on work; matches come to you.
How and when does Turing pay contractors?
Monthly, in USD, via Deel, Payoneer, or direct bank transfer. You're engaged as an independent contractor responsible for your own local taxes. Plan your cash flow around a monthly cycle if you're used to weekly payouts elsewhere.
What does task-based AI training work look like?
Practical, hands-on data work: recording short videos, categorizing images, rating text responses, or analyzing data. Tasks are designed to be short and distinct, typically 5 to 60 minutes each.
What does Data Science work look like for a LLM Trainer – Terminal-Bench (Python & Linux Systems)?
Tasks here are scoped to Data Science, not generic labeling. As a LLM Trainer – Terminal-Bench (Python & Linux Systems), expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Data Science) rather than following a one-size-fits-all rubric. If you don't have hands-on Data Science background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding and Python are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
The listing doesn't state a rate. The $12–$30/hr shown here is our estimate from the role type and location (see /pay-methodology), so treat it as a rough guide and confirm the actual rate with the platform before committing time.
What happens when I click Apply on this listing?
You'll be taken to Turing's external site to complete your application there. This listing links through a referral, but the process is identical to applying directly; the link just routes you correctly. Create an account on their site and follow their onboarding steps.
Can I apply from outside Bangladesh, Egypt, India and 6 other countries?
This specific role is open only to people based in Bangladesh, Egypt, India, Kenya, Mexico, Nigeria, Pakistan, Turkey, and Ghana. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply, since they can differ by country.