QA / Evaluation Lead
innodata • Hybrid - Washington D.C
Education
Bachelor
Type
Hourly
Pay Rate
$45–$50/hr
Listed
85d ago
About this role
From the innodata listing
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers. About the Program: Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust & Safety work. Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment. Over an intensive 20-week phase, you'll help stand up a data services storefront, a DataCard governance framework, synthetic data integration, and Databricks write-back capabilities. About the Role: As the QA/Evaluation Lead, you'll own quality and evaluation across the platform. You'll design the evaluation framework that measures whether our data services and outputs meet the bar, build repeatable test and validation processes, and give the team an objective read on readiness at each milestone. Partnering with the Delivery Owner and engineering leads, you'll turn quality from an afterthought into a measurable, demonstrable strength. It's a role for someone who thinks rigorously about evaluation and takes pride in evidence-backed quality. Key Responsibilities: Design and own the inter-annotator agreement (IAA) methodology for the Phase 1 demonstration corpus — metric selection (Cohen's kappa, Fleiss, Krippendorff's alpha), sampling design, adjudication workflow, and agreement thresholds Define evaluation framework architecture: test and evaluation plans, IAA targets, drift detection gates, and model performance metrics per SOW Section 2.9 Configure and operate sampling-based quality control across the self-service and white-glove annotation paths during Phase D corpus production Design and implement confidence-threshold escalation routing from automated annotation to senior-annotator adjudication Validate quality scoring and IAA computation within the Innodata data layer Support AI Solutions Engineer on evaluation design for SAM 2 and Frontier model API validation — define what 'good enough' looks like quantitatively Produce evaluation framework documentation for the Phase 1 NPP closeout package, including per-DataCard documentation with the SA Must-Have Qualifications: Bachelor's degree in Statistics, Data Science, Computer Science, or related quantitative field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis. 6+ years total professional experience, 4+ years in data quality, evaluation methodology, or QA on AI/ML programs IAA methodology expertise — Cohen's kappa, Fleiss' kappa, Krippendorff's alpha: hands-on, not theoretical Evaluation framework design for AI/ML training data programs QC process design: sampling methodology, escalation workflows, adjudication protocols Python for QC tooling, metric computation, and statistical analysis Active Secret clearance with TS/SCI eligibility Nice-to-Have Qualifications: Prior DoD or IC data quality program experience CVAT or equivalent annotation platform QC workflow configuration Drift detection and model monitoring methodology Experience with FMV / video annotation quality standards The expected hourly salary range for this position is $45 to $50 p/hour, based on experience, skills, and qualifications. Note to Candidates: This role is not a project manager with QC responsibilities — it is a methodology expert who owns the intellectual framework behind data quality on a federal AI program. The right candidate can walk into a meeting with Government evaluators and explain exactly
Requirements
- Must be eligible to work in Hybrid - Washington D.C
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
- Bachelor's degree or equivalent professional experience
- Demonstrated expertise in Software Engineering
Interview Prep
This listing calls for these tools directly. Prep for the technical screen:
Why this role
This QA / Evaluation Lead position pays $45–$50/hr. The bar for it is solid Software Engineering knowledge, and the AI training workflow itself gets taught on the job.
Skills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from innodata
Discover more opportunities on innodata that match your skills and interests.
View All innodata Jobs →Verified Reviews
Community Reviews
Share your experience with innodata
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
What does asynchronous AI training work mean in practice?
No set hours, no check-ins, no meetings. You log in when you want, pick up an available task, complete it, and submit; nobody is waiting on you in real time. That's different from remote employment, where you're expected online during business hours. The tradeoff: you're competing with others for available tasks, so an empty queue means there's simply nothing to do until more work is released.
Do task-based AI training roles require an interview?
Most skip a live interview entirely and gate access through an automated assessment or qualification task instead. Where we've confirmed the specific requirement for this listing, it's called out above in What We Know About This Role.
What does Software Engineering work look like for a QA / Evaluation Lead?
Tasks here are scoped to Software Engineering, not generic labeling. As a QA / Evaluation Lead, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to Software Engineering) rather than following a one-size-fits-all rubric. If you don't have hands-on Software Engineering background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding, Data Annotation, Python, and TypeScript are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
This listing is posted at $45–$50/hr, an hourly rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.
Can I apply from outside Hybrid - Washington D.C?
This specific role is open only to people based in Hybrid - Washington D.C. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply.