Data & Annotation Engineer
innodata • Hybrid - Washington D.C
Education
Not stated
Type
Hourly
Pay Rate
$55–$60/hr
Listed
86d ago
About this role
From the innodata listing
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers. About the Program: Innodata's Federal Practice builds the trusted data layer for critical infrastructure Trust & Safety work. Partnering with a leading systems integrator, we're delivering a modern, governed data services platform in a secure federal (IL4) environment. Over an intensive 20-week phase, you'll help stand up a data services storefront, a DataCard governance framework, synthetic data integration, and Databricks write-back capabilities. About the Role: As the Data/Annotation Engineer, you'll be hands-on with the data itself. You'll administer the annotation toolchain, manage annotation workflows across the corpus, and produce the per-dataset documentation that feeds our governance framework. You'll work with the AI Solutions Engineer to ensure the data going into our models is accurate, well-labeled, and fully traceable. This role is for someone detail-obsessed who understands that great AI starts with disciplined, well-governed data. Key Responsibilities: Receive, validate, ingest, and ontology-map the ODIN mission-aligned corpus from AFS delivery Produce the ODIN load report: corpus description, ontology mapping, readiness state Configure CVAT annotation pipeline against the Phase 1 starter kit rule pack Operate both self-service and lightweight white-glove annotation paths during Phase D corpus production Produce 50–100 label demonstration corpus across synthetic and mission-aligned content Support QA/Evaluation Lead on QC execution and corpus annotation dry-runs Associate DataCard provenance records with annotated and synthetic outputs in coordination with the Solution Architect Must-Have Qualifications: Bachelor's degree in Data Science, Computer Science, or related field preferred. Equivalent experience may substitute for degree on a 2-for-1 basis. 5+ years total professional experience, 3+ years in data engineering or annotation operations CVAT — deployment and day-to-day operation required; this is not a nice-to-have Annotated dataset ingest pipelines: schema mapping, format validation, ontology alignment Full-motion video (FMV) annotation concepts and tooling Python scripting for data wrangling, validation, and format conversion Active Secret clearance with TS/SCI eligibility Nice-to-Have Qualifications: Bachelor's degree in Computer Science, Machine Learning, Data Science, or related field required; Master's degree preferred. Equivalent experience may substitute for degree on a 2-for-1 basis CVAT annotation platform — AI feature configuration and operation DoD or IC data program experience: CUI, distribution statements, federal data governance Evaluation design for AI/ML training data: IAA methodology, drift detection, model performance measurement Video understanding or FMV annotation experience DataCard or ML data provenance framework familiarity The expected hourly salary range for this position is $55 to $60 p/hour, based on experience, skills, and qualifications. Note to Candidates: Phase D corpus production (Weeks 17–19) is the core demonstration deliverable. Candidates must be genuinely comfortable operating CVAT at production quality against a mission dataset under a milestone deadline Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application proce
Requirements
- Must be eligible to work in Hybrid - Washington D.C
- Fluent proficiency in English (Written & Verbal)
- Reliable high-speed internet connection
- Bachelor's degree or equivalent professional experience
- Demonstrated expertise in STEM
Interview Prep
Sample questions for a Data Engineer role, written in-house to help you prepare.
How do you design an ETL pipeline to be idempotent so reruns don't duplicate data?
I design loads around upserts keyed on a natural or surrogate key rather than plain inserts, and I make each pipeline stage stateless with respect to its inputs so rerunning it with the same input produces the same output. Partitioning by a processing date and overwriting the whole partition on rerun is another reliable pattern, especially for batch pipelines.
When would you choose a star schema over a snowflake schema for a data warehouse?
A star schema keeps dimension tables denormalized, which trades storage and some data redundancy for simpler, faster queries. That's the right call when query performance and ease of use for analysts matter more than storage efficiency. A snowflake schema is worth the extra join complexity when dimension tables are large, change frequently, and normalization meaningfully reduces storage and update anomalies.
How do you handle schema evolution in a pipeline that ingests data from a source you don't control?
I treat the ingestion layer as schema-tolerant by landing raw data first, before any strict schema is enforced, and applying schema validation at a later transformation stage. This way a new or renamed field upstream doesn't break ingestion outright. I also version the schema explicitly so downstream consumers can detect and handle changes rather than silently breaking.
What's your approach to choosing between a batch and a streaming architecture for a new data pipeline?
The deciding factor is how fresh the data needs to be for the downstream use case, not how technically interesting streaming is. If consumers only need daily or hourly updates, batch is simpler to build, debug, and operate. Streaming is worth the added complexity only when there's a real requirement for near-real-time data, like fraud detection or live dashboards.
This listing calls for these tools directly. Prep for the technical screen:
Why this role
This Data & Annotation Engineer position pays $55–$60/hr. The bar for it is solid STEM knowledge, and the AI training workflow itself gets taught on the job.
Talent pool
We're light on STEM candidates
We've matched 65 people with a STEM background against 726 STEM listings we've tracked, so most go out without one. Set up a profile and we'll consider you for a role like this one.
Set up your profileSkills and categories
Explore other opportunities in related specializations:
Related jobs
Browse All Jobs from innodata
Discover more opportunities on innodata that match your skills and interests.
View All innodata Jobs →Verified Reviews
Community Reviews
Share your experience with innodata
Help other candidates make better decisions by leaving a review.
Sign in to leave a reviewLeave your review
Common questions
Is academic-niche AI training just data labeling?
No, it's closer to academic research. Expect to write or verify complex proofs, solve advanced equations, or check the logic behind a model's step-by-step reasoning. The goal is teaching AI systems to reason deeply within your specific field.
Do I need a PhD for academic-niche AI training roles?
For the top pay tiers, a PhD or current enrollment is usually expected. But the domain assessment is what decides it: if you can solve the problems, the degree becomes secondary.
What does STEM work look like for a Data & Annotation Engineer?
Tasks here are scoped to STEM, not generic labeling. As a Data & Annotation Engineer, expect to draw on real domain judgment (evaluating outputs, correcting errors, or providing expert reasoning specific to STEM) rather than following a one-size-fits-all rubric. If you don't have hands-on STEM background, this is likely not the right listing to start with.
What specific skills does this listing call for?
Coding, Python, and TypeScript are named directly in the listing. If you don't have hands-on experience with these, expect the screening process to test for them directly rather than accepting adjacent experience as a substitute.
How much does this specific role pay?
This listing is posted at $55–$60/hr, an hourly rate. The range reflects experience level and negotiated terms, not a placeholder, so where you land in it depends on your background and the assessment. Pay can change between when we last checked the listing and when you apply, so confirm the current number on the platform's own application page before committing time.
Can I apply from outside Hybrid - Washington D.C?
This specific role is open only to people based in Hybrid - Washington D.C. If you are somewhere else, applying is unlikely to lead to an offer even if you pass the assessment, because the restriction is usually about where the work can legally be contracted rather than your skills. Read the full description for any tax-residency or right-to-work caveats before you apply.