Skip to content
aitrainer.work - AI Training Jobs Platform
Engineering Full-Time
Minerva

Data Engineer (US)

Minerva • New York, NY

Company

Minerva

Annual salary

$200k – $250k/yr

Location

New York, NY

Listed

75d ago

Experience:
3+ years
Workplace:
In-person 5 days/week in our Williamsburg, NYC office.
Visa:
Visa sponsorship available
Equity:
Competitive Equity.

On-site in New York, NY. Visa sponsorship available.

Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.

Apply for a referral →

The recruiter emails you before anything happens and may suggest other jobs that suit you better.

About this Role

We are looking for an experienced Data Engineer to design, operate, and scale the pipelines, storage layers, and standardization systems that power our data product and AI platform. You should be comfortable owning analytics engineering and data modeling within a Snowflake-powered warehouse and have a track record of building fast, resilient data infrastructure. Bonus points if you have experience enabling AI and agentic access layers.

What you'll do

  • Own ingestion, data modeling and transformation on a people data product consisting of ~260m individuals, building and improving features, attributes, and pipelines.
  • Scale first-party integrations with sales and marketing systems of record, supporting a growing multi-terabyte data footprint.
  • Partner with talented Data Scientists to solve challenging data problems: productionizing complicated pipelines, scaling, preserving data quality.
  • Enable AI and agentic access layers, from natural language retrieval APIs to structured data orchestration.
  • Have the opportunity to work on tactical teams building out data products that directly unlock 7 figure contracts.
  • Work with data warehouses (Snowflake, Lakehouse on S3), data pipelines and orchestration tools (Dagster, dbt), and AWS infrastructure.

Frequently Asked Questions

How do I apply for the Data Engineer (US) role at Minerva? +

Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.

What does this Minerva role pay? +

The listing gives $200k – $250k/yr.

Is this role remote? +

The listing gives the location as New York, NY. In-person 5 days/week in our Williamsburg, NYC office. Visa sponsorship available.

Interview Prep

Sample questions for a Data Engineer role, written in-house to help you prepare.

How do you design an ETL pipeline to be idempotent so reruns don't duplicate data?

I design loads around upserts keyed on a natural or surrogate key rather than plain inserts, and I make each pipeline stage stateless with respect to its inputs so rerunning it with the same input produces the same output. Partitioning by a processing date and overwriting the whole partition on rerun is another reliable pattern, especially for batch pipelines.

When would you choose a star schema over a snowflake schema for a data warehouse?

A star schema keeps dimension tables denormalized, which trades storage and some data redundancy for simpler, faster queries. That's the right call when query performance and ease of use for analysts matter more than storage efficiency. A snowflake schema is worth the extra join complexity when dimension tables are large, change frequently, and normalization meaningfully reduces storage and update anomalies.

How do you handle schema evolution in a pipeline that ingests data from a source you don't control?

I treat the ingestion layer as schema-tolerant by landing raw data first, before any strict schema is enforced, and applying schema validation at a later transformation stage. This way a new or renamed field upstream doesn't break ingestion outright. I also version the schema explicitly so downstream consumers can detect and handle changes rather than silently breaking.

What's your approach to choosing between a batch and a streaming architecture for a new data pipeline?

The deciding factor is how fresh the data needs to be for the downstream use case, not how technically interesting streaming is. If consumers only need daily or hourly updates, batch is simpler to build, debug, and operate. Streaming is worth the added complexity only when there's a real requirement for near-real-time data, like fraud detection or live dashboards.

See all 10 questions for this role →

Related Roles

Minerva

Browse all startup roles

Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.

View all startup roles →