Company
OurFirm.ai
Annual salary
$180k – $220k/yr
Location
New York, NY
Listed
8d ago
- Experience:
- 4 to 8 years
- Workplace:
- Remote US-based with a strong preference for NYC. Eastern Time overlap preferred (4+ hours daily).
- Equity:
- 0.15%, 0.25% equity
Remote (listed in New York, NY).
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
About this Role
We're looking for a **Data Engineer **with 4+ YOE to own the caselaw and docket data layer at OurFirm.ai. You'll build production pipelines that ingest, normalize, and enrich data from PACER, NYSCEF, state court systems, and published opinions, turning raw legal data into the structured signal that feeds our AI strategy layer. Every downstream product surface depends on what you build. Legal data is uniquely messy, and prior fluency with it will save months of ramp.
What you'll do
- Take full ownership of the caselaw and docket data layer, build and maintain production pipelines ingesting from PACER, NYSCEF, state courts, and published opinions at scale across millions of court records
- Wrangle messy semi-structured inputs (PDFs, scanned filings, XML/HTML) into clean, queryable structures
- Design LLM-assisted extraction workflows using Claude, Gemini, or OpenAI to turn unstructured legal text into reliable structured signal
- Work closely with full-stack engineers to feed the judicial behavioral intelligence and AI strategy layers
- Run statistical analyses, optimize SQL queries, and use RAG and semantic search tooling (Pinecone, Voyage AI) to maximize data impact
- Operate and optimize the AWS data stack (S3, RDS) in line with SOC 2 subservice architecture guardrails
Frequently Asked Questions
How do I apply for the Data Engineer (US) role at OurFirm.ai? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this OurFirm.ai role pay? +
The listing gives $180k – $220k/yr.
Is this role remote? +
Yes, the listing is remote. Remote US-based with a strong preference for NYC. Eastern Time overlap preferred (4+ hours daily).
Interview Prep
Sample questions for a Data Engineer role, written in-house to help you prepare.
How do you design an ETL pipeline to be idempotent so reruns don't duplicate data?
I design loads around upserts keyed on a natural or surrogate key rather than plain inserts, and I make each pipeline stage stateless with respect to its inputs so rerunning it with the same input produces the same output. Partitioning by a processing date and overwriting the whole partition on rerun is another reliable pattern, especially for batch pipelines.
When would you choose a star schema over a snowflake schema for a data warehouse?
A star schema keeps dimension tables denormalized, which trades storage and some data redundancy for simpler, faster queries. That's the right call when query performance and ease of use for analysts matter more than storage efficiency. A snowflake schema is worth the extra join complexity when dimension tables are large, change frequently, and normalization meaningfully reduces storage and update anomalies.
How do you handle schema evolution in a pipeline that ingests data from a source you don't control?
I treat the ingestion layer as schema-tolerant by landing raw data first, before any strict schema is enforced, and applying schema validation at a later transformation stage. This way a new or renamed field upstream doesn't break ingestion outright. I also version the schema explicitly so downstream consumers can detect and handle changes rather than silently breaking.
What's your approach to choosing between a batch and a streaming architecture for a new data pipeline?
The deciding factor is how fresh the data needs to be for the downstream use case, not how technically interesting streaming is. If consumers only need daily or hourly updates, batch is simpler to build, debug, and operate. Streaming is worth the added complexity only when there's a real requirement for near-real-time data, like fraud detection or live dashboards.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →