Company
Replay
Annual salary
$250k – $330k/yr
Location
New York, NY
Listed
11d ago
- Experience:
- 3 to 6 years
- Workplace:
- 5 days in-office in Dumbo, Brooklyn (non-negotiable)
- Equity:
- Competitive equity
On-site in New York, NY.
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
About this Role
We are looking for a **Machine Learning Engineer with 3+ years of experience **to join a ~5-person engineering team at Replay and improve the de-identification pipeline that turns sensitive enterprise data into AI-training-ready datasets. You'll be an applied engineer first, studying model failures, building datasets, running experiments, and shipping improvements into a live production pipeline. Replay is the primary source of real, proprietary enterprise data for the world's leading frontier AI labs, having scaled from $0 to a multi-eight-figure run rate in months. This is a high-ownership role where you'll work directly with a wealth of real-world data across multiple modalities, tackling hard problems in NER, entity resolution, and precision-recall tradeoffs where both over-redaction and missed sensitive data have real consequences.
What will you be doing?
Own and improve the de-identification model, retrain, evaluate where it fails to generalize, and ship measurable improvements to precision and recall in production
Design and scale labeling strategies, including active learning loops and LLM-assisted review, to use Replay's massive wealth of proprietary data
Expand de-identification capabilities into new modalities (audio, video, tabular data) that haven't been tackled yet
Combine deterministic rules, classical ML, fine-tuning, and LLM-based approaches to solve NER and entity resolution problems, choosing the right tool based on evidence, not hype
Build evaluation benchmarks and experiment infrastructure that make the path from error discovery to trustworthy production improvement faster and more repeatable
Frequently Asked Questions
How do I apply for the Machine Learning Engineer (US) role at Replay? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this Replay role pay? +
The listing gives $250k – $330k/yr.
Is this role remote? +
The listing gives the location as New York, NY. 5 days in-office in Dumbo, Brooklyn (non-negotiable).
Interview Prep
Sample questions for a Machine Learning Engineer role, written in-house to help you prepare.
How do you choose between a simpler interpretable model and a more complex model that performs marginally better?
The decision depends on how the model's output is used downstream. If a human needs to act on and justify individual predictions, like in lending or healthcare, interpretability often outweighs a small accuracy gain. If the model feeds an automated system where explainability isn't a hard requirement, I'll take the complexity if the performance gain is meaningful and validated, not just noise.
Explain how you would set up a proper train, validation, and test split for a time series problem.
Random splitting leaks future information into training for time series, so I split chronologically instead, training on the earliest period, validating on the next, and testing on the most recent. Any cross-validation also needs to respect time order, using something like rolling-origin validation, rather than standard k-fold, which would otherwise let the model see the future during validation.
What statistical test would you use to determine if an A/B test result is significant, and what assumptions does it rely on?
For a conversion-rate comparison, a two-proportion z-test or chi-squared test is standard, and it assumes independent observations and a large enough sample size for the normal approximation to hold. If the sample is small or the metric isn't binary, I'd switch to a more appropriate test, since applying the wrong test's assumptions is a common source of false confidence in A/B results.
How do you approach hyperparameter tuning efficiently when the model is expensive to train?
I use a coarse random search first to identify the promising region of the hyperparameter space, since grid search wastes compute on combinations unlikely to matter. From there, Bayesian optimization narrows in on the best configuration with far fewer expensive training runs than an exhaustive search would require, which matters a lot when each run is costly.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →