Skip to content
aitrainer.work - AI Training Jobs Platform
Engineering Full-Time
Plastic Labs

Machine Learning Engineer - Evals (US)

Plastic Labs • New York, NY

Company

Plastic Labs

Annual salary

$220k – $300k/yr

Location

New York, NY

Listed

39d ago

Experience:
3+ years
Workplace:
5 days in-office in New York, NY
Equity:
Competitive equity

On-site in New York, NY.

Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.

Apply for a referral →

The recruiter emails you before anything happens and may suggest other jobs that suit you better.

About this Role

You'll own evaluation of Honcho end to end, the instrument that tells a research-driven team whether their identity representations are getting better. Plastic Labs builds the memory and identity layer for the agentic world, and the product works. It's a complex multi-agent harness with real black-box behavior, so the job is defining what "better" means for representations that change over time, without ground truth, and building the machinery that measures it.

If you love constructing measurements from nothing, reading traces to find what's broken, and shipping the fix yourself, you'll feel at home here.

What You'll Do

Design the evals. Define the scores and decide what "better" means for an entity representation that changes over time, and keep that definition current as the product and methods move.

Build the pipelines and harnesses. Data in, labels, versions, reruns, judges: the machinery that lets the team ask a new question this week and get an answer this week.

Run them and harvest insights. Read the traces and results, find what's broken, not what's easy to measure, and propose fixes that produce higher-fidelity representations.

Build simulation agents at scale. The ceiling on iteration speed is how many entities can be modeled and measured at once. You raise it.

Own the loop end to end. The question, the pipeline, the rerun, the writeup. Nothing gets scoped and handed off, and nothing waits on someone else's sprint.

Frequently Asked Questions

How do I apply for the Machine Learning Engineer - Evals (US) role at Plastic Labs? +

Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.

What does this Plastic Labs role pay? +

The listing gives $220k – $300k/yr.

Is this role remote? +

The listing gives the location as New York, NY. 5 days in-office in New York, NY.

Interview Prep

Sample questions for a Machine Learning Engineer role, written in-house to help you prepare.

How do you choose between a simpler interpretable model and a more complex model that performs marginally better?

The decision depends on how the model's output is used downstream. If a human needs to act on and justify individual predictions, like in lending or healthcare, interpretability often outweighs a small accuracy gain. If the model feeds an automated system where explainability isn't a hard requirement, I'll take the complexity if the performance gain is meaningful and validated, not just noise.

Explain how you would set up a proper train, validation, and test split for a time series problem.

Random splitting leaks future information into training for time series, so I split chronologically instead, training on the earliest period, validating on the next, and testing on the most recent. Any cross-validation also needs to respect time order, using something like rolling-origin validation, rather than standard k-fold, which would otherwise let the model see the future during validation.

What statistical test would you use to determine if an A/B test result is significant, and what assumptions does it rely on?

For a conversion-rate comparison, a two-proportion z-test or chi-squared test is standard, and it assumes independent observations and a large enough sample size for the normal approximation to hold. If the sample is small or the metric isn't binary, I'd switch to a more appropriate test, since applying the wrong test's assumptions is a common source of false confidence in A/B results.

How do you approach hyperparameter tuning efficiently when the model is expensive to train?

I use a coarse random search first to identify the promising region of the hyperparameter space, since grid search wastes compute on combinations unlikely to matter. From there, Bayesian optimization narrows in on the best configuration with far fewer expensive training runs than an exhaustive search would require, which matters a lot when each run is costly.

See all 10 questions for this role →

Related Roles

Plastic Labs

Browse all startup roles

Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.

View all startup roles →