Company
Amboras
Annual salary
$150k – $250k/yr
Location
San Francisco, CA
Listed
11d ago
- Experience:
- 2 to 4 years
- Workplace:
- Onsite in SF
- Equity:
- 0.1% - 0.5%
On-site in San Francisco, CA.
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
About this Role
You'll own the AI layer of the product, the models, the pipelines, and the evals that decide whether what we ship works for merchants. This is not a research seat. You'll be in production code, shipping to real customers, and you'll be the person the rest of the team asks when something model-shaped breaks.
You're one of the first engineers here, so the architecture decisions you make in month one are the ones we live with for years.
What you'll do
- Own the end-to-end AI systems: prompting, retrieval, tool use, agent orchestration, fine-tuning where it earns its keep
- Build the eval harness before the feature, we don't ship model changes on vibes
- Get latency and cost to numbers that work at merchant scale, not demo scale
- Take an ambiguous product problem and decide whether it's a model problem, a data problem, or a UX problem, then fix it
- Ship to production weekly and watch what happens in the wild
Frequently Asked Questions
How do I apply for the AI Engineer (US) role at Amboras? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this Amboras role pay? +
The listing gives $150k – $250k/yr.
Is this role remote? +
The listing gives the location as San Francisco, CA. Onsite in SF.
Interview Prep
Sample questions for a AI Engineer role, written in-house to help you prepare.
How do you decide which evaluation metric to optimize for a classification model when the classes are imbalanced?
Accuracy is misleading on imbalanced data because a model can score well by always predicting the majority class. I look at precision, recall, and F1 for the minority class specifically, and pick the metric that matches the real cost of false positives versus false negatives in the deployment context, since those costs are rarely symmetric in practice.
What is your process for debugging a deep learning model that trains fine but performs poorly at inference time?
I start by checking for a train/inference mismatch: different preprocessing, batch normalization behaving differently in eval mode, or data leakage during training that inflated validation scores. Comparing the exact input pipeline used at training time against the one used at inference usually surfaces the discrepancy before I need to touch the model architecture itself.
How do you approach feature selection when working with high-dimensional tabular data?
I start with correlation and mutual information analysis to drop obviously redundant features, then use a tree-based model's feature importance as a first pass filter. From there, recursive feature elimination with cross-validation tells me where the marginal value of additional features drops off, which keeps the final feature set both smaller and more robust to overfitting.
Explain how you would detect and address data drift in a production machine learning pipeline.
I monitor the statistical distribution of input features and model outputs over time, using something like population stability index or KL divergence against a reference window. When drift crosses a threshold, I retrain on recent data rather than the full historical set, and I keep an alert in place so drift is caught before it shows up as a quality regression.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →