Company
Inception
Annual salary
$200k – $300k/yr
Location
Palo Alto, CA
Listed
269d ago
- Experience:
- 2 to 5 years
- Workplace:
- 5 days in-office in Bay Area
- Visa:
- Visa sponsorship available
- Equity:
- Competitive equity
On-site in Palo Alto, CA. Visa sponsorship available.
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
About this Role
Inception creates the world’s fastest, most efficient AI models. Today’s autoregressive LLMs generate tokens sequentially, which makes them painfully slow and expensive. Inception’s diffusion-based LLMs (dLLMs) generate answers in parallel. They are up to 10X faster and more efficient, while delivering quality. Inception pioneered the application of diffusion to language, launching the world’s first commercially available dLLM, Mercury, in early 2025, and is currently deploying large-scale diffusion LLMs at Fortune 500 companies. Diffusion is the technology behind today’s image and video AI, and Inception making it the standard for LLMs as well.
Frequently Asked Questions
How do I apply for the Machine Learning Systems Engineer (US) role at Inception? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this Inception role pay? +
The listing gives $200k – $300k/yr.
Is this role remote? +
The listing gives the location as Palo Alto, CA. 5 days in-office in Bay Area. Visa sponsorship available.
Interview Prep
Sample questions for a Machine Learning Engineer role, written in-house to help you prepare.
How do you choose between a simpler interpretable model and a more complex model that performs marginally better?
The decision depends on how the model's output is used downstream. If a human needs to act on and justify individual predictions, like in lending or healthcare, interpretability often outweighs a small accuracy gain. If the model feeds an automated system where explainability isn't a hard requirement, I'll take the complexity if the performance gain is meaningful and validated, not just noise.
Explain how you would set up a proper train, validation, and test split for a time series problem.
Random splitting leaks future information into training for time series, so I split chronologically instead, training on the earliest period, validating on the next, and testing on the most recent. Any cross-validation also needs to respect time order, using something like rolling-origin validation, rather than standard k-fold, which would otherwise let the model see the future during validation.
What statistical test would you use to determine if an A/B test result is significant, and what assumptions does it rely on?
For a conversion-rate comparison, a two-proportion z-test or chi-squared test is standard, and it assumes independent observations and a large enough sample size for the normal approximation to hold. If the sample is small or the metric isn't binary, I'd switch to a more appropriate test, since applying the wrong test's assumptions is a common source of false confidence in A/B results.
How do you approach hyperparameter tuning efficiently when the model is expensive to train?
I use a coarse random search first to identify the promising region of the hyperparameter space, since grid search wastes compute on combinations unlikely to matter. From there, Bayesian optimization narrows in on the best configuration with far fewer expensive training runs than an exhaustive search would require, which matters a lot when each run is costly.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →