Company
Dedalus Labs
Annual salary
$180k – $300k/yr
Location
San Francisco, CA
Listed
15d ago
- Experience:
- 0 to 12 years
- Workplace:
- Full-time, 5 days per week on-site in San Francisco. Relocation support available (flights + 6 months rent covered).
- Equity:
- 0.5%, 1%
On-site in San Francisco, CA.
Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.
Apply for a referral →The recruiter emails you before anything happens and may suggest other jobs that suit you better.
What you'll do
- Build core compute and runtime primitives for long-running, stateful AI agents, including virtualization, sandboxing, and isolation systems for secure multi-tenant workloads
- Take ownership of entire subsystems end-to-end (storage, scheduling, networking, virtualization) and ship them to production at an 8-person team where your work matters immediately
- Work closely with the CTO and engineering team to design and implement persistent state, snapshotting, and recovery mechanisms for always-on agent workloads
- Reduce sandbox startup latency to sub-50ms and optimize performance across every syscall and scheduling decision in the critical path
- Design and build distributed storage layers, schedulers, and networking infrastructure that will scale across data centers globally
- Profile, debug, and benchmark systems under real-world scale, latency, correctness, and reliability constraints, measuring before optimizing, then optimizing relentlessly
- Contribute to the longer-term roadmap including on-prem GPU clusters and model training infrastructure as Dedalus expands beyond cloud compute
What you need
- We need a systems engineer with depth in OS/kernel internals, virtualization, file systems, or distributed infrastructure who has owned technically ambitious projects end-to-end, in production, open source, or research. You should be comfortable working at the lowest layers of the stack (kernel, hypervisor, storage, networking) and have a verifiable body of work that demonstrates genuine depth. Bonus points if you have experience with GPU infrastructure, AI/ML systems, or distributed storage/consensus.
Frequently Asked Questions
How do I apply for the Systems Engineer (US) role at Dedalus Labs? +
Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.
What does this Dedalus Labs role pay? +
The listing gives $180k – $300k/yr.
Is this role remote? +
The listing gives the location as San Francisco, CA. Full-time, 5 days per week on-site in San Francisco. Relocation support available (flights + 6 months rent covered).
Interview Prep
Sample questions for a Systems Engineer role, written in-house to help you prepare.
How do you decide between round-robin and least-connections load balancing for a given service?
I use round-robin when requests are roughly uniform in processing time, and least-connections when request duration varies significantly, since round-robin can overload a server handling several slow requests while it keeps receiving new ones at the same rate as faster servers.
What's your process for diagnosing intermittent network latency that only occurs under peak load?
I check for saturation on specific links or interfaces during the peak window using monitoring data, rather than testing during off-peak hours when the issue won't reproduce. I also look at whether the latency correlates with a specific service or is spread across the network, which points to either an application issue or a capacity issue.
How do you approach capacity planning for network infrastructure that needs to support future growth?
I project growth based on historical trend data combined with known upcoming initiatives, and I build in headroom above the projection since infrastructure upgrades often have long lead times. I avoid provisioning purely for current load, since that leaves no margin for unexpected spikes or faster-than-projected growth.
What steps do you take to ensure a load balancer configuration change doesn't cause an outage during deployment?
I test the configuration change in a staging environment that mirrors production traffic patterns first, and I roll out changes gradually, shifting a small percentage of traffic before a full cutover. I also keep the previous configuration ready to roll back to quickly if something unexpected happens.
Related Roles
Browse all startup roles
Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.
View all startup roles →