Skip to content
aitrainer.work - AI Training Jobs Platform
Engineering Full-Time
Nuance Labs

Member of Technical Staff - Model Optimization and Inference (Experienced) (US)

Nuance Labs • Seattle, WA

Company

Nuance Labs

Annual salary

$250k – $350k/yr

Location

Seattle, WA

Listed

225d ago

Experience:
2+ years
Workplace:
This is a full-time, in-office role at our Seattle, WA location (5 days/week). Relocation assistance is available.
Visa:
Visa sponsorship available

On-site in Seattle, WA. Visa sponsorship available.

Send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role.

Apply for a referral →

The recruiter emails you before anything happens and may suggest other jobs that suit you better.

What you'll do

  • Own end-to-end inference optimization across our model stack, LLMs, audio models, and diffusion-based components
  • Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
  • Evaluate, deploy, and extend inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) for our specific workloads
  • Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
  • Build internal tooling that makes optimization work faster and more rigorous, profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
  • Accelerate diffusion model inference, consistency models, step distillation, caching strategies, and custom kernel optimizations
  • Apply and develop quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
  • Work closely with research and infrastructure to ensure new models ship with optimized serving from day one

What you need

  • We are seeking an experienced ML Infrastructure/Systems Engineer with at least 2 years of full-time experience in building and maintaining production-level ML systems. You should be comfortable designing scalable infrastructure from scratch, making informed design decisions by comparing various technologies, and have a track record of optimizing systems for latency, throughput, and cost. We're looking for someone with a broad understanding of the ML infra space, including inference infrastructure, real-time video streaming, and data engineering, who can own complex projects and debug distributed systems. Bonus points if you have experience with video or audio models and low-level optimization techniques like CUDA kernels.

Frequently Asked Questions

How do I apply for the Member of Technical Staff - Model Optimization and Inference (Experienced) (US) role at Nuance Labs? +

Use the Apply for a referral button on this page to send us your LinkedIn and CV. If your experience fits, we'll introduce you to the recruiter filling this role. They'll email you to check you're interested, then put you forward for this job or others that suit you better. It's free.

What does this Nuance Labs role pay? +

The listing gives $250k – $350k/yr.

Is this role remote? +

The listing gives the location as Seattle, WA. This is a full-time, in-office role at our Seattle, WA location (5 days/week). Relocation assistance is available. Visa sponsorship available.

Related Roles

Nuance Labs

Browse all startup roles

Salaried roles at startups, AI companies and established businesses, all filled by referral. One application covers every role.

View all startup roles →