About the Role
You will deploy models to production and own the path from research checkpoints to serving infrastructure. You will optimize latency, throughput, and cost using inference techniques, build high performance systems for streaming workloads, and create tooling for safe and reliable model releases.
Requirements
- No formal certifications or degrees required
- Experience deploying and serving machine learning models in production
- Experience with latency sensitive or real time applications preferred
- Strong engineering skills in GPU programming and inference optimization
- Experience with CUDA Triton TensorRT vLLM SGLang or serving frameworks
- Ability to profile diagnose and eliminate serving stack bottlenecks
- Ability to build tooling to measure performance
Responsibilities
- Deploy models to production and connect research checkpoints to serving infrastructure
- Optimize inference latency throughput and cost
- Apply quantization distillation KV cache optimization batching and custom kernels
- Build and tune high performance serving systems for real time streaming workloads
- Create tooling and infrastructure for fast and reliable model deployment
- Profile diagnose and eliminate bottlenecks across the serving stack
Benefits
- Annual discretionary professional development stipend
- Annual discretionary social travel stipend
- Annual company offsite
- Monthly co-working stipend