Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build, operate, and optimize a production AI inference platform. You will onboard and configure new models, profile performance across GPUs and distributed systems, improve reliability and cost efficiency, manage model serving lifecycles, and operate multi-tenant GPU scheduling and isolation.
Requirements
- Hands-on experience serving LLMs on NVIDIA GPUs.
- Deep expertise in TensorRT-LLM, vLLM, SGLang, or another modern inference runtime.
- Knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
- Understanding of KV cache architecture.
- Experience measuring model quality equivalence using evaluation harnesses, benchmarks, and regression detection.
- Experience designing and tuning distributed inference systems.
- Experience with multi-tenant GPU scheduling, workload isolation, and QoS.
- Proficiency with Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
- Strong Linux and systems performance fundamentals.
- Production experience with model serving infrastructure, CI/CD, automated testing, observability, and production readiness.
Responsibilities
- Optimize LLM inference performance across NVIDIA GPU architectures and inference runtimes.
- Build performance profiling and observability across GPU kernels and multi-node inference systems.
- Design and optimize KV cache and distributed inference architectures.
- Own day-zero model onboarding and select serving configurations.
- Maintain validated performance profiles and quality regression testing.
- Measure and monitor quality equivalence across serving configurations.
- Manage model versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
- Automate model deployment, distribution, production readiness, observability, and reliable operation.
- Optimize model placement, scaling, and resource allocation across the inference fleet.
- Design and operate multi-tenant GPU scheduling and workload isolation.
Benefits
- Annual discretionary bonus
- Group medical, pharmacy, dental, and vision insurance
- 401k with discretionary employer match
- Short-term and long-term disability insurance
- Life and AD&D insurance
- Health savings accounts
- Flexible spending accounts