Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build end-to-end inference capabilities, deploy and serve models across distributed clusters and heterogeneous hardware, and make deployments production-ready with monitoring, gateways, and endpoints. You will evaluate serving frameworks, optimize performance, implement autoscaling and KV-cache orchestration, and troubleshoot customer inference issues.
Requirements
- Strong general inference background across the request-to-token path
- Deep production Kubernetes experience
- Knowledge of TTFT, disaggregated inference, speculative decoding, and KV cache
- Familiarity with modern inference frameworks and serving engines
- Working knowledge of NVIDIA Dynamo in distributed serving architectures
- Experience setting up monitoring, gateways, and endpoints for production inference services
- Ability to build a product end to end and serve real traffic
Responsibilities
- Build inference capabilities on the unified control plane
- Deploy and serve models across distributed clusters and heterogeneous hardware
- Evaluate inference frameworks and serving engines
- Set up monitoring, gateways, and endpoints for production inference services
- Optimize inference performance and implement autoscaling and KV-cache orchestration
- Debug customer inference and performance issues