Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and optimize model-serving stacks and inference microservices, develop GPU kernels, tune LLM inference and KV-cache strategies, operate distributed multi-node GPU serving systems, build serving-platform components and OpenAI-compatible endpoints, and improve performance, reliability, observability, and model support.
Requirements
- Experience with production model-serving frameworks
- Proficiency in C++, Python, and Rust
- Experience writing GPU kernels using CUDA or ROCm
- Understanding of LLM inference internals, attention mechanisms, KV-cache management, continuous batching, and quantization
- Experience with distributed multi-node, multi-GPU serving environments
- Experience deploying and managing services on Kubernetes, OpenShift, or similar platforms
- Experience with performance profiling, benchmarking, and debugging latency or throughput issues
Responsibilities
- Build, operate, and optimize production model-serving stacks
- Develop and maintain high-throughput model-inference microservices
- Write and optimize custom GPU kernels
- Optimize LLM prefill, decoding, attention, and continuous batching
- Implement and tune quantization, speculative decoding, tensor parallelism, pipeline parallelism, and MoE serving
- Design and implement KV-cache optimization strategies
- Build and operate fault-tolerant serving systems on orchestration platforms
- Implement distributed computing across multi-node, multi-GPU clusters
- Contribute to distributed serving architecture components
- Build and maintain OpenAI-compatible endpoints
- Conduct profiling and benchmarking to resolve latency and throughput regressions
- Build telemetry-driven observability platforms
- Support a broad range of production model classes