Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will develop core frameworks for model training and evaluation that make models performant, resource-efficient, and reliable at scale. You will build and optimize the training stack, profile GPU-cluster workloads, remove bottlenecks, and ensure fault-tolerant, deterministic execution for long-running jobs.
Requirements
- Deep industry experience with AI/ML teams
- Software system design skills
- Python proficiency
- PyTorch or JAX fluency
- Demonstrated exceptional impact on real-world problems and systems
Responsibilities
- Build the training stack across model, layer, and kernel levels
- Optimize workloads through parallelism, quantization, and custom kernels
- Profile end-to-end training runs on large GPU clusters
- Eliminate bottlenecks and failures
- Monitor throughput, utilization, and uptime
- Ensure model architectures and training recipes scale efficiently
- Implement fault tolerance, checkpointing, and deterministic orchestration for large-scale jobs