Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own the lifecycle of bare metal GPU nodes from provisioning and delivery through operations break fix firmware management and decommissioning. You will automate node delivery manage GPU and server firmware improve fleet reliability participate in on call operations and define acceptance standards for new hardware and regions.
Requirements
- 3+ years in large scale bare metal or server fleet operations HPC or cloud infrastructure or 6+ years for Senior level
- Experience operating GPU servers at scale including driver CUDA and firmware management
- Strong Linux systems skills
- Experience with PXE IPMI Redfish OS imaging and automated provisioning
- Familiarity with DPU SmartNIC and bare metal networking
- Infrastructure automation skills with Ansible Terraform Python or Go
- Comfort with on call incident management and operational runbooks
- Multi region or large fleet operations experience is a plus
Responsibilities
- Own bare metal GPU node provisioning delivery operation break fix and decommissioning
- Build automated and repeatable node delivery pipelines
- Manage DPU SmartNIC and server firmware version baselines upgrades and validation
- Drive fleet reliability through incident response root cause analysis and health monitoring
- Participate in a 7 by 24 multi region on call rotation
- Build operational runbooks and tooling
- Partner with Storage Image and Network teams on provisioning and handoff
- Define bring up rack capacity and acceptance standards for GPU SKUs and regions
Benefits
- Remote work within San Jose or Austin