Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will integrate and operate distributed and parallel file systems within a large GPU cloud. You will manage provisioning multi tenant isolation quotas and lifecycle operations tune storage performance architect multi region data placement and maintain golden images drivers CUDA and container registries.
Requirements
- 3+ years in storage engineering or platform infrastructure or 6+ years for Senior level
- Hands on distributed and parallel file system experience
- Understanding of distributed file system internals data and metadata separation replication consistency and POSIX and object semantics
- Production experience with Ceph Lustre GPFS Spectrum Scale BeeGFS JuiceFS or MinIO
- Performance tuning for high throughput parallel I/O
- Familiarity with NVMe RDMA RoCE storage networking and caching
- Strong Linux systems and automation skills with Python or Go and CI and CD
- HPC AI storage or multi region storage experience is a plus
Responsibilities
- Design and integrate distributed and parallel file systems for AI training and inference
- Own storage provisioning mounting tenant isolation quotas and lifecycle management
- Tune storage throughput and latency for parallel workloads
- Architect multi region storage data locality replication consistency and durability
- Manage golden images templates GPU drivers CUDA and container registries
- Build monitoring capacity planning and operational runbooks
- Partner with Compute Network and Control Plane teams
Benefits
- Remote work within San Jose or Austin