Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the architecture, capacity planning, and lifecycle management of a global Ceph topology. You will create documentation, post-mortems, and runbooks; partner on storage solutions; mentor engineers; lead zero-downtime upgrades; and resolve complex Ceph recovery scenarios while communicating during high-pressure events.
Requirements
- 8+ years of infrastructure engineering experience
- 5+ years focused on multi-petabyte production Ceph clusters
- Ceph expertise
- Technical documentation
- Technical communication
- Cross-team collaboration
- Pure native Ceph operations outside cloud wrappers such as OpenStack
- Linux storage stack
- Device mapper
- NVMe-over-Fabrics
- Asynchronous I/O
- Kernel and user-space boundaries
Responsibilities
- Lead the architectural design, capacity planning, and lifecycle management of global Ceph topology
- Write accessible storage documentation, post-mortems, and runbooks
- Partner with compute, networking, and platform teams on storage solutions
- Mentor mid-level and senior engineers
- Lead orchestration and testing pipelines for zero-downtime major-release upgrades
- Resolve complex Ceph recovery scenarios
- Communicate incident status to leadership during high-pressure events
Benefits
- Equity compensation
- Healthcare
- Lunch benefits
- Wellbeing benefits