Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and deliver systems spanning job scheduling and execution runtime. You will automate cluster operations, diagnose distributed-system failures, optimize GPU workloads, guide HPC infrastructure planning, mentor engineers, and collaborate with research staff on system designs.
Requirements
- 8+ years of professional experience developing business-critical software and operating large-scale compute infrastructure
- Go and/or Python proficiency
- Bachelor's degree in a related field, or equivalent advanced-degree experience
- Linux internals knowledge
- Knowledge of Docker container runtimes
- Experience designing, debugging, and optimizing distributed systems and databases
- Writing and consensus-building skills
Responsibilities
- Design and deliver systems spanning job scheduling and execution runtime
- Build tooling and software-defined infrastructure for cluster health management
- Analyze distributed-system failures and optimize distributed workloads
- Contribute to large-scale HPC roadmap planning
- Review code and design documents
- Mentor team members and improve team processes
- Collaborate with research staff on system designs and implementation
Benefits
- Medical, dental, vision, and employee assistance coverage for team members and families
- Health savings account, healthcare reimbursement arrangement, and flexible spending accounts
- 401k plan
- $125 monthly commuting or internet expense assistance
- $200 monthly fitness and wellbeing expense assistance
- Up to 10 sick days, 7 personal days, 20 vacation days, and 12 paid holidays annually
- Annual bonuses
- Long-term incentive plan