Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and maintain research-oriented software and infrastructure for AI evaluations. You may develop Inspect sandbox plugins, add support for open-weight models, design custom research infrastructure, collaborate with the open-source community, and debug evaluation issues.
Requirements
- Write production-quality code efficiently
- Demonstrate Python expertise and familiarity with its ecosystem and tooling
- Communicate effectively in writing and verbally
- Work effectively with others and solve challenging problems
Responsibilities
- Develop features for Inspect sandbox plugins that support agentic evaluations
- Implement support for open-weight models on the model-hosting platform
- Design custom infrastructure for research projects
- Collaborate with the open-source community on Inspect and its plugin ecosystem
- Debug Inspect issues during frontier-model evaluation testing
Benefits
- Pre-release access to multiple frontier models and ample compute
- Operational support
- Annual learning and development stipends
- Conference and external-collaboration funding
- Hybrid working and flexibility for occasional remote work abroad
- Work-from-home equipment stipend
- At least 25 days of annual leave
- 8 public holidays
- Extra team-wide breaks
- 3 volunteering days
- Paid parental leave
- Employer pension contribution of 28.97% of base salary
- Cycling, donation, retail, and gym discounts and benefits