Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead RL-based post-training for video diffusion and flow-matching models at multi-node scale. You will build video reward models, design preference-data workflows, train and validate learned judges, prevent reward hacking, run human and automated evaluations, and distill aligned models into efficient samplers.
Requirements
- 2+ years of hands-on research experience in post-training or generative modeling
- RL or preference-optimization experience on generative models with evidence of model improvement
- Knowledge of diffusion or flow-matching models
- PyTorch proficiency
- Multi-node distributed training experience
Responsibilities
- Run RL post-training for video diffusion and flow-matching models at multi-node scale
- Build video reward models and define evaluation dimensions
- Design and configure preference-data collection workflows
- Train and validate learned judges and safeguard against reward hacking
- Own post-training evaluation using human preference studies and automated metrics
- Distill RL-tuned models into efficient few-step samplers while preserving alignment gains
Benefits
- Equity
- Health benefits
- 401k matching
- Flexible onsite/remote hybrid work