About the Role
You will develop data mixes for LLM training and build pipelines that process petabyte-scale datasets. You will create systems for web crawling, ingestion, storage, retrieval, and versioning; evaluate data quality and diversity; and ensure data collection follows privacy regulations.
Requirements
- BS, MS, or PhD in Computer Science, Machine Learning, or a related field, or equivalent experience
- 3 or more years of experience building data-processing pipelines at scale for AI or ML applications
- Proficiency in Python and experience with Apache Spark, Beam, and Airflow
- Familiarity with synthetic data generation and data augmentation
- Familiarity with web scraping, crawling technologies, and Common Crawl datasets
- Understanding of machine-learning fundamentals and experience with PyTorch or TensorFlow
- Experience with SQL and NoSQL databases
Responsibilities
- Develop data mixes for LLM training using open-source datasets, synthetic data, and curated human feedback
- Design and implement data pipelines for petabyte-scale datasets
- Build systems for web crawling, data ingestion, and real-time data processing
- Develop tools and frameworks for data storage, retrieval, and versioning across distributed systems
- Create evaluation frameworks for data diversity, quality, and representativeness
- Ensure data collection adheres to privacy regulations
Benefits
- Equity
- Flexible vacation and paid time off
- Health, dental, and vision insurance
- 401k match
- Catered meals
- Commuter subsidies