Research Pre Training Data
Thinking Machines LabVisit Thinking Machines Lab website
Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.
Distributed
Funding history
About Thinking Machines Lab
Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will research, source, curate, filter, and analyze large-scale text, code, and multimodal training data. You will develop data-quality metrics, improve scalable data-processing systems, assess privacy, safety, and licensing risks, evaluate downstream model effects, and publish research, code, datasets, and insights.
Requirements
- Proficiency in Python
- Familiarity with PyTorch, TensorFlow, or JAX
- Ability to debug distributed training and write scalable code
- Bachelor’s degree or equivalent experience in a relevant discipline
- Written technical communication
- Probability, statistics, and machine learning fundamentals
- Experience curating, preprocessing, or analyzing large-scale datasets
- Knowledge of data ethics, safety, and licensing frameworks for AI datasets
Responsibilities
- Design and implement techniques for curating, sourcing, and filtering large-scale text, code, and multimodal data
- Develop data-quality metrics and analysis for coverage, diversity, and representativeness
- Scale data-processing systems efficiently and reproducibly
- Investigate and mitigate privacy, safety, and licensing risks
- Evaluate dataset improvements through downstream model learning and behavior
- Publish and present research and share code, datasets, and insights
Benefits
- Health, dental, and vision benefits
- Unlimited PTO
- Paid parental leave
- Relocation support as needed
