Research Safety

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will investigate how models handle harmful and dual-use requests and design experiments that guide training and evaluation. Depending on project needs, you will build filtering pipelines and classifiers, apply safety-focused post-training methods, create evaluations and synthetic data, red-team models, and design mitigations.

Requirements

  • Bachelor’s degree or equivalent experience in a relevant discipline
  • AI safety research experience
  • Experience with RLHF, RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red-teaming
  • Proficiency in Python
  • Familiarity with PyTorch, TensorFlow, or JAX
  • Ability to debug distributed training and write scalable code
  • Written technical communication
  • Experience with long-horizon, multi-step, or agentic evaluations
  • Synthetic data generation
  • Red-teaming and jailbreaking
  • Knowledge of scalable oversight, reward hacking, and jailbreak robustness

Responsibilities

  • Build data-filtering pipelines and quality classifiers for pre-training corpora
  • Apply safety-focused post-training techniques
  • Design, build, and maintain safety evaluations for long-horizon and agentic tasks
  • Generate and curate synthetic data for training and evaluation
  • Red-team models and products to identify failure modes, jailbreaks, and risks
  • Design mitigations for identified risks

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
Research Safety at Thinking Machines Lab | JobStash