Research RL Scaling

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will co-design reinforcement-learning recipes and training systems at frontier scale. You will advance asynchronous RL, improve rollout generation and its training integration, run and stabilize large training jobs, optimize compute and training efficiency, and conduct well-instrumented ablations and scaling studies.

Requirements

  • Proficiency in Python
  • Familiarity with PyTorch, TensorFlow, or JAX
  • Ability to debug distributed training and write scalable code
  • Bachelor’s degree or equivalent experience in a relevant discipline
  • Written technical communication
  • Research judgment, ablation design, and baseline evaluation
  • Reinforcement learning for large language models
  • Asynchronous reinforcement learning
  • Distributed training
  • Inference systems
  • Low-precision training and inference
  • Quantization
  • LLM serving

Responsibilities

  • Co-design reinforcement-learning recipes and systems and validate them at frontier scale
  • Advance asynchronous reinforcement-learning algorithms
  • Improve rollout-generation efficiency and integration with training
  • Run frontier-scale reinforcement learning end to end and maintain stable training runs
  • Optimize accelerator utilization, memory, communication, and low-precision numerics
  • Conduct ablations and scaling studies backed by reliable instrumentation

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
Research RL Scaling at Thinking Machines Lab | JobStash