Research Engineer Infrastructure Training Systems

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own and optimize distributed training systems for large-scale workloads. You will improve GPU throughput, efficiency, reproducibility, reliability, maintainability, security, and reusable training infrastructure.

Requirements

  • Strong engineering, performance, maintainability, and debugging skills
  • Understanding of PyTorch, JAX, and their underlying system architectures
  • Ability to work with cross-functional partners and subject matter experts

Responsibilities

  • Design, implement, and optimize distributed training systems across thousands of GPUs and nodes
  • Develop high-performance optimizations for throughput and efficiency
  • Develop reusable frameworks and libraries for training reproducibility, reliability, and scalability
  • Establish reliability, maintainability, and security standards
  • Build scalable infrastructure with researchers and engineers
  • Publish and share learnings through documentation, open-source libraries, or technical reports

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship