Research Engineer Infrastructure Numerics

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will improve the numerical foundations of distributed large-model training. You will optimize precision formats, kernels, communication primitives, parallelism strategies, orchestration, monitoring, stability, and reproducibility.

Requirements

  • Understanding of PyTorch, JAX, and their underlying system architectures
  • Strong engineering, performance, maintainability, and debugging skills in floating-point numerics, low-precision arithmetic, and distributed systems
  • Ability to work with cross-functional partners and subject matter experts

Responsibilities

  • Design and optimize distributed training infrastructure for large-scale LLMs
  • Implement and evaluate low-precision numerics
  • Develop kernels and communication primitives for mixed- and low-precision arithmetic
  • Co-design model architectures and training recipes with research teams
  • Prototype and benchmark data, tensor, and pipeline parallelism strategies
  • Contribute to orchestration and monitoring systems for distributed experiments
  • Publish and share learnings through documentation, open-source libraries, or technical reports

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship