Research Engineer Infrastructure Kernels

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop, profile, and optimize ML kernels for large-scale language-model training. You will improve low-precision computation, GPU efficiency, distributed infrastructure, reusable kernel libraries, benchmarks, and system reproducibility.

Requirements

  • Understanding of PyTorch, JAX, and their underlying system architectures
  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks
  • Ability to analyze, profile, and optimize compute-intensive workloads
  • Strong engineering, performance, maintainability, and debugging skills
  • Ability to work with cross-functional partners and subject matter experts

Responsibilities

  • Design and implement custom ML kernels for LLM operations
  • Develop compute primitives that reduce memory-bandwidth bottlenecks
  • Align kernel-level optimizations with model architecture and algorithmic goals
  • Develop and maintain reusable kernel libraries and performance benchmarks
  • Improve infrastructure stability, scalability, reproducibility, and compute utilization
  • Document and share technical insights through talks, papers, or open-source contributions

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship
Research Engineer Infrastructure Kernels at Thinking Machines Lab | JobStash