Research Engineer Infrastructure Inference

Artificial-intelligence research and product company building customizable AI systems, including the Tinker training API and Inkling open-weight models.

Distributed
About Thinking Machines Lab

Thinking Machines Lab develops AI products that let researchers and developers fine-tune and use models, while also releasing open-weight multimodal models.

View jobs by Thinking Machines Lab

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, optimize, and scale infrastructure for large-model inference. You will improve performance, latency, throughput, reliability, and reproducibility; extend distributed inference and serving frameworks; optimize GPU utilization; and document and share technical learnings.

Requirements

  • Understanding of PyTorch, JAX, and their underlying system architectures
  • Experience with inference serving systems optimized for throughput and latency
  • Strong engineering, performance, maintainability, and debugging skills
  • Ability to work with cross-functional partners and subject matter experts

Responsibilities

  • Bring AI models into production with researchers and engineers
  • Enable high-performance inference for novel architectures
  • Design and implement techniques, tools, and architectures that improve inference performance, latency, throughput, and efficiency
  • Optimize the codebase and GPU compute fleet
  • Extend Kubernetes, Ray, and SLURM orchestration frameworks for distributed inference, evaluation, and large-batch serving
  • Establish reliability, observability, and reproducibility standards across the inference stack
  • Publish and share learnings through documentation, open-source libraries, or technical reports

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support
  • Visa sponsorship