Research Engineer - ML Infrastructure

AI molecular-design company building models and a computer-aided design suite for drug discovery.

Series CRecently funded0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

San Francisco, United States
About Chai Discovery

Chai Discovery develops generative AI software that predicts and reprograms biochemical molecular interactions to help scientists design biomolecules and medicines.

View jobs by Chai Discovery

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop core frameworks for model training and evaluation that make models performant, resource-efficient, and reliable at scale. You will build and optimize the training stack, profile GPU-cluster workloads, remove bottlenecks, and ensure fault-tolerant, deterministic execution for long-running jobs.

Requirements

  • Deep industry experience with AI/ML teams
  • Software system design skills
  • Python proficiency
  • PyTorch or JAX fluency
  • Demonstrated exceptional impact on real-world problems and systems

Responsibilities

  • Build the training stack across model, layer, and kernel levels
  • Optimize workloads through parallelism, quantization, and custom kernels
  • Profile end-to-end training runs on large GPU clusters
  • Eliminate bottlenecks and failures
  • Monitor throughput, utilization, and uptime
  • Ensure model architectures and training recipes scale efficiently
  • Implement fault tolerance, checkpointing, and deterministic orchestration for large-scale jobs