Member of Technical Staff ML Performance

Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.

New York City, United States
About Modal

Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.

View jobs by Modal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will optimize machine-learning systems for throughput and latency. You will contribute to open-source projects and container runtime performance work, improve language and diffusion model execution, and diagnose GPU performance bottlenecks across algorithms and host overhead.

Requirements

  • 5+ years of experience writing high-quality high-performance code
  • Experience with Torch, high-level ML frameworks, and inference engines
  • Familiarity with Nvidia GPU architecture and CUDA
  • Experience with machine-learning performance engineering

Responsibilities

  • Improve machine-learning system performance at scale
  • Contribute to open-source projects and container runtime development
  • Optimize language and diffusion model throughput and latency
  • Debug GPU performance bottlenecks
  • Improve algorithm efficiency and reduce host overhead

Benefits

  • Equity
Member of Technical Staff ML Performance at Modal | JobStash