Member of Technical Staff ML Performance
Modal Labs, Inc. operates Modal, a serverless cloud and AI infrastructure platform for developers running inference, training, batch processing, and isolated sandboxes.
Funding history
About Modal
Modal provides code-first, elastic CPU/GPU compute infrastructure for AI workloads, including model inference, fine-tuning and training, large-scale batch jobs, and secure ephemeral execution environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will optimize machine-learning systems for throughput and latency. You will contribute to open-source projects and container runtime performance work, improve language and diffusion model execution, and diagnose GPU performance bottlenecks across algorithms and host overhead.
Requirements
- 5+ years of experience writing high-quality high-performance code
- Experience with Torch, high-level ML frameworks, and inference engines
- Familiarity with Nvidia GPU architecture and CUDA
- Experience with machine-learning performance engineering
Responsibilities
- Improve machine-learning system performance at scale
- Contribute to open-source projects and container runtime development
- Optimize language and diffusion model throughput and latency
- Debug GPU performance bottlenecks
- Improve algorithm efficiency and reduce host overhead
Benefits
- Equity
