Post-Training Research Engineer
BasetenVisit Baseten website
Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.
San Francisco, United States
Funding history
About Baseten
Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build in-house tooling for efficiently training different model architectures and post-training techniques at scale. You will work across systems infrastructure, distributed tensor computation, GPU kernels, storage, networking, and cluster platforms to improve training performance.
Requirements
- Knowledge of modern ML techniques and transformer training
- Advanced experience with PyTorch, TensorFlow, JAX, or a similar tensor computation library
- Knowledge of transformer parallelism strategies
- Ability to profile and improve distributed GPU programs
- Ability to perform roofline analysis
- Knowledge of HPC and distributed computing platforms including Slurm, Ray, Kubernetes, or Dask
- Knowledge of InfiniBand, RoCE, or GPUDirect
- Knowledge of operating systems, containerization, and networking protocols
Responsibilities
- Build in-house tooling for post-training models
- Support efficient training of model architectures and techniques at scale
- Work across systems infrastructure and distributed tensor computation
- Profile and improve distributed GPU program performance
Benefits
- Equity
- 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
- Flexible PTO
- Company-wide Winter Break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) for U.S. employees
