Post-Training Research Engineer

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build in-house tooling for efficiently training different model architectures and post-training techniques at scale. You will work across systems infrastructure, distributed tensor computation, GPU kernels, storage, networking, and cluster platforms to improve training performance.

Requirements

  • Knowledge of modern ML techniques and transformer training
  • Advanced experience with PyTorch, TensorFlow, JAX, or a similar tensor computation library
  • Knowledge of transformer parallelism strategies
  • Ability to profile and improve distributed GPU programs
  • Ability to perform roofline analysis
  • Knowledge of HPC and distributed computing platforms including Slurm, Ray, Kubernetes, or Dask
  • Knowledge of InfiniBand, RoCE, or GPUDirect
  • Knowledge of operating systems, containerization, and networking protocols

Responsibilities

  • Build in-house tooling for post-training models
  • Support efficient training of model architectures and techniques at scale
  • Work across systems infrastructure and distributed tensor computation
  • Profile and improve distributed GPU program performance

Benefits

  • Equity
  • 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
  • Flexible PTO
  • Company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees
Post-Training Research Engineer at Baseten | JobStash