Software Engineer - Model Performance

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

About the Role

You will optimize and productionize ML inference techniques, investigate performance issues in core ML libraries, and scale improvements across large language models. You will design solutions with engineers and own projects from initial idea through production deployment.

Requirements

  • Experience with Python or C++
  • Familiarity with LLM optimization techniques
  • Strong familiarity with PyTorch, TensorRT, or TensorRT-LLM
  • Experience and interest in LLMs
  • Deep understanding of GPU architecture

Responsibilities

  • Implement, refine, and productionize ML inference techniques and infrastructure
  • Debug ML performance issues in TensorRT, PyTorch, TensorRT-LLM, vLLM, SGLang, CUDA, and related libraries
  • Apply optimization techniques across large language models
  • Design and implement solutions with engineers
  • Own projects from idea to production

Benefits

  • Equity
  • 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
  • Flexible PTO
  • Company-wide winter break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees
Software Engineer - Model Performance at Baseten | JobStash