Software Engineer - Model Performance
31 minutes agoSalary: 180K - 360KSan Francisco; Montreal; New York; Seattle; TorontoHybridFull TimeAiJobs by Baseten
BasetenVisit Baseten website
Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.
San Francisco, United States
Funding history
About Baseten
Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.
Skills
About the Role
You will optimize and productionize ML inference techniques, investigate performance issues in core ML libraries, and scale improvements across large language models. You will design solutions with engineers and own projects from initial idea through production deployment.
Requirements
- Experience with Python or C++
- Familiarity with LLM optimization techniques
- Strong familiarity with PyTorch, TensorRT, or TensorRT-LLM
- Experience and interest in LLMs
- Deep understanding of GPU architecture
Responsibilities
- Implement, refine, and productionize ML inference techniques and infrastructure
- Debug ML performance issues in TensorRT, PyTorch, TensorRT-LLM, vLLM, SGLang, CUDA, and related libraries
- Apply optimization techniques across large language models
- Design and implement solutions with engineers
- Own projects from idea to production
Benefits
- Equity
- 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
- Flexible PTO
- Company-wide winter break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) for U.S. employees
