Software Engineer - Model Performance Systems
31 minutes agoSeniorSalary: 165K - 330KSan Francisco; Montreal; New York; Seattle; TorontoHybridFull TimeAiJobs by Baseten
BasetenVisit Baseten website
Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.
San Francisco, United States
Funding history
About Baseten
Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.
Skills
About the Role
You will build automated LLM benchmarking and evaluation systems, GPU-enabled development environments, profiling tools, and monitoring dashboards. You will automate performance testing and release workflows, identify bottlenecks, and develop tools that optimize latency, cost, and model quality.
Requirements
- Python
- GPU memory subsystems
- InfiniBand
- Transformer mathematics
- FLOPs
- Memory requirements
- Stress testing
- Fuzz testing
Responsibilities
- Evaluate and automate LLM quality benchmarks and workload performance suites
- Develop and maintain GPU-enabled development environments
- Build open-source tools for model evaluation, benchmarking, and analysis
- Profile systems, identify bottlenecks, and debug compute and networking stacks
- Develop dashboards and alerts for system and runtime performance
- Automate performance testing and release workflows through CI/CD
- Build optimization tools for latency, cost, and quality tradeoffs
Benefits
- Equity
- 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
- Flexible PTO
- Company-wide winter break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) for U.S. employees
