Software Engineer - Inference Performance

Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.

San Francisco, United States
About Baseten

Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.

View jobs by Baseten

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will optimize inference across runtime internals, kernels, scheduling, serving, and routing. You will productionize inference techniques, investigate latency and memory bottlenecks, benchmark performance, tune models on new hardware, and contribute to open-source inference engines.

Requirements

  • Experience with Python or C++
  • Familiarity with LLM optimization techniques
  • Familiarity with PyTorch, TensorRT, or TensorRT-LLM
  • Experience with LLMs
  • Understanding of GPU architecture

Responsibilities

  • Implement and productionize inference techniques in runtime internals
  • Profile and optimize inference from kernels through scheduling and routing
  • Improve GPU utilization, latency, throughput, and cost
  • Bring up and tune model architectures on new hardware
  • Build benchmarking frameworks across models and hardware configurations
  • Contribute to open-source inference engines

Benefits

  • Equity
  • Medical, dental, and vision insurance for U.S. employees and dependents
  • Flexible PTO
  • Company-wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k) for U.S. employees