Software Engineer - GPU Networking and Distributed Systems
31 minutes agoLeadSalary: 165K - 330KSan Francisco; Montreal; New York; Seattle; TorontoHybridFull TimeEngineeringJobs by Baseten
BasetenVisit Baseten website
Baseten is an AI inference platform for deploying, optimizing, and scaling custom, open-source, and fine-tuned models in production.
San Francisco, United States
Funding history
About Baseten
Baseten provides model runtimes, inference infrastructure, developer workflows, and deployment options including managed cloud, self-hosted, and hybrid environments.
Skills
About the Role
You will integrate RDMA networking into inference infrastructure, optimize communication for distributed model serving, and improve model startup speeds. You will validate hardware performance, build interconnect observability, and develop communication-library and kernel optimizations that overlap compute with data transfer.
Requirements
- InfiniBand
- RoCE v2
- C++
- Python
- NVIDIA architecture
- Memory hierarchy
- TensorRT-LLM
- C++ binding
- Python binding
- NVLink
- Kubernetes networking
Responsibilities
- Integrate RDMA, RoCE, and InfiniBand capabilities into the inference stack
- Implement and tune networking layers for distributed inference
- Develop checkpointing and storage mechanisms for faster model startup
- Characterize and validate networking performance on GPU clusters
- Write hardware acceptance tests for throughput and latency
- Design observability tools for packet flow, congestion, and bandwidth
- Optimize communication libraries and kernels to overlap compute and data transfer
Benefits
- Equity
- 100% medical, dental, and vision insurance coverage for U.S. employees and dependents
- Flexible PTO
- Company-wide winter break
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k) for U.S. employees
