Senior HPC Engineer Fleet Engineering
30 minutes agoSeniorSalary: 227K - 356KSan Francisco Office (Fremont St); Bellevue Office; Remote, USA; San Jose Office (First St)HybridFull TimeDevopsJobs by Lambda
LambdaVisit Lambda website
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
San Francisco, United States
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and monitor large-scale HPC clusters for AI workloads. You will automate deployment and lifecycle management, create runbooks and remediations, troubleshoot networking and GPU infrastructure, participate in on-call incident response, and improve operational procedures and engineering requirements.
Requirements
- 7+ years of experience in site reliability engineering, HPC engineering, DevOps, or a similar role.
- Knowledge of AI infrastructure, GPU architectures, and hardware performance optimization.
- Knowledge of Linux-based distributed systems.
- Experience configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE, Ethernet, switching, GPU-direct, and NCCL.
- Knowledge of Python and Go.
- Experience with Prometheus, Grafana, and ClickHouse.
- Proficiency with Ansible and Terraform.
- Problem-solving and troubleshooting skills.
Responsibilities
- Build and operate cluster-health monitoring and alerting.
- Deploy and configure large-scale HPC clusters for AI workloads.
- Automate operating systems, firmware, drivers, and networking with configuration-as-code tools.
- Create runbooks and automated remediations for cluster failure modes.
- Troubleshoot cluster issues across fabrics, networking, GPUs, switching, and power.
- Participate in on-call rotations and lead incident response for cluster-level problems.
- Maintain standard operating procedures and provide requirements to improve stability and operational efficiency.
Benefits
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401k plan with 2% company match for USA employees.
- Flexible paid time off.
