Senior Site Reliability Engineer SDN
48 minutes agoSeniorSalary: 240K - 312KSan Francisco Office (Fremont St); Bellevue Office; San Jose Office (First St)HybridFull TimeDevopsJobs by Lambda
LambdaVisit Lambda website
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
San Francisco, United States
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and scale multi-tenant cloud networking and SDN infrastructure, improve Kubernetes control-plane and SmartNIC dataplane software, and automate operational workflows. You will maintain monitoring tools, improve deployment safety, drive incident and capacity practices, and participate in on-call rotation.
Requirements
- 5+ years in site reliability engineering, production engineering, or a similar role
- Experience supporting large-scale distributed systems in production
- Kubernetes lifecycle management and production operations experience
- On-call and incident-response experience
- Linux, Kubernetes, distributed-systems, and networking troubleshooting skills
- Observability, monitoring, alerting, and metrics experience
- Linux networking-stack knowledge
- Multi-datacenter and hybrid-cloud experience
- Python, Ansible, or similar infrastructure automation experience
- CI/CD and GitOps deployment workflow experience
Responsibilities
- Operate and scale multi-tenant cloud networking and SDN infrastructure
- Operate and improve Kubernetes control-plane services and SmartNIC dataplane software
- Develop tooling and automation to improve reliability
- Collaborate on service reliability and deployment workflows
- Deploy and maintain network monitoring, observability, and management tools
- Improve deployment safety through CI/CD, GitOps, testing, and progressive rollouts
- Drive observability, incident management, capacity planning, postmortems, and on-call operations
Benefits
- Cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match for USA employees
- Flexible paid time off
