Senior Site Reliability Engineer SDN

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and scale multi-tenant cloud networking and SDN infrastructure, improve Kubernetes control-plane and SmartNIC dataplane software, and automate operational workflows. You will maintain monitoring tools, improve deployment safety, drive incident and capacity practices, and participate in on-call rotation.

Requirements

  • 5+ years in site reliability engineering, production engineering, or a similar role
  • Experience supporting large-scale distributed systems in production
  • Kubernetes lifecycle management and production operations experience
  • On-call and incident-response experience
  • Linux, Kubernetes, distributed-systems, and networking troubleshooting skills
  • Observability, monitoring, alerting, and metrics experience
  • Linux networking-stack knowledge
  • Multi-datacenter and hybrid-cloud experience
  • Python, Ansible, or similar infrastructure automation experience
  • CI/CD and GitOps deployment workflow experience

Responsibilities

  • Operate and scale multi-tenant cloud networking and SDN infrastructure
  • Operate and improve Kubernetes control-plane services and SmartNIC dataplane software
  • Develop tooling and automation to improve reliability
  • Collaborate on service reliability and deployment workflows
  • Deploy and maintain network monitoring, observability, and management tools
  • Improve deployment safety through CI/CD, GitOps, testing, and progressive rollouts
  • Drive observability, incident management, capacity planning, postmortems, and on-call operations

Benefits

  • Cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match for USA employees
  • Flexible paid time off
Senior Site Reliability Engineer SDN at Lambda | JobStash