Staff Software Engineer Managed Kubernetes

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will set technical direction and build GPU-aware managed Kubernetes and orchestration services. You will design scalable control planes, platform services, networking and storage requirements, resilience automation, and operational practices while mentoring engineers and working with customers and infrastructure partners.

Requirements

  • 10+ years of software engineering, platform engineering, or SRE experience
  • 5+ years focused on Kubernetes at scale
  • Expert knowledge of Kubernetes API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns
  • Go and Python software engineering expertise
  • GPU orchestration experience including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling
  • Technical leadership and cross-team design influence experience
  • Managed services or multi-tenant platform experience
  • Distributed systems knowledge including consensus, fault tolerance, consistency, and graceful degradation
  • Prometheus, Grafana, distributed tracing, and alerting experience
  • Linux, L2-L7 networking, RDMA, InfiniBand, and RoCE knowledge
  • Infrastructure-as-code and GitOps experience

Responsibilities

  • Drive technical vision for the bare-metal Managed Kubernetes platform
  • Integrate and extend NVIDIA GPU, network, and orchestration tooling
  • Design GPU-aware orchestration systems and managed services
  • Define networking and storage requirements for AI workloads
  • Build Managed Slurm on Kubernetes foundations
  • Design inference platform services, autoscaling, and multi-model deployment patterns
  • Design self-healing automation, incident response, root-cause analysis, and platform resilience
  • Lead chaos engineering and establish upgrade, patching, and zero-downtime maintenance practices
  • Translate platform requirements across Network, Storage, and Security
  • Set Kubernetes technical direction, lead reviews, mentor engineers, and establish engineering practices
  • Engage with customers, NVIDIA, and the open-source community
  • Design AIOps systems for capacity planning, anomaly detection, and predictive maintenance

Benefits

  • Equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness stipend for select roles
  • Commuter stipend for select roles
  • 401k plan with 2% company match for US employees
  • Flexible paid time off