Staff Software Engineer Managed Kubernetes
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will set technical direction for a bare-metal managed Kubernetes platform. You will develop GPU-aware orchestration and managed services, integrate NVIDIA tooling, and design resilient systems for AI workloads. You will guide cross-functional infrastructure decisions, lead design reviews and chaos engineering, establish operational automation, mentor engineers, and engage customers and open-source communities.
Requirements
- 10+ years of software engineering, platform engineering, or SRE experience
- At least 5 years focused on Kubernetes at scale
- Expert knowledge of Kubernetes internals, controllers, schedulers, operators, CRDs, CSI, and CNI
- Infrastructure expertise across compute, networking, storage, and security
- Production software engineering skills in Go and Python
- Experience with NVIDIA GPU orchestration in Kubernetes
- Technical leadership experience across teams
- Experience designing and operating managed services or multi-tenant platforms
- Knowledge of distributed systems principles
- Experience with Prometheus, Grafana, distributed tracing, and alerting
- Knowledge of Linux, networking, RDMA, InfiniBand, and RoCE
- Experience with infrastructure-as-code and GitOps
Responsibilities
- Drive technical vision for the Managed Kubernetes platform
- Design control-plane scalability, multi-tenancy, lifecycle management, and high availability
- Integrate NVIDIA GPU and network orchestration tooling
- Design GPU-aware orchestration systems
- Lead development of managed-service platforms
- Define networking and storage requirements for AI workloads
- Build Managed Slurm on Kubernetes
- Design inference-serving and autoscaling services
- Design self-healing and incident-response automation
- Lead chaos engineering
- Establish upgrade, patching, and zero-downtime maintenance practices
- Lead technical reviews and mentor engineers
Benefits
- Cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match for USA employees
- Flexible paid time off
