Senior Site Reliability Engineer Managed Kubernetes

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and maintain large bare-metal Kubernetes clusters, handle incidents and recovery, support customers, and participate in on-call coverage. You will build control-plane services and controllers, automate cluster lifecycle management, and define reliability indicators and objectives for Kubernetes services.

Requirements

  • 6+ years in SRE, operations engineering, or a similar role
  • Knowledge of Linux clusters and systems
  • Go and Python programming skills
  • GitOps, ArgoCD, Helm, and Kubernetes operator experience
  • Production Kubernetes operations experience
  • Customer incident-support experience
  • Prometheus, Grafana, FluentBit, and CI/CD familiarity
  • Experience provisioning Kubernetes with kubeadm, Cluster API, or similar tools

Responsibilities

  • Operate and maintain bare-metal Kubernetes clusters
  • Handle cluster degradation, recovery, resizing, and incident response
  • Participate in on-call rotation for critical incidents
  • Assist customers with Kubernetes integration, storage, and authentication
  • Create tooling and automate platform-quality validation using Python and Go
  • Design and maintain Kubernetes control-plane services, operators, and controllers
  • Automate provisioning, upgrades, patching, and deletion
  • Define and implement SLOs and SLIs for Kubernetes services

Benefits

  • Cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match for USA employees
  • Flexible paid time off
Senior Site Reliability Engineer Managed Kubernetes at Lambda | JobStash