Senior Incident Manager

Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.

San Francisco, United States
About Lambda

Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.

View jobs by Lambda

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead the full lifecycle of critical incidents affecting AI infrastructure and data center services. You will coordinate triage, escalation, resolution, communication, post-incident reviews, reliability improvements, operational reporting, and on-call incident response.

Requirements

  • 8+ years of experience in incident management, site reliability engineering, or infrastructure operations
  • Experience managing incidents in large-scale distributed infrastructure environments
  • Understanding of data center operations, GPU compute clusters, networking, storage, and cloud or hybrid infrastructure
  • Experience leading high-pressure incident response situations
  • Experience with ITIL, SRE, or equivalent incident management frameworks
  • Experience with PagerDuty, ServiceNow, Jira, Datadog, Prometheus, or Grafana

Responsibilities

  • Lead responses to critical incidents affecting AI infrastructure, GPU clusters, networking, storage, and data center operations
  • Serve as Incident Commander during major outages
  • Coordinate engineering, networking, facilities, and vendor teams
  • Own incident triage, escalation, coordination, resolution, and post-incident review
  • Maintain incident documentation and operational playbooks
  • Analyze incident patterns and improve system reliability
  • Lead post-incident reviews and root cause analysis
  • Track incident metrics, including MTTR, MTTD, and recurrence rates
  • Improve incident processes, escalation paths, tooling, runbooks, and reliability frameworks
  • Provide executive incident summaries and maintain operational health reporting
  • Participate in an on-call rotation

Benefits

  • Cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match for USA employees
  • Flexible paid time off
Senior Incident Manager at Lambda | JobStash