Senior Incident Manager
48 minutes agoLeadSalary: 125K - 195KRemote, USA; San Jose Office (Zanker)RemoteFull TimeOperationsJobs by Lambda
LambdaVisit Lambda website
Lambda is an AI infrastructure company providing GPU supercomputers and cloud capacity for AI training and inference.
San Francisco, United States
About Lambda
Lambda, Inc. builds and operates AI-focused compute infrastructure, including single-tenant Superclusters, deployable 1-Click Clusters, and on-demand GPU Instances for researchers, enterprises, and hyperscalers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the full lifecycle of critical incidents affecting AI infrastructure and data center services. You will coordinate triage, escalation, resolution, communication, post-incident reviews, reliability improvements, operational reporting, and on-call incident response.
Requirements
- 8+ years of experience in incident management, site reliability engineering, or infrastructure operations
- Experience managing incidents in large-scale distributed infrastructure environments
- Understanding of data center operations, GPU compute clusters, networking, storage, and cloud or hybrid infrastructure
- Experience leading high-pressure incident response situations
- Experience with ITIL, SRE, or equivalent incident management frameworks
- Experience with PagerDuty, ServiceNow, Jira, Datadog, Prometheus, or Grafana
Responsibilities
- Lead responses to critical incidents affecting AI infrastructure, GPU clusters, networking, storage, and data center operations
- Serve as Incident Commander during major outages
- Coordinate engineering, networking, facilities, and vendor teams
- Own incident triage, escalation, coordination, resolution, and post-incident review
- Maintain incident documentation and operational playbooks
- Analyze incident patterns and improve system reliability
- Lead post-incident reviews and root cause analysis
- Track incident metrics, including MTTR, MTTD, and recurrence rates
- Improve incident processes, escalation paths, tooling, runbooks, and reliability frameworks
- Provide executive incident summaries and maintain operational health reporting
- Participate in an on-call rotation
Benefits
- Cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match for USA employees
- Flexible paid time off
