Incident Manager
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead critical production incidents, coordinate cross-functional responses, provide timely internal and customer communications, investigate root causes, and drive reliability improvements. You will also develop incident practices and mentor peers in technical response and communications.
Requirements
- 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale cloud-native systems.
- Ability to lead and coordinate high-severity incidents.
- Understanding of AWS, Azure, or GCP infrastructure.
- Expertise in log analysis and debugging.
- Experience with log aggregation, observability, metrics, logging, and tracing tools.
- Proficiency in Python, Go, or Bash.
- Experience developing and maintaining incident playbooks and communication templates.
- Excellent interpretation, writing, and communication skills.
- Degree in Computer Science, Computer Engineering, or a related engineering field.
Responsibilities
- Lead critical incidents and coordinate multi-disciplinary response efforts.
- Drive technical root cause analysis and reliability improvements.
- Communicate incident updates to internal stakeholders and customers.
- Mentor and train peers in incident communication and technical response.
