Senior Platform Monitoring Engineer
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead platform incident investigations from detection through mitigation and resolution. You will perform root-cause analyses, design alerting and observability workflows, build monitoring automation, address reliability gaps, mentor junior engineers, and participate in an on-call rotation.
Requirements
- SRE
- DevOps
- Production engineering
- AWS
- Azure
- GCP
- Docker
- Kubernetes
- ELK
- Prometheus
- Grafana
- PagerDuty
- Monitoring
- Logging
- Alerting
- Metric
- Log
- Trace
- Python
- Automation
- Incident management
- Root cause analysis
- Computer Science
- Computer Engineering
Responsibilities
- Lead platform incident investigations and coordinate rapid resolution
- Conduct post-incident root-cause analysis
- Design and implement alerting pipelines and observability workflows
- Build automation tools and reusable monitoring patterns
- Resolve reliability gaps affecting customer experience
- Mentor junior engineers on observability and service health metrics
- Participate in the on-call rotation
Benefits
- Annual performance bonus eligibility
- Equity eligibility
