Senior Software Engineer - Observability
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Maintainer signals as of 9/23/2026
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will establish logging, metrics, and tracing standards. You will build infrastructure for emitting, aggregating, storing, and displaying metrics; improve scalability and reliability; participate in on-call rotations; and reduce observability costs through analysis and resource optimization.
Requirements
- 5+ years of production-level experience with Python, Java, Scala, C++, or a similar language.
- Experience developing large-scale distributed systems.
- Familiarity with metrics collection, health monitoring, and observability tools.
Responsibilities
- Establish standards for logging, metrics, and tracing.
- Collaborate with teams to identify system-performance metrics.
- Build infrastructure to emit, aggregate, store, and display metrics.
- Support dashboards and alerting.
- Contribute to the technical roadmap for scalability, performance, and reliability.
- Participate in on-call rotations and reduce incident response times.
- Analyze observability expenses and optimize retention, queries, and resources.
