Senior Staff Technical Program Manager - Reliability

Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.

Series F+0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/23/2026

160 Spear Street, Suite 1300, San Francisco, CA 94105, United States
About Databricks

Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.

View jobs by Databricks

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead multi-quarter reliability roadmaps and critical infrastructure programs. You will manage planning, risks, dependencies, trade-offs, reporting, and delivery; align engineering partners; improve reliability practices; and establish scalable governance, metrics, processes, and documentation.

Requirements

  • 10+ years managing and delivering large-scale technical programs in cloud infrastructure, distributed systems, SRE, or platform engineering
  • Experience developing infrastructure at two or more hyperscale cloud providers
  • Knowledge of cloud primitives, multi-availability-zone and regional architecture, and control-plane and data-plane patterns
  • Experience leading large-scale reliability programs
  • Understanding of infrastructure, distributed systems, or SRE practices
  • Experience partnering with senior engineering leadership on multi-team initiatives
  • Ability to create program plans with milestones, KPIs, and success metrics
  • Experience managing cross-organizational dependencies, technical risks, and multi-quarter timelines
  • Experience delivering multi-cloud or large-scale cloud-native service programs
  • Experience building engineering processes and operational frameworks
  • Experience with compute fleets, container orchestration, autoscaling, or control-plane architecture preferred
  • Knowledge of SLOs, error budgets, chaos engineering, failure mode analysis, and incident management preferred
  • Jira or equivalent program-tracking tools preferred

Responsibilities

  • Define long-term reliability roadmaps and align engineering teams
  • Own end-to-end program planning, risk management, dependency mapping, reporting, and delivery
  • Identify process and architecture gaps and drive improvements
  • Facilitate cross-functional technical alignment and prioritization
  • Diagnose reliability bottlenecks and improve scalability, fault tolerance, automation, and operational tooling
  • Drive adoption of error budgets, incident reviews, resilience patterns, and operational readiness
  • Implement program governance, processes, metrics, and documentation

Benefits

  • Annual performance bonus eligibility
  • Equity eligibility