Staff Software Engineer Distributed Data Systems
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Maintainer signals as of 9/24/2026
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build next-generation distributed data storage and processing systems for diverse workloads from ETL to data science. You may develop Apache Spark, cloud storage services and client libraries, Delta Lake, Delta Pipelines, or query optimization and execution-engine capabilities.
Requirements
- BS in Computer Science, a related technical field, or equivalent practical experience
- 8+ years of production experience with Java, Scala, or C++
- Foundation in algorithms and data structures
- Experience with distributed systems, databases, and big-data systems including Apache Spark and Hadoop
- MS or PhD in databases or distributed systems is optional
Responsibilities
- Build distributed data storage and processing systems
- Develop Apache Spark
- Deliver cloud storage services and client libraries
- Develop Delta Lake storage-management capabilities
- Develop and operate data-pipeline orchestration capabilities
- Build query optimization and execution-engine capabilities
Benefits
- Annual performance bonus eligibility
- Equity eligibility
