Staff Software Engineer, Model Serving
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Maintainer signals as of 9/23/2026
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design and build high-throughput, low-latency inference systems for CPU and GPU workloads. You will define serving architecture and roadmaps, optimize autoscaling and operational efficiency, improve latency and availability, mentor engineers, and deliver reliable model-serving capabilities.
Requirements
- Large-scale distributed system
- Model serving
- Inference system
- Routing
- Scheduling
- Autoscaling
- Observability
- Algorithm
- Data structure
- System design
- CPU
- GPU
- Communication
- Mentoring
Responsibilities
- Design and implement scalable model-serving systems and APIs
- Define the technical roadmap and long-term serving architecture
- Optimize performance, throughput, autoscaling, and operational efficiency
- Build model container, deployment, routing, caching, observability, and autoscaling capabilities
- Translate customer needs into reliable and performant systems
- Improve latency, availability, and cost-effectiveness
- Establish code quality, testing, and operational-readiness practices
- Mentor engineers through design reviews and technical guidance
Benefits
- Annual performance bonus eligibility
- Equity eligibility
