Senior Engineering Manager AI Runtime
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Maintainer signals as of 9/23/2026
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead and develop engineers responsible for the product experience and foundational infrastructure for managed GPU training. You will set the product and technical roadmap, make architecture and product-design decisions, and deliver solutions with platform, product, infrastructure, research, and customers. You will improve observability and reliability for multi-node training jobs, advocate for customer needs, and help recruit and develop engineering talent.
Requirements
- 8+ years of software engineering experience
- 3+ years of engineering management experience
- Experience building and operating managed GPU training infrastructure at scale
- Familiarity with PyTorch, DeepSpeed, Composer, Megatron-LM, FSDP, tensor parallelism, and pipeline parallelism
- Experience with checkpointing, elastic training, and automated failure recovery
- Understanding of NCCL, interconnect topologies, and memory optimization
- Experience building platform products with SLAs and owning customer experience
- Cross-functional leadership across platform, product, and research
- Collaboration and communication skills
- BS/MS in Computer Science, Electrical Engineering, or a related technical field
Responsibilities
- Lead, mentor, and grow the engineering team responsible for Custom Training and its foundational infrastructure
- Define and own the product and technical roadmap
- Collaborate across platform, product, infrastructure, research, and customer stakeholders to deliver products
- Drive architectural decisions and product design for managed GPU training at scale
- Advocate for customer needs through direct engagement
- Build observability and reliability practices for multi-node training jobs
- Partner with recruiting to attract, hire, and develop engineering talent
Benefits
- Annual performance bonus eligibility
- Equity eligibility
