Staff Software Engineer AI Research Infrastructure

Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.

160 Spear Street, Suite 1300, San Francisco, CA 94105, United States
About Databricks

Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.

View jobs by Databricks

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will develop and operate the research stack for large-scale training and inference experiments. You will build scheduling, orchestration, monitoring, developer tooling, and repeatable research pipelines. You will improve infrastructure reliability, efficiency, and security while mentoring engineers and shaping the research-computing roadmap.

Requirements

  • BS, MS, or PhD in Computer Science or a related field
  • 5+ years of software engineering experience, including large-scale distributed systems or infrastructure
  • Experience building and operating distributed systems, data pipelines, or large-scale backend services
  • Proficiency in C++, Rust, Go, Java, Scala, or another systems programming language
  • Experience contributing to cluster schedulers, resource managers, or job orchestration systems such as Kubernetes, Slurm, or Ray
  • Understanding of modern ML training and inference workflows, including distributed training, model parallelism, fine-tuning, and evaluation
  • Experience taking complex systems from prototype to stable services
  • Clear communication with researchers and engineers

Responsibilities

  • Design and implement infrastructure for large-scale experiments, data processing, and model training
  • Build abstractions for job submission, scheduling, and monitoring
  • Create experiment-management systems, CI/testing infrastructure, and research workflows
  • Shape the long-term roadmap for research computation
  • Mentor engineers working on compute, infrastructure, and AI systems

Benefits

  • Eligibility for annual performance bonus
  • Equity