Staff Backline Engineer ML AI
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will serve as a senior escalation point for complex ML and AI issues. You will investigate and reproduce customer problems, identify root causes, troubleshoot training and inference failures, and work with Engineering and Product to drive resolutions. You will improve diagnostics, documentation, tooling, automation, and support capabilities while mentoring other engineers.
Requirements
- Troubleshooting experience with distributed ML and AI systems
- Python
- PyTorch, TensorFlow, or Scikit-Learn
- Databricks ML and AI technologies including MLflow, Model Serving, Feature Engineering, Spark MLlib, and model lifecycle management
- Apache Spark knowledge including DataFrames, query execution, distributed computing, memory management, shuffles, and performance optimisation
- Training and inference performance troubleshooting
- ML deployment and infrastructure experience including Kubernetes, cloud ML platforms, CI/CD, model monitoring, and production ML systems
- Ability to analyze code, logs, stack traces, metrics, traces, execution plans, and system behavior
- Technical communication skills
Responsibilities
- Serve as a senior escalation point for complex ML and AI issues
- Investigate logs, traces, metrics, profiling, configuration, source code, and customer workloads
- Reproduce customer issues through experimentation, Python and Spark development, and performance analysis
- Troubleshoot model training, inference, resource, distributed execution, and deployment failures
- Partner with Engineering and Product to resolve issues and improve products
- Create improved diagnostics, documentation, tooling, automation, and Claude skill capabilities
- Mentor engineers and improve Support troubleshooting capabilities
- Contribute as a technical subject matter expert to cross-functional initiatives
