Staff Software Engineer - GenAI Inference
Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.
Funding history
Investors
About Databricks
Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead the architecture, design, and implementation of an inference engine for large-scale LLMs. You will optimize latency, throughput, memory efficiency, and hardware utilization; build scalable scheduling and memory mechanisms; ensure reliable inference pipelines; and collaborate across research, platform, infrastructure, and security functions.
Requirements
- Software engineering experience in performance-critical systems
- Experience owning complex system components and architectural decisions
- Understanding of ML inference internals
- CUDA and GPU programming experience
- Experience with cuBLAS, cuDNN, and NCCL
- Distributed systems design experience
- Experience solving performance bottlenecks across kernels, memory, networking, and scheduling
- Experience building instrumentation, tracing, and profiling tools for ML models
- Ability to translate ML research ideas into production systems
- Communication and leadership skills
Responsibilities
- Own the architecture, design, and implementation of the inference engine
- Collaborate on an LLM model-serving stack
- Bring new model architectures and features into the engine
- Optimize latency, throughput, memory efficiency, and hardware utilization
- Build instrumentation, profiling, and tracing tooling
- Architect routing, batching, scheduling, memory management, and dynamic loading mechanisms
- Ensure reliability, reproducibility, and fault tolerance in inference pipelines
- Integrate with distributed inference infrastructure
- Drive cross-functional collaboration
- Represent the team through benchmarks, whitepapers, and open-source contributions
Benefits
- Annual performance bonus eligibility
- Equity eligibility
