Staff Software Engineer - GenAI Performance and Kernel

10 hours agoLeadSalary: 191K - 233KSan Francisco, USAAiJobs by Databricks

Databricks is a data and AI platform that lets organizations build analytics, AI agents, and applications on a unified, governed lakehouse.

160 Spear Street, Suite 1300, San Francisco, CA 94105, United States
About Databricks

Data engineers, analysts, and AI teams use Databricks to process large datasets, build reliable pipelines, and train models on a single governed platform. Users can run SQL analytics, serve ML predictions in real time, and deploy AI agents grounded in enterprise data. Its open lakehouse architecture provides consistent security and governance across analytical and operational workloads.

View jobs by Databricks

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the design, implementation, optimization, and correctness of GPU kernels for the GenAI inference stack. You will lead kernel performance investigations, develop profiling and verification tooling, integrate optimizations into ML systems, influence system architecture, mentor engineers, and monitor production impact.

Requirements

  • Deep hands-on experience writing and tuning compute kernels for ML workloads
  • Knowledge of GPU and accelerator architecture
  • Experience with optimization techniques including tiling, vectorization, fusion, and auto-tuning
  • Familiarity with ML kernel libraries or open kernels
  • Debugging and profiling skills
  • Experience with numerical stability, mixed precision, quantization, and error propagation
  • Experience integrating optimized kernels into ML inference systems
  • Experience building high-performance GPU-accelerated products
  • Communication and leadership skills
  • Track record of shipping performance-critical production software

Responsibilities

  • Lead the design, implementation, benchmarking, and maintenance of optimized compute kernels
  • Drive kernel-level performance improvements
  • Integrate kernel optimizations with higher-level ML systems
  • Build profiling, instrumentation, and verification tooling
  • Lead performance investigations and root-cause analysis for inference bottlenecks
  • Establish reusable and portable kernel abstractions and frameworks
  • Influence system architecture decisions
  • Mentor engineers and provide code reviews
  • Collaborate to deploy kernel optimizations to production and monitor their impact

Benefits

  • Annual performance bonus eligibility
  • Equity eligibility