Member of Technical Staff, Kernels

49 minutes agoSalary: 200K - 350KBay AreaOnsiteFull TimeAiJobs by Inception

AI research and product company building diffusion-based language models for production applications.

Palo Alto, United States

Funding history

About Inception

Inception develops and deploys the Mercury family of diffusion LLMs, which generate and refine output in parallel to target lower-latency, lower-cost production AI workloads.

View jobs by Inception

Skills

About the Role

You will implement optimized GPU kernels and compute primitives for core language model operations. You will reduce memory bottlenecks, support low-precision arithmetic, and improve compute reliability, reproducibility, scalability, and utilization.

Requirements

  • CUDA
  • CuTe
  • Triton
  • GPU programming
  • PyTorch
  • TensorFlow
  • Performance optimization
  • Profiling
  • FP8
  • INT8
  • Block floating point
  • XLA
  • TVM
  • Distributed training
  • Data parallelism
  • Model parallelism
  • Pipeline parallelism
  • Python
  • C++
  • Rust
  • Go
  • Docker
  • Kubernetes
  • CI/CD

Responsibilities

  • Design and implement optimized machine learning kernels for GPU architectures
  • Develop compute primitives that reduce memory bandwidth bottlenecks
  • Improve kernel efficiency for core language model operations
  • Ensure reproducibility and precision consistency across compute infrastructure
  • Improve compute resource utilization and infrastructure scalability

Benefits

  • Equity
  • Flexible vacation and paid time off
  • Health insurance
  • Dental insurance
  • Vision insurance
  • 401k match
  • Catered meals
  • Commuter subsidies
Member of Technical Staff, Kernels at Inception | JobStash