Member of Technical Staff Performance Optimization

Fireworks AI operates an AI platform for production inference and training of open-source models.

San Mateo, United States
About Fireworks AI

Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.

View jobs by Fireworks AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will optimize AI workloads across GPU kernels and distributed systems. You will profile bottlenecks, improve latency and resource efficiency, implement CUDA and Triton optimizations, benchmark performance, and scale training and inference across multi-GPU and multi-node environments.

Requirements

  • Performance optimization
  • High-performance computing
  • CUDA
  • ROCm
  • GPU profiling
  • Nsight
  • nvprof
  • CUPTI
  • PyTorch
  • Distributed system
  • GPU architecture
  • Parallel programming
  • Compute kernel

Responsibilities

  • Optimize system and GPU performance for AI workloads
  • Improve latency, throughput, memory usage, and compute efficiency
  • Profile and resolve GPU and kernel bottlenecks
  • Implement low-level optimizations with CUDA and Triton
  • Improve execution speed and resource utilization
  • Co-design hardware-efficient model architectures
  • Build performance benchmarking and monitoring infrastructure
  • Scale inference and training across multi-GPU and multi-node environments
  • Evaluate emerging hardware accelerators and runtimes

Benefits

  • Equity
Member of Technical Staff Performance Optimization at Fireworks AI | JobStash