Member of Technical Staff Performance Optimization
Fireworks AIVisit Fireworks AI website
Fireworks AI operates an AI platform for production inference and training of open-source models.
Fireworks AI on X (Twitter)Fireworks AI on DiscordFireworks AI on GitHubFireworks AI on Documentation
San Mateo, United States
Funding history
About Fireworks AI
Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will optimize AI workloads across GPU kernels and distributed systems. You will profile bottlenecks, improve latency and resource efficiency, implement CUDA and Triton optimizations, benchmark performance, and scale training and inference across multi-GPU and multi-node environments.
Requirements
- Performance optimization
- High-performance computing
- CUDA
- ROCm
- GPU profiling
- Nsight
- nvprof
- CUPTI
- PyTorch
- Distributed system
- GPU architecture
- Parallel programming
- Compute kernel
Responsibilities
- Optimize system and GPU performance for AI workloads
- Improve latency, throughput, memory usage, and compute efficiency
- Profile and resolve GPU and kernel bottlenecks
- Implement low-level optimizations with CUDA and Triton
- Improve execution speed and resource utilization
- Co-design hardware-efficient model architectures
- Build performance benchmarking and monitoring infrastructure
- Scale inference and training across multi-GPU and multi-node environments
- Evaluate emerging hardware accelerators and runtimes
Benefits
- Equity
