Member of Technical Staff AI Training Infrastructure

Fireworks AI operates an AI platform for production inference and training of open-source models.

San Mateo, United States
About Fireworks AI

Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.

View jobs by Fireworks AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design, build, and optimize infrastructure for large-scale model training. You will develop distributed training pipelines, optimize workloads across GPUs and data centers, manage training data storage, automate provisioning and orchestration, implement observability tools, and troubleshoot performance issues.

Requirements

  • 3+ years of experience with distributed systems and ML infrastructure
  • Experience with PyTorch
  • Proficiency with AWS, GCP, or Azure
  • Experience with Kubernetes and Docker
  • Knowledge of data parallelism, model parallelism, and FSDP

Responsibilities

  • Design and implement scalable infrastructure for large-scale model training
  • Develop and maintain distributed training pipelines for LLMs and multimodal models
  • Optimize training performance across GPUs, nodes, and data centers
  • Implement monitoring, logging, and debugging tools
  • Architect and maintain storage for training datasets
  • Automate infrastructure provisioning, scaling, and orchestration
  • Analyze and improve training-system efficiency, scalability, and cost-effectiveness
  • Troubleshoot distributed training performance issues

Benefits

  • Equity
Member of Technical Staff AI Training Infrastructure at Fireworks AI | JobStash