Member of Technical Staff Cloud Infrastructure

Fireworks AI operates an AI platform for production inference and training of open-source models.

San Mateo, United States
About Fireworks AI

Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.

View jobs by Fireworks AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will architect and build backend infrastructure for distributed training, inference, and data processing. You will design core services, lead technical discussions, optimize compute, storage, and networking, improve reliability and observability, and deploy systems that support large-scale machine learning workloads.

Requirements

  • 5+ years of backend infrastructure experience in cloud environments
  • Experience with ML infrastructure and tooling
  • Strong Python or C++ development skills
  • Knowledge of distributed systems scheduling, orchestration, storage, networking, and compute optimization
  • Experience with AWS, GCP, or Azure
  • Experience with PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes

Responsibilities

  • Architect and build backend infrastructure for distributed training, inference, and data processing
  • Lead technical design discussions and mentor engineers
  • Design backend services including job schedulers, resource managers, autoscalers, and model serving layers
  • Optimize compute costs, storage lifecycles, and network performance
  • Integrate cloud-native and open-source infrastructure technologies
  • Own systems from design through deployment and observability
  • Ensure availability, fault tolerance, disaster recovery, scalability, and performance
  • Develop monitoring, alerting, logging, and tracing solutions

Benefits

  • Equity
Member of Technical Staff Cloud Infrastructure at Fireworks AI | JobStash