Member of Technical Staff Cloud Infrastructure
Fireworks AIVisit Fireworks AI website
Fireworks AI operates an AI platform for production inference and training of open-source models.
Fireworks AI on X (Twitter)Fireworks AI on DiscordFireworks AI on GitHubFireworks AI on Documentation
San Mateo, United States
Funding history
About Fireworks AI
Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will architect and build backend infrastructure for distributed training, inference, and data processing. You will design core services, lead technical discussions, optimize compute, storage, and networking, improve reliability and observability, and deploy systems that support large-scale machine learning workloads.
Requirements
- 5+ years of backend infrastructure experience in cloud environments
- Experience with ML infrastructure and tooling
- Strong Python or C++ development skills
- Knowledge of distributed systems scheduling, orchestration, storage, networking, and compute optimization
- Experience with AWS, GCP, or Azure
- Experience with PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes
Responsibilities
- Architect and build backend infrastructure for distributed training, inference, and data processing
- Lead technical design discussions and mentor engineers
- Design backend services including job schedulers, resource managers, autoscalers, and model serving layers
- Optimize compute costs, storage lifecycles, and network performance
- Integrate cloud-native and open-source infrastructure technologies
- Own systems from design through deployment and observability
- Ensure availability, fault tolerance, disaster recovery, scalability, and performance
- Develop monitoring, alerting, logging, and tracing solutions
Benefits
- Equity
