Member of Technical Staff Reliability Engineering
Fireworks AIVisit Fireworks AI website
Fireworks AI operates an AI platform for production inference and training of open-source models.
Fireworks AI on X (Twitter)Fireworks AI on DiscordFireworks AI on GitHubFireworks AI on Documentation
San Mateo, United States
Funding history
About Fireworks AI
Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define reliability standards, build observability and incident-management tooling, investigate cross-system failures, automate operational work, and partner with infrastructure, inference, training, performance, and product functions to improve customer-facing reliability.
Requirements
- Have 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals
- Have 5+ years writing production tools and systems in Python, Go, C++, or Rust
- Operate and debug Kubernetes, Terraform, and Docker in high-throughput production
- Understand distributed systems, microservices, or multi-region systems
- Understand fault-tolerant design, SLO and SLA management, automated failover, and high availability
- Use TCP/IP, HTTP, and gRPC
- Have observability experience with Prometheus, Grafana, and OpenTelemetry
Responsibilities
- Define SLOs, error budgets, and production-readiness criteria
- Own logging, telemetry, alerting, failure testing, load testing, and self-healing automation
- Identify and resolve customer-impacting reliability failures
- Find cross-system failure modes and drive fixes
- Coordinate production incidents and run blameless postmortems
- Automate repetitive operational work
- Partner on capacity, multi-region risk, serving and training failures, rollouts, and customer-facing reliability
Benefits
- Equity
