Member of Technical Staff Reliability Engineering

Fireworks AI operates an AI platform for production inference and training of open-source models.

San Mateo, United States
About Fireworks AI

Fireworks AI provides serverless and dedicated model inference, model deployment, and supervised and reinforcement fine-tuning for developers and enterprises building AI applications.

View jobs by Fireworks AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will define reliability standards, build observability and incident-management tooling, investigate cross-system failures, automate operational work, and partner with infrastructure, inference, training, performance, and product functions to improve customer-facing reliability.

Requirements

  • Have 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals
  • Have 5+ years writing production tools and systems in Python, Go, C++, or Rust
  • Operate and debug Kubernetes, Terraform, and Docker in high-throughput production
  • Understand distributed systems, microservices, or multi-region systems
  • Understand fault-tolerant design, SLO and SLA management, automated failover, and high availability
  • Use TCP/IP, HTTP, and gRPC
  • Have observability experience with Prometheus, Grafana, and OpenTelemetry

Responsibilities

  • Define SLOs, error budgets, and production-readiness criteria
  • Own logging, telemetry, alerting, failure testing, load testing, and self-healing automation
  • Identify and resolve customer-impacting reliability failures
  • Find cross-system failure modes and drive fixes
  • Coordinate production incidents and run blameless postmortems
  • Automate repetitive operational work
  • Partner on capacity, multi-region risk, serving and training failures, rollouts, and customer-facing reliability

Benefits

  • Equity
Member of Technical Staff Reliability Engineering at Fireworks AI | JobStash