Machine Learning Engineer Reliability

fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.

San Francisco, United States
About fal

Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.

View jobs by fal

Skills

About the Role

You will own the reliability, security, and safety of production generative media model APIs. You will build observability and safe deployment systems, lead incident response, improve GPU capacity management, and help incorporate reliability requirements into new model onboarding.

Requirements

  • Production ML or high-scale API operations
  • Diffusion models
  • Distributed systems
  • Networking
  • Observability
  • Incident management
  • Generative models
  • Security
  • ML safety
  • Python
  • PyTorch
  • Diffusers
  • Kubernetes

Responsibilities

  • Own availability, latency, and throughput SLOs for generative media model APIs
  • Build monitoring, alerting, and observability for ML-specific failures and model regressions
  • Harden deployments with canary releases, shadow testing, automated rollbacks, and validation gates
  • Drive secure model serving, abuse detection, rate limiting, and adversarial-use protection
  • Operationalize content moderation, safety classifiers, and inference-time guardrails
  • Lead incident response, postmortems, and recurrence-prevention work
  • Improve capacity planning, autoscaling, and GPU fleet efficiency
  • Incorporate reliability, security, and safety requirements into model onboarding

Benefits

  • Equity
  • Regular team events and offsites
Machine Learning Engineer Reliability at fal | JobStash