Software Engineer Site Reliability

fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.

San Francisco, United States
About fal

Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.

View jobs by fal

Skills

About the Role

You will operate Kubernetes infrastructure, build CI/CD and deployment systems, automate production issue analysis and resolution, and create dashboards and alerts. You will define SLOs, improve incident response, manage networking and service-mesh configurations, and drive reliability through automation, runbooks, and chaos engineering.

Requirements

  • 5+ years managing critical production systems and software development workflows
  • Kubernetes
  • Terraform
  • Ansible
  • Linux networking
  • Container networking
  • CNI plugins
  • VXLAN
  • BGP
  • DNS
  • CI/CD
  • GitOps
  • FluxCD or ArgoCD
  • Python
  • Go or Bash
  • Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, or Datadog
  • Technical decision-making

Responsibilities

  • Operate Kubernetes infrastructure
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Automate production issue analysis and resolution with AI
  • Build dashboards, alerting, and anomaly detection
  • Define SLOs and incident response processes
  • Manage networking, load balancing, and service mesh configurations
  • Drive reliability improvements through automation, runbooks, and chaos engineering

Benefits

  • Regular team events and offsites
Software Engineer Site Reliability at fal | JobStash