Software Engineer Site Reliability
fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.
Funding history
About fal
Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate Kubernetes infrastructure, build CI/CD and deployment systems, automate production issue analysis and resolution, and create dashboards and alerts. You will define SLOs, improve incident response, manage networking and service-mesh configurations, and drive reliability through automation, runbooks, and chaos engineering.
Requirements
- 5+ years managing critical production systems and software development workflows
- Kubernetes
- Terraform
- Ansible
- Linux networking
- Container networking
- CNI plugins
- VXLAN
- BGP
- DNS
- CI/CD
- GitOps
- FluxCD or ArgoCD
- Python
- Go or Bash
- Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, or Datadog
- Technical decision-making
Responsibilities
- Operate Kubernetes infrastructure
- Build and maintain CI/CD pipelines and deployment infrastructure
- Automate production issue analysis and resolution with AI
- Build dashboards, alerting, and anomaly detection
- Define SLOs and incident response processes
- Manage networking, load balancing, and service mesh configurations
- Drive reliability improvements through automation, runbooks, and chaos engineering
Benefits
- Equity
- Health insurance
- Dental insurance
- Vision insurance
- Regular team events and offsites
