Senior Staff Software Engineer Kubernetes Infrastructure
fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.
San Francisco, United States
Funding history
About fal
Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will design, provision, operate, upgrade, recover, and decommission customer compute environments. You will build Kubernetes and Slurm clusters, Linux provisioning workflows, GPU infrastructure, networking, storage, observability, and reusable operational tooling.
Requirements
- Linux
- Kubernetes
- Bare metal
- etcd
- containerd
- CNI
- CSI
- KVM
- QEMU
- libvirt
- VFIO
- NVIDIA GPU
- TCP/IP
- VLAN
- routing
- tcpdump
- Wireshark
- Ansible
- Slurm
- Python
- Go
Responsibilities
- Design and deliver the full lifecycle of customer compute environments
- Automate infrastructure delivery and operations with AI
- Provision dedicated Kubernetes and Slurm clusters
- Build Linux images and automated OS-provisioning workflows
- Operate NVIDIA GPU infrastructure
- Design Kubernetes and data-center networking
- Configure distributed and shared storage
- Build monitoring, alerting, diagnostics, and automated recovery
- Develop reusable tooling, standards, documentation, and runbooks
- Translate workload requirements into infrastructure designs
Benefits
- Equity
- Visa sponsorship
- Relocation assistance to San Francisco
- Health insurance
- Dental insurance
- Vision insurance
- Regular team events and offsites
