Software Engineer Infrastructure
fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.
San Francisco, United States
Funding history
About fal
Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build fleet-management software and tooling for server provisioning, health monitoring, diagnostics, recovery, alerting, security, storage, and Linux performance tuning. You will also work with partners to resolve technical issues that automation cannot address.
Requirements
- Python
- Linux
- Ansible
- Terraform
- cloud-init
- LVM
- RAID
- NVMe
- NFS
- Lustre
- GPFS
- GPU
- NVIDIA
- CUDA
- SELinux
- AppArmor
- SSH
- SOC 2
- ISO 27001
Responsibilities
- Build and maintain a Python fleet-tracking system
- Build server-management tooling for provisioning, health checks, diagnostics, recovery, and alerting
- Create metrics, dashboards, and alerts for hardware health
- Automate tooling, alerting, and recovery with AI
- Implement OS-level security and compliance automation
- Manage storage systems for model weights, checkpoints, and scratch workloads
- Tune Linux systems and GPU driver stacks for AI workloads
- Develop automated error-detection and recovery processes
- Resolve technical issues with partners
Benefits
- Equity
- Relocation assistance to San Francisco
- Health insurance
- Dental insurance
- Vision insurance
- Regular team events and offsites
