Software Engineer, Platform
fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.
San Francisco, United States
Funding history
About fal
Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.
Skills
About the Role
You will build Python systems and tooling that manage the full lifecycle of GPU servers, from procurement and provisioning through health monitoring, diagnostics, recovery, and alerting. You will secure and tune Linux infrastructure, manage storage for AI workloads, create fleet health dashboards, and work with partners to resolve technical issues.
Requirements
- 3+ years of experience managing bare-metal and cloud server fleets at scale
- Production Python software engineering skills
- Deep Linux systems knowledge
- Experience with configuration management and infrastructure as code, including Ansible, Terraform, and cloud-init
- Understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O tuning
- Familiarity with hardware diagnostics and failure modes
- Experience building internal infrastructure tools or dashboards
- Communication skills and ability to drive technical decisions across teams
- Familiarity with network configuration and diagnostics
- Experience with NVIDIA GPU infrastructure
- Experience with AMD GPUs
- Experience with bare-metal and VM provisioning
- Experience with cloud-provider compliance frameworks
Responsibilities
- Build and maintain a Python fleet tracking system
- Build server management tooling for provisioning, health checks, GPU diagnostics, recovery, and alerting
- Create hardware-health metrics, dashboards, and alerts
- Automate alerting and recovery using AI
- Implement OS-level security controls
- Manage and optimize distributed and local storage systems
- Tune Linux systems and GPU driver stacks for AI workloads
- Develop automated error detection and recovery processes
- Work with partners to resolve technical issues
Benefits
- Regular team events and offsites
