Software Engineer Infrastructure

fal is an active generative-media AI platform for developers, providing optimized model APIs, serverless deployment, and GPU compute.

San Francisco, United States
About fal

Founded in 2021 by Burkay Gur and Gorkem Yurtseven, fal provides infrastructure for production generative-media applications, including image, video, audio, 3D, and multimodal models.

View jobs by fal

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build fleet-management software and tooling for server provisioning, health monitoring, diagnostics, recovery, alerting, security, storage, and Linux performance tuning. You will also work with partners to resolve technical issues that automation cannot address.

Requirements

  • Python
  • Linux
  • Ansible
  • Terraform
  • cloud-init
  • LVM
  • RAID
  • NVMe
  • NFS
  • Lustre
  • GPFS
  • GPU
  • NVIDIA
  • CUDA
  • SELinux
  • AppArmor
  • SSH
  • SOC 2
  • ISO 27001

Responsibilities

  • Build and maintain a Python fleet-tracking system
  • Build server-management tooling for provisioning, health checks, diagnostics, recovery, and alerting
  • Create metrics, dashboards, and alerts for hardware health
  • Automate tooling, alerting, and recovery with AI
  • Implement OS-level security and compliance automation
  • Manage storage systems for model weights, checkpoints, and scratch workloads
  • Tune Linux systems and GPU driver stacks for AI workloads
  • Develop automated error-detection and recovery processes
  • Resolve technical issues with partners

Benefits

  • Equity
  • Relocation assistance to San Francisco
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Regular team events and offsites
Software Engineer Infrastructure at fal | JobStash