Production Engineer Compute Team Lead

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will lead production engineering for a large GPU fleet. You will define availability objectives, build fleet lifecycle automation, and establish on-call and escalation practices. You will hire and develop engineers while improving reliability and automated remediation.

Requirements

  • Lead SRE or production engineering teams running large fleets.
  • Improve service availability.
  • Ship automated remediation that replaces runbooks.
  • Hire and develop engineers.

Responsibilities

  • Lead the compute production engineering team.
  • Own fleet availability for compute.
  • Define SLOs and build reliability tooling.
  • Build automation for node provisioning, health checks, remediation, and return to service.
  • Set the on-call and escalation model.

Benefits

  • Equity
Production Engineer Compute Team Lead at Fluidstack | JobStash