Customer Reliability Engineer

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own reliability for assigned customer workloads, including clusters, SLAs, and escalations. You will debug issues across hardware, fabric, and scheduling layers, communicate incidents clearly, and turn recurring pain points into engineering fixes.

Requirements

  • Experience supporting large-scale HPC cloud or AI compute customers
  • Distributed systems debugging
  • Incident communication
  • Root-cause remediation
  • GPU training workload knowledge
  • InfiniBand
  • RoCE
  • Slurm
  • Kubernetes
  • NCCL

Responsibilities

  • Own reliability for named customer workloads
  • Debug workload degradation across the full stack
  • Run customer-facing incident communications
  • Turn recurring customer issues into engineering fixes

Benefits

  • Equity
Customer Reliability Engineer at Fluidstack | JobStash