Customer Reliability Engineer
58 minutes agoSalary: 173K - 224KSan Francisco, CA; Austin, TX; New York, NY; Seattle, WAOnsiteFull TimeDevopsJobs by Fluidstack
FluidstackVisit Fluidstack website
AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.
New York City, United States
Funding history
About Fluidstack
Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own reliability for assigned customer workloads, including clusters, SLAs, and escalations. You will debug issues across hardware, fabric, and scheduling layers, communicate incidents clearly, and turn recurring pain points into engineering fixes.
Requirements
- Experience supporting large-scale HPC cloud or AI compute customers
- Distributed systems debugging
- Incident communication
- Root-cause remediation
- GPU training workload knowledge
- InfiniBand
- RoCE
- Slurm
- Kubernetes
- NCCL
Responsibilities
- Own reliability for named customer workloads
- Debug workload degradation across the full stack
- Run customer-facing incident communications
- Turn recurring customer issues into engineering fixes
Benefits
- Equity
