Production Engineer Compute

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate automation for compute fleet health, GPU repair, qualification, telemetry, and reliability. You will develop metrics and alerting, manage Redfish and BMC tooling, respond to incidents, and improve fleet-scale operations.

Requirements

  • Experience shipping production automation used by other teams
  • Ability to reason about firmware- and silicon-level hardware failure modes
  • Experience with LLM APIs, MCP servers, and agentic frameworks
  • Experience using AI coding tools
  • Ability to run incidents and write postmortems

Responsibilities

  • Build metrics pipelines, alerting, and unified compute fleet health views
  • Automate failure detection, triage, parts management, and return to service
  • Design and expand the GPU qualification platform
  • Own Redfish and BMC tooling for telemetry and log collection
  • Own compute fleet reliability, scalability, and operations
  • Run incidents, write postmortems, and fix systemic causes

Benefits

  • Equity
  • Retirement or pension plan
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Generous PTO policy