Software Engineer GPU Infrastructure

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build fleet-health metrics, alerting, repair automation, GPU qualification systems, and low-level hardware tooling. You will manage fault detection, triage, parts management, return to service, firmware telemetry, log collection, and compute-fleet reliability at scale.

Requirements

  • Experience shipping production automation used by other teams
  • Experience with LLM APIs, MCP servers, agentic frameworks, and AI coding tools
  • Ability to reason about hardware failure modes at firmware and silicon levels
  • Ability to run incidents, write postmortems, and fix systemic causes

Responsibilities

  • Build metrics pipelines, alerting, and a unified GPU fleet-health view
  • Build repair automation from fault detection through triage, parts management, and return to service
  • Design and expand GPU burn-in, performance-baselining, and qualification systems
  • Own Redfish and BMC tooling, firmware telemetry, and fleet-scale log collection
  • Own compute-fleet reliability, scalability, and operations

Benefits

  • Equity
  • Retirement or pension plan
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Generous PTO policy