Production Engineer Compute
1 day agoSalary: 208K - 269KSan Francisco, CA; Austin, TX; New York, NY; Seattle, WAOnsiteFull TimeDevopsJobs by Fluidstack
FluidstackVisit Fluidstack website
AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.
New York City, United States
Funding history
About Fluidstack
Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and operate automation for compute fleet health, GPU repair, qualification, telemetry, and reliability. You will develop metrics and alerting, manage Redfish and BMC tooling, respond to incidents, and improve fleet-scale operations.
Requirements
- Experience shipping production automation used by other teams
- Ability to reason about firmware- and silicon-level hardware failure modes
- Experience with LLM APIs, MCP servers, and agentic frameworks
- Experience using AI coding tools
- Ability to run incidents and write postmortems
Responsibilities
- Build metrics pipelines, alerting, and unified compute fleet health views
- Automate failure detection, triage, parts management, and return to service
- Design and expand the GPU qualification platform
- Own Redfish and BMC tooling for telemetry and log collection
- Own compute fleet reliability, scalability, and operations
- Run incidents, write postmortems, and fix systemic causes
Benefits
- Equity
- Retirement or pension plan
- Health insurance
- Dental insurance
- Vision insurance
- Generous PTO policy
