Software Engineer Compute GPU
1 day agoSalary: 208K - 269KSan Francisco, CA; Austin, TX; New York, NY; Seattle, WAOnsiteFull TimeDevopsJobs by Fluidstack
FluidstackVisit Fluidstack website
AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.
New York City, United States
Funding history
About Fluidstack
Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build metrics, alerting, fleet-health, repair, and qualification systems for production GPUs. You will develop Redfish and BMC tooling, automate detection through return-to-service workflows, operate fleet-scale infrastructure, and lead incident response and systemic remediation.
Requirements
- Experience with firmware- and silicon-level hardware failure modes
- Experience shipping production automation used by other teams
- Fluency with LLM APIs, MCP servers, agentic frameworks, and AI coding tools
- Ability to operate incident response and postmortem processes
Responsibilities
- Build metrics pipelines, alerting, and a unified GPU fleet-health view
- Automate compute failure detection, triage, parts management, and return to service
- Design and expand GPU burn-in, performance-baselining, and NPI qualification systems
- Own Redfish and BMC tooling for telemetry and fleet-scale log collection
- Own reliability, scalability, and operation of the compute fleet
- Run incidents, write postmortems, and fix systemic causes
Benefits
- Equity
- Retirement or pension plan
- Health insurance
- Dental insurance
- Vision insurance
- Generous PTO
