Decision Engineer Compute Operations

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build production systems for machine telemetry, health checks, incident response, repair workflows, hardware qualification, and facility maintenance. You will work alongside production engineers and facility operators, own asset models and migrations, and turn operational procedures and records into auditable structured data.

Requirements

  • Go
  • Python
  • TypeScript
  • LLM API
  • MCP server
  • Agentic framework
  • Claude Code
  • Cursor
  • On-call
  • Production engineering

Responsibilities

  • Build fleet telemetry, health-check, alerting, and incident systems
  • Automate repair and RMA workflows from failure detection through return to service
  • Build rack-level hardware qualification and burn-in workflows
  • Own facility asset models, maintenance-system rollout, and inventory migration
  • Turn runbooks, training records, and technician qualifications into auditable procedures

Benefits

  • Equity
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Retirement plan
  • Generous PTO policy
Decision Engineer Compute Operations at Fluidstack | JobStash