Distributed Systems Engineer

AI infrastructure company that builds and operates large-scale compute and data-center infrastructure for frontier AI workloads.

New York City, United States
About Fluidstack

Fluidstack deploys AI compute infrastructure, including custom data centers and large-scale compute capacity, for AI labs, governments, and enterprises.

View jobs by Fluidstack

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will build and operate observability pipelines, health checks, infrastructure APIs, and the production control plane. You will maintain fleet state as a source of truth, integrate new hardware, operate Kubernetes infrastructure, and address incidents and systemic causes.

Requirements

  • API design and versioning
  • AI tooling
  • LLM API
  • MCP
  • agentic framework
  • Production service engineering
  • Distributed systems engineering
  • Data pipeline engineering
  • Prometheus
  • Thanos
  • VictoriaMetrics
  • Temporal
  • Cadence
  • BMC
  • Redfish
  • Go
  • Python
  • Postgres

Responsibilities

  • Build and operate the observability platform
  • Define and build infrastructure APIs
  • Build the production control plane
  • Own fleet state as a source of truth
  • Integrate new hardware and sites into the platform
  • Run incidents write postmortems and fix systemic causes

Benefits

  • Equity
  • Retirement or pension plan
  • Health insurance
  • Dental insurance
  • Vision insurance
  • Generous PTO policy
Distributed Systems Engineer at Fluidstack | JobStash