Senior Staff Site Reliability Engineer AI Infrastructure

AI-infrastructure company producing memory-centric hardware, networking, and software for low-latency generative-AI inference in data centers.

Series C0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/25/2026

Santa Clara, United States
About d-Matrix

d-Matrix develops digital in-memory-compute technology and a full-stack inference platform comprising Corsair accelerators, JetStream networking accelerators, and Aviator software.

View jobs by d-Matrix

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own reliability, automation, and observability across colocation, on-premises, cloud, and customer-facing environments. You will provision and troubleshoot infrastructure, build infrastructure-as-code and operational automation, maintain monitoring, participate in on-call response, and document incident root causes and runbooks.

Requirements

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience
  • 7+ years of SRE, infrastructure engineering, or systems administration experience
  • Linux systems knowledge
  • Hands-on colocation or on-premises server infrastructure experience
  • Terraform or Ansible experience
  • Kubernetes operational experience
  • Observability tooling experience
  • Python or Bash scripting
  • Incident response and root-cause analysis experience

Responsibilities

  • Own reliability and availability across server fleets, lab clusters, cloud environments, and platform services
  • Provision servers and configure operating systems, networking, storage, and Kubernetes infrastructure
  • Lead capacity planning and hardware lifecycle management
  • Track cloud spend for FinOps and workload-placement decisions
  • Drive provisioning and operational changes through Terraform or Ansible
  • Build automation for host lifecycle management, fleet health, remediation, self-service, and networking
  • Design and maintain monitoring, alerting, and SLIs
  • Participate in on-call incident triage and resolution
  • Produce root-cause analyses for P0 and P1 incidents
  • Support platform services and document operational runbooks

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k)
  • Equity
  • Bonus