Site Reliability Engineer for AI Accelerator Infrastructure
d-MatrixVisit d-Matrix website
AI-infrastructure company producing memory-centric hardware, networking, and software for low-latency generative-AI inference in data centers.
Santa Clara, United States
Funding history
About d-Matrix
d-Matrix develops digital in-memory-compute technology and a full-stack inference platform comprising Corsair accelerators, JetStream networking accelerators, and Aviator software.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and automate infrastructure across colocation, on-premises labs, cloud platforms, and customer-facing services. You will provision systems, manage networking and storage, build infrastructure as code, maintain observability, respond to incidents, and document operational procedures.
Requirements
- 5+ years of SRE, infrastructure engineering, or systems administration experience
- Linux systems knowledge
- Colocation or on-premises server infrastructure experience
- Terraform or Ansible experience
- Kubernetes operational experience
- Prometheus and Grafana or Datadog experience
- Python or Bash scripting experience
- Incident response and root-cause analysis experience
Responsibilities
- Own reliability and availability for assigned infrastructure domains
- Provision servers and configure operating systems, networks, storage, and hardware
- Operate high-speed interconnect environments
- Conduct capacity planning and hardware lifecycle management
- Maintain Terraform and Ansible configurations
- Build automation, remediation workflows, and self-service tooling
- Design monitoring dashboards, alerts, and SLIs
- Participate in on-call incident response and produce root-cause analyses
- Support customer-facing platform services
- Maintain runbooks, diagrams, and troubleshooting guides
