Director Site Reliability Engineer AI Infrastructure
AI-infrastructure company producing memory-centric hardware, networking, and software for low-latency generative-AI inference in data centers.
Maintainer signals as of 9/25/2026
Funding history
About d-Matrix
d-Matrix develops digital in-memory-compute technology and a full-stack inference platform comprising Corsair accelerators, JetStream networking accelerators, and Aviator software.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will build and lead the Site Reliability Engineering function while remaining hands-on with critical infrastructure. You will define reliability practices, lead incident response, establish observability and automation, manage capacity and costs, and guide storage modernization across cloud, on-premises, colocation, and customer-facing services.
Requirements
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
- 15+ years of experience in SRE, infrastructure engineering, or production engineering.
- 5+ years of experience leading SRE or infrastructure engineering teams.
- Experience building or significantly rebuilding an SRE or infrastructure function.
- Linux systems expertise, including bare-metal operations, enterprise shared storage, and hybrid-cloud storage integration.
- Experience operating colocation and on-premises hardware at scale.
- Production-scale Infrastructure as Code experience with Terraform and Ansible.
- Kubernetes cluster operations experience.
- Observability stack experience with Prometheus, Grafana, Datadog, Splunk, or equivalent.
- Python or Go scripting experience for production services and infrastructure automation.
- Executive communication skills.
Responsibilities
- Build and lead the SRE function and hire, develop, and retain SRE engineers.
- Direct the Data Center and Lab Technician team and set operational priorities and standards.
- Own 24×7 reliability across colocation, on-premises clusters, cloud environments, and customer-facing platform services.
- Define SLIs, SLOs, error budgets, on-call rotations, and incident and root-cause-analysis processes.
- Own the observability stack and design metrics, traces, logs, alerts, and SLO visibility.
- Drive infrastructure-as-code automation and build self-healing infrastructure.
- Own FinOps and capacity planning across cloud, colocation, and on-premises infrastructure.
- Lead the migration to an enterprise-grade shared storage platform, including architecture, vendor selection, and disaster-recovery design.
Benefits
- Equity
- Bonus
- Medical insurance
- Dental insurance
- Vision insurance
- 401(k)
- Incentive opportunities
