Vice President Site Reliability Engineering (Data Centers)
Galaxy Digital is a full-service digital asset firm offering institutional OTC trading, lending, derivatives, staking, tokenization, asset management, investment banking, and venture funding.
Projects
About Galaxy Digital
Galaxy Digital is a technology-driven financial services and investment management firm that offers institutions and direct clients a full suite of financial solutions spanning the digital assets ecosystem. The company operates three complementary businesses: Global Markets, Asset Management, and Digital Infrastructure Solutions.
Skills
About the Role
You will lead the design, deployment, and maintenance of automation toolsets and the systems they interact with. You will establish and enforce Infrastructure as Code standards, manage configuration and image pipelines, and implement monitoring and SLIs/SLOs for automation platforms. You will drive automated lifecycle management for physical and virtual assets, develop custom tooling and scripts (Python, Go, PowerShell, Bash), analyze capacity and performance in virtual environments, collaborate with datacenter stakeholders to enable self-service workflows, and provide technical mentorship to SREs.
Requirements
- 6-10 years experience in Infrastructure SRE or DevOps focused on infrastructure automation at scale
- Deep proficiency with Terraform including providers modules and state management
- Deep proficiency with Ansible including roles playbooks and Tower/AWX
- Hands-on experience with image creation tools such as Packer Ansible and SCCM for Windows and Linux
- Experience managing and automating virtual platforms such as VMware vSphere vCenter and ESXi
- Experience with cloud providers such as Azure and AWS
- Scripting skills in Python Go PowerShell and Bash
- Experience with observability tools such as Splunk ELK Prometheus or Grafana
- Understanding of network topology and experience with Juniper or Palo Alto platforms
- Mastery of Git branching strategies PR workflows and CI/CD platforms such as Jenkins GitLab CI or GitHub Actions
- Experience managing Windows Server and Linux performance and troubleshooting
Responsibilities
- Oversee a specialized SRE team focused on automation platforms
- Establish and enforce Infrastructure as Code standards and governance
- Lead configuration management and image pipelines using Ansible and Packer for Windows and Linux
- Manage monitoring and observability of automation platforms and implement SLIs/SLOs
- Drive automated lifecycle management of physical and virtual assets
- Develop custom tooling and scripts in Python Go PowerShell and Bash
- Collaborate with datacenter teams and stakeholders to enable self service workflows
- Analyze capacity and performance to optimize automated deployments
- Mentor and provide technical guidance to SREs
