Manager Production Site Reliability Engineering
Vultr is an active independent cloud-infrastructure company providing global compute, GPU, bare-metal, storage, Kubernetes, and serverless AI inference services.
Maintainer signals as of 9/25/2026
About Vultr
Vultr provides globally available cloud infrastructure for developers, enterprises, and AI innovators, including Cloud Compute, Cloud GPU, Bare Metal, Cloud Storage, managed Kubernetes, and serverless inference.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead production site reliability operations for the production control plane, including availability, performance, incident response, monitoring, configuration management, capacity planning, and infrastructure re-architecture. You will hire and develop SREs and database engineers, establish operational practices, and ensure services are production-ready, observable, and resilient.
Requirements
- 10+ years of professional experience in site reliability engineering or infrastructure operations
- At least 2 years in a team lead or management role
- Experience operating production web stacks at scale
- Linux systems knowledge, including networking, systemd, package management, firewall configuration, and performance tuning
- Experience building and operating monitoring and alerting systems
- Experience with configuration management at scale, such as Puppet, Ansible, or Chef
- Written and verbal communication skills for runbooks, postmortems, incident communications, and cross-functional coordination
- Experience hiring and building engineering teams
- Experience with database replication, backup strategies, and failure modes
- Experience leading production migrations or major infrastructure transitions
- Experience with Harvester or similar HCI platforms for production workloads on VMs
- Experience with Redis cluster architecture, failover, and performance tuning
- Experience with HAProxy, keepalived, and load-balancing strategies
- Experience with PCI compliance environments
- Experience writing PHP
- Experience with Cloudflare or similar CDN or edge platforms and DNS cutovers
Responsibilities
- Build, hire, mentor, and lead a Production Site Reliability Engineering team
- Own the availability, performance, and operability of the production web stack
- Lead incident response, postmortems, runbook improvements, and alerting improvements
- Manage configuration management for consistent and auditable production infrastructure
- Partner with engineering teams to ensure services are observable, deployable, and documented before release
- Drive control plane re-architecture, parallel runs, cutover, and production-readiness validation
- Set the operational roadmap for capacity planning, scaling, disaster recovery, and resilient architecture
- Establish and maintain monitoring, alerting, and observability across the production stack
- Manage production database and caching environments, including replication, tuning, backups, and failover testing
- Foster operational excellence through SLOs, operational reviews, blameless postmortems, and deployment-process improvements
Benefits
- 100% company-paid medical, dental, and vision insurance premiums
- 401(k) matching 100% up to 4% with immediate vesting
- $2,500 annual professional development reimbursement
- 11 holidays, paid time off accrual, and PTO rollover
- Increased PTO at 3- and 10-year anniversaries
- One-month paid sabbatical every five years
- Annual anniversary bonus
- Remote office setup stipend
- Monthly internet reimbursement up to $75
- Monthly gym membership reimbursement up to $50
- Company-paid Wellable subscription
