Principal Staff Site Reliability Engineer
DigitalBridge Group, Inc. is a global alternative asset manager focused on investing in, owning, operating, and building digital infrastructure. It serves limited partners, shareholders, corporate borrowers, entrepreneurs, and technology and telecommunications companies.
Funding history
Investors
About DigitalBridge Group, Inc.
DigitalBridge manages investment strategies spanning digital infrastructure equity, core-plus assets, private credit, liquid public-market strategies, and software-defined infrastructure ventures. Its focus includes data centers, cell towers, fiber networks, small cells, edge infrastructure, and related technologies. The firm invests in and actively manages portfolio companies while providing financing and capital solutions to institutional investors, companies, and entrepreneurs.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own reliability for critical services and set technical direction for hybrid cloud infrastructure. You will build infrastructure automation, improve observability, lead AI-assisted incident-response capabilities, strengthen production security, participate in on-call, and mentor senior engineers.
Requirements
- 10+ years building and operating production infrastructure at scale
- Hybrid cloud and on-premises infrastructure experience
- AWS and Azure expertise
- Linux expertise
- Terraform
- Ansible
- CI/CD
- Python
- Go
- Kubernetes
- Service mesh
- Container security
- Datadog
- Prometheus
- Grafana
- OpenTelemetry
- ELK
- Splunk
- Incident command
Responsibilities
- Own end-to-end reliability for business-critical services, including SLOs, error budgets, capacity planning, disaster recovery, and incident command
- Design and evolve multi-cloud and on-premises infrastructure across AWS, Azure, and colocated environments
- Build and maintain Terraform, Ansible, and CI/CD infrastructure
- Advance observability through metrics, logs, traces, and profiling
- Lead AI-assisted SRE capabilities for triage, incident summaries, runbooks, and root-cause analysis
- Harden production systems with security, data platform, and application teams
- Participate in on-call, run blameless postmortems, and drive systemic fixes
- Mentor senior engineers and lead infrastructure code and production-readiness reviews
