Director, Site Reliability Engineering
Stellar is an open-source blockchain network designed to facilitate fast, low-cost cross-border payments and financial transactions globally.
Maintainer signals as of 9/2/2026
Funding history
Investors
Projects
About Stellar Development Foundation
Stellar is an open-source blockchain network focused on improving global financial access by enabling fast, secure, and low-cost cross-border transactions. It connects financial institutions, payment systems, and individuals to transfer digital and fiat assets. With a mission to promote financial inclusion, Stellar is designed to serve both developed and underserved markets, offering interoperability between traditional and blockchain-based financial systems.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
This senior engineering leadership role reports to the CTO and sets the vision, operating model, and culture for SRE. The role owns cloud foundations, Kubernetes, CI/CD, observability, infrastructure automation, service ownership standards, reliability practices, and operational maturity across engineering.
Requirements
- 10+ years of experience in SRE, infrastructure, platform, cloud infrastructure, production operations, or related engineering roles.
- 5+ years leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
- Experience defining team charters, operating models, roadmaps, success measures, and engineering practices.
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability, automation, and operational risk.
- 3+ years with AWS, GCP, or similar cloud infrastructure.
- 3+ years with Kubernetes, container orchestration, infrastructure-as-code, CI/CD, and deployment safety.
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
- Experience improving service ownership, operational readiness, and production accountability.
- Pragmatic tooling judgment and ability to operate effectively in a small or mid-sized engineering organization.
- Clear executive communication and ability to partner with a CTO and senior engineering leaders.
Responsibilities
- Lead, coach, and develop a distributed SRE team.
- Define and roll out a Service Ownership & Maturity Framework across engineering.
- Own cloud foundations, Kubernetes, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
- Improve service ownership through standards, dashboards, runbooks, alerting, escalation paths, and deployment practices.
- Measure reliability, operational maturity, infrastructure health, and developer productivity.
- Improve deployment automation, resilience, self-healing, disaster recovery, and service reliability.
- Mature incident response, escalation, postmortems, and on-call health.
- Build paved paths and self-service infrastructure to reduce toil and cognitive load.
- Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT.
- Evaluate AI-assisted and agentic workflows for infrastructure operations and developer productivity.
Benefits
- Competitive health, dental, and vision coverage.
- Flexible time off and 15 company holidays.
- Paid parental leave and pregnancy disability leave.
- $80 monthly gym reimbursement.
- Life and AD&D insurance up to $50,000.
- Short- and long-term disability coverage.
- 401(k) with 4% match.
- Health and dependent care FSA accounts.
- $250 monthly commuter contribution.
- Health savings account contributions.
- Family building benefits through Kindbody.
- Wellbeing benefits including One Medical, Rightway, and Headspace.
- $1,500 annual learning and development budget.
- Daily lunch and snacks in office.
- Company retreats.
- Lumen-denominated grants.
