Director, Site Reliability Engineering
Stellar is an open-source blockchain network designed to facilitate fast, low-cost cross-border payments and financial transactions globally.
Funding
Investors
Projects
About Stellar
Stellar is an open-source blockchain network focused on improving global financial access by enabling fast, secure, and low-cost cross-border transactions. It connects financial institutions, payment systems, and individuals to transfer digital and fiat assets. With a mission to promote financial inclusion, Stellar is designed to serve both developed and underserved markets, offering interoperability between traditional and blockchain-based financial systems.
Skills
About the Role
You will lead, coach, and develop a distributed SRE team, setting the vision, operating model, and priorities for the function. You will own and improve core engineering infrastructure services, including cloud foundations, Kubernetes, CI/CD, observability, secrets management, and infrastructure automation. You will define and roll out a Service Ownership & Maturity Framework across engineering, help teams strengthen their operational practices, and make reliability and infrastructure health measurable through trusted metrics. You will improve deployment automation, resilience, and disaster recovery readiness, and mature incident response, escalation, and on-call practices across a geographically distributed team. You will also partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT, and evaluate AI-assisted and agentic workflows to improve infrastructure operations and developer workflows. You will report directly to the CTO.
Requirements
- 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles
- 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers
- Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk
- 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments
- 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices
- Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability
- A pragmatic approach to tooling: understanding when to build, buy, adapt, simplify, or retire systems
- Ability to operate effectively in a small or mid-sized engineering organization where influence comes from credibility, judgment, and outcomes
- Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders
Responsibilities
- Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures
- Define and roll out a Service Ownership & Maturity Framework across engineering
- Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation
- Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices
- Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence
- Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk
- Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team
- Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability
- Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT
- Pragmatically evaluate AI-assisted and agentic workflows to improve infrastructure operations, service ownership, developer workflows, or toil reduction
Benefits
- Competitive health, dental & vision coverage with most plans covered at 100% for the employee + any dependents
- Flexible time off + 15 company holidays including a company-wide holiday break
- Generous paid parental leave for all parents, plus paid pregnancy disability leave for birthing parents
- Gym reimbursement ($80 per month)
- Life & AD&D (up to $50K)
- Short & Long term disability
- 401K with 4% match
- Health & Dependent Care FSA Accounts
- Commuter benefits with $250/month employer contribution
- Health Savings Account (HSA) with monthly employer contribution
- Family building benefits through Kindbody
- Wellbeing benefits (One Medical, Rightway, Headspace)
- L&D budget of $1,500/year
- Daily lunch and snacks in office
- Company retreats
- Lumen-denominated grants
