Search...

Director of Site Reliability Engineering

Stellar logo
Stellar

Stellar is an open-source blockchain network designed to facilitate fast, low-cost cross-border payments and financial transactions globally.

P.O. Box 77105, San Francisco, CA 94107, United States‍, United States

Investors

About Stellar

Stellar is an open-source blockchain network focused on improving global financial access by enabling fast, secure, and low-cost cross-border transactions. It connects financial institutions, payment systems, and individuals to transfer digital and fiat assets. With a mission to promote financial inclusion, Stellar is designed to serve both developed and underserved markets, offering interoperability between traditional and blockchain-based financial systems.

View jobs by Stellar

Skills

About the Role

You will lead and develop a distributed Site Reliability Engineering function, define its operating model, and establish practical standards for service ownership. You will own and improve shared infrastructure, strengthen observability and incident practices, automate deployments and recovery, and enable engineering teams to operate reliable services with confidence.

Requirements

  • 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or related engineering roles.
  • 5+ years of experience leading, managing, or developing infrastructure, SRE, platform, or reliability engineers.
  • Experience defining team charters, operating models, roadmaps, success measures, and engineering practices.
  • Deep technical judgment in cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
  • 3+ years of experience with modern cloud infrastructure such as AWS or GCP.
  • 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety.
  • Experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
  • Experience improving service ownership, operational readiness, and production accountability for product or application engineering teams.
  • Executive communication skills and ability to partner with a CTO and senior engineering leaders.

Responsibilities

  • Lead, coach, and develop the distributed SRE team.
  • Set the SRE vision, charter, operating model, priorities, and success measures.
  • Define and roll out a Service Ownership and Maturity Framework.
  • Own and improve cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
  • Enable engineering teams through standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
  • Measure reliability, operational maturity, infrastructure health, and developer productivity.
  • Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability.
  • Mature incident response, escalation, postmortems, and on-call health.
  • Build paved paths and self-service infrastructure to reduce toil and cognitive load.
  • Partner with Security, Compliance, Legal, Finance, Procurement, and Corporate IT on infrastructure and controls.
  • Evaluate AI-assisted and agentic workflows for infrastructure operations and developer productivity.

Benefits

  • Health, dental, and vision coverage with most plans fully covered for employees and dependents.
  • Flexible time off and 15 company holidays, including a company-wide holiday break.
  • Paid parental leave and paid pregnancy disability leave.
  • Gym reimbursement of $80 per month.
  • Life and accidental death and dismemberment insurance.
  • Short-term and long-term disability coverage.
  • 401K with 4% match.
  • Health and Dependent Care FSA accounts.
  • Commuter benefits with a $250 monthly employer contribution.
  • Health Savings Account with monthly employer contribution.
  • Family-building benefits through Kindbody.
  • Wellbeing benefits through One Medical, Rightway, and Headspace.
  • Daily lunch and snacks in the office.
  • Company retreats.