Senior Site Reliability Engineer, Observability
Ripple provides payments, custody and stablecoin solutions that help financial institutions integrate blockchain and digital assets.
Funding history
Investors
Projects
About Ripple
Ripple helps financial institutions transform global payments by providing blockchain-powered infrastructure for cross-border payments, digital asset custody, and stablecoin solutions. With it, users can enable instant settlements, reduce costs, and access new markets. The company was originally founded as OpenCoin in 2012 and rebranded to Ripple in 2015.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
This engineering-first role focuses on hands-on observability and reliability engineering across Azure and AWS environments, with coaching and consulting responsibilities. The role includes designing monitoring and alerting systems, defining reliability objectives, managing observability infrastructure, establishing incident management practices, and enabling engineering teams through workshops, documentation, and guidance.
Requirements
- 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering focused on observability and production operations.
- Ability to deliver hands-on engineering work while coaching and mentoring teams.
- Experience working in Agile/Scrum environments.
- Expert-level New Relic experience and strong NRQL proficiency.
- Deep understanding of structured logging, RED/USE metrics, distributed tracing, dashboards, and alert design.
- Expertise defining and implementing SLOs, SLIs, and error budgets.
- Hands-on experience with Incident.IO, PagerDuty, OpsGenie, or similar platforms.
- Experience designing incident response workflows, on-call rotations, escalation policies, and post-incident reviews.
- Ability to troubleshoot complex production issues using observability data.
- Strong Terraform experience and familiarity with IaC governance.
- Proficiency with PowerShell scripting.
- Strong Azure experience and working knowledge of AWS.
- Experience with Azure DevOps for CI/CD pipeline authoring and troubleshooting.
- Experience with Octopus Deploy.
- Comfort working across Windows and Linux server environments.
- Familiarity with Slack for operational workflows and incident communication.
- Experience with alert noise reduction and observability cost optimization is desired.
- Chaos engineering or game day experience is desired.
- VM-hosted SQL Server monitoring and performance optimization knowledge is desired.
- Familiarity with FinTech compliance requirements such as SOC 2 and ISO 27001 is desired.
- Python or Bash scripting experience is desired.
- Familiarity with Jira is desired.
Responsibilities
- Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS.
- Write NRQL queries for troubleshooting, analysis, and reporting.
- Define and implement SLOs, SLIs, and error budgets while coaching teams on reliability practices.
- Lead alert noise reduction and signal quality engineering.
- Optimize observability costs through log ingestion management and pipeline rules.
- Improve structured logging, metrics instrumentation, distributed tracing, and dashboard patterns with engineering teams.
- Develop and maintain Terraform infrastructure as code for monitoring and observability resources.
- Establish and enforce IaC governance standards.
- Author and troubleshoot Azure DevOps pipelines.
- Administer and configure Incident.IO, including alert routing, notification workflows, integrations, and runbooks.
- Build incident management foundations including postmortem processes, on-call rotations, escalation policies, severity classification, and response playbooks.
- Track MTTR, MTTD, incident frequency, and reliability trends.
- Respond to and debrief production incidents and facilitate post-incident reviews.
- Enable engineering teams through workshops, consultation, and hands-on guidance.
- Collaborate with the Subsystems Platform Team on self-service observability capabilities.
- Build team competency through documentation, training materials, and knowledge sharing.
Benefits
- Professional development budget
- Flexible in-office schedule
- Bi-weekly all-company meetings with leadership
- Team offsites, team bonding activities, and happy hours
- Equity
- Competitive benefits covering physical and mental healthcare, retirement, family forming, and family support
- Employee giving match
- Mobile phone stipend
- R&R days
- Generous wellness reimbursement and weekly onsite and virtual programming
- Generous vacation policy
- Industry-leading parental leave policies and family planning benefits
- Catered lunches and fully stocked kitchens with premium snacks and beverages
