Staff Site Reliability Engineer - Release Engineering
Plaid is a financial technology company providing APIs and network connectivity for businesses to build financial products. Its platform supports bank-account linking, financial data access, identity verification, fraud and risk tools, credit underwriting, and bank payments.
Maintainer signals as of 8/12/2026
Funding history
Projects
About Plaid
Plaid operates a financial data network and API platform that lets businesses connect to financial institutions and build financial experiences. Its products support account and identity verification, real-time balance and transaction data, investment and liability data, income and underwriting workflows, fraud and AML risk checks, and multi-rail bank payments. It serves developers, businesses, financial institutions, platforms, lenders, banks, and consumer-facing financial-product providers.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define and scale reliability practices across product engineering. You will architect SLO and error-budget programs, promote progressive delivery and automated safety gates, guide teams toward production readiness, build self-service deployment features, lead critical incident response, and improve release safety for high-velocity development.
Requirements
- Over 8 years of professional experience in backend systems, SRE, or platform engineering
- Experience designing reliability programs such as service maturity models or SLI frameworks
- Experience building or operating canary rollout systems, metric-gated analysis, or automated rollback infrastructure
- Technical proficiency in software development
- Ability to drive organizational change without formal authority
- Technical judgment in high-stakes production scenarios
- Exposure to Kubernetes, service mesh technologies, Prometheus, or ArgoCD
Responsibilities
- Lead the expansion of reliability standards across product engineering
- Architect and manage SLO and error-budget frameworks
- Promote progressive delivery and automated safety gates
- Guide product teams toward production readiness
- Collaborate with Platform and Infrastructure teams on self-service platform features
- Direct critical incident response and post-mortem improvements
- Scale safety systems for increased code-change volume
Benefits
- Equity
