Senior Site Reliability Engineer APAC
Reap is a stablecoin-native financial infrastructure company providing business accounts, Visa corporate cards, global fiat payments, expense management, and embedded card and payment services.
Funding history
About Reap
Reap is a Hong Kong-founded fintech that bridges traditional finance and digital assets for global businesses. Its platform supports stablecoin-funded corporate cards, business payments settled in fiat, expense controls, treasury operations, and embedded card-issuing and payment infrastructure. Payward completed its acquisition of Reap on July 1, 2026; Reap continues operating as a standalone brand within the Payward ecosystem.
Skills
About the Role
You will define reliability standards, establish SLIs, SLOs, and error budgets, and participate in on-call and incident-response practices. You will expand infrastructure as code, build self-service and ephemeral environments, improve observability, operate regulated cloud systems, and embed security controls into the infrastructure layer.
Requirements
- SRE experience defining SLOs, operating error budgets, or building incident and postmortem processes
- Strong Linux, systems, networking, and administration fundamentals
- Deep Terraform experience including state management and module design
- Strong AWS experience across multi-account and multi-region environments
- Experience operating ECS, Fargate, Kubernetes, Lambda, SQS, and EventBridge
- Experience with GitOps, Argo CD, and GitHub Actions
- Experience with observability tooling
- Python, Go, or Bash automation experience
- Experience operating regulated systems under PCI DSS and financial compliance
- Experience managing production incidents
Responsibilities
- Define SLIs, SLOs, and error budgets with product and engineering teams
- Participate in on-call rotation and build incident response and blameless postmortem practices
- Bring legacy infrastructure under infrastructure as code coverage
- Consolidate Terraform modules, governance, and drift detection
- Automate account provisioning and environment setup
- Build self-service infrastructure and ephemeral production-like environments
- Implement logging, metrics, tracing, and actionable alerting
- Own cloud operations including uptime, failover, capacity, disaster recovery, and incident response
- Embed secrets management, least-privilege IAM, network segmentation, and compliance controls
- Partner with product and engineering teams on reliability work
