Senior Site Reliability Engineer LATAM
Reap is a stablecoin-native financial infrastructure company providing business accounts, Visa corporate cards, global fiat payments, expense management, and embedded card and payment services.
Funding history
About Reap
Reap is a Hong Kong-founded fintech that bridges traditional finance and digital assets for global businesses. Its platform supports stablecoin-funded corporate cards, business payments settled in fiat, expense controls, treasury operations, and embedded card-issuing and payment infrastructure. Payward completed its acquisition of Reap on July 1, 2026; Reap continues operating as a standalone brand within the Payward ecosystem.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will define reliability standards, establish SLIs, SLOs, and error budgets, and participate in on-call and incident-response practices. You will expand infrastructure as code, build self-service and ephemeral environments, improve observability, operate regulated cloud systems, and embed security controls into the infrastructure layer.
Requirements
- SRE experience defining SLOs, operating error budgets, or building incident and postmortem processes
- Strong Linux, systems, networking, and administration fundamentals
- Deep Terraform experience including state management and module design
- Strong AWS experience across multi-account and multi-region environments
- Experience operating ECS, Fargate, Kubernetes, Lambda, SQS, and EventBridge
- Experience with GitOps, Argo CD, and GitHub Actions
- Experience with observability tooling
- Python, Go, or Bash automation experience
- Experience operating regulated systems under PCI DSS and financial compliance
- Experience managing production incidents
Responsibilities
- Define SLIs, SLOs, and error budgets with product and engineering teams
- Participate in on-call rotation and build incident response and blameless postmortem practices
- Bring legacy infrastructure under infrastructure as code coverage
- Consolidate Terraform modules, governance, and drift detection
- Automate account provisioning and environment setup
- Build self-service infrastructure and ephemeral production-like environments
- Implement logging, metrics, tracing, and actionable alerting
- Own cloud operations including uptime, failover, capacity, disaster recovery, and incident response
- Embed secrets management, least-privilege IAM, network segmentation, and compliance controls
- Partner with product and engineering teams on reliability work
