Senior Manager Site Reliability Engineering
Nium provides global real-time payments infrastructure for businesses, banks, fintechs, marketplaces, payroll providers, and travel companies. Its platform supports cross-border payouts, multi-currency accounts, foreign exchange, account verification, and physical, virtual, and stablecoin-backed card programs.
Maintainer signals as of 8/23/2026
About Nium Pte. Ltd.
Nium is a global payments infrastructure company that enables businesses and financial institutions to collect, hold, convert, and disburse funds across borders. Its offerings include multi-currency accounts, global payouts, foreign exchange, beneficiary account verification, physical and virtual card issuance, stablecoin-backed cards, the Nium Portal, and APIs for payment integration. Nium serves banks, money transfer operators, marketplaces, payroll providers, spend-management companies, creator-economy platforms, airlines, hotels, and online travel agencies.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will lead reliability engineers across multiple time zones and define reliability standards for critical payment and card services. You will manage incidents, observability, capacity planning, automation, disaster recovery, compliance practices, budgets, and production readiness across engineering.
Requirements
- 10 or more years of experience in software engineering, infrastructure, or site reliability engineering
- At least 4 years of people management or technical leadership experience
- Experience operating and scaling high-availability, transaction-heavy production systems
- Expertise with AWS, Kubernetes, container orchestration, and infrastructure as code
- Experience with observability tools such as Prometheus, Grafana, Datadog, ELK, or OpenSearch
- Experience defining and operationalizing SLOs and error budgets
- Understanding of distributed systems, databases, caching, messaging queues, and microservices
- Experience leading major incident response and effective postmortems
- Familiarity with PCI-DSS, SOC 2, and ISO 27001
- Excellent communication skills
- Metrics-driven engineering decision-making
- Experience with Python, Go, Bash, and CI/CD pipelines
- Experience with payment rails, card networks, banking integrations, hypergrowth, FinOps, or building SRE functions is advantageous
Responsibilities
- Lead and develop SRE and reliability engineering teams
- Define SLIs, SLOs, and error budgets for critical services
- Drive incident management, escalation, response, and postmortems
- Embed reliability and operational readiness into software development
- Build observability across metrics, logs, traces, and alerts
- Lead capacity planning and performance engineering
- Develop automation and self-healing systems
- Own disaster recovery, business continuity, and chaos engineering
- Ensure infrastructure practices meet security and regulatory requirements
- Manage reliability tooling, cloud costs, and staffing plans
- Report technical risk and system health to executives and the board
- Establish production readiness reviews, runbooks, and operational standards
Benefits
- Performance bonuses
- Sales commissions
- Equity for specific roles
- Recognition programs
- Medical coverage
- 24/7 employee assistance program
- Generous vacation programs including a year-end shutdown
- Hybrid work with 3 days per week in the office
- Role-specific training
- Internal workshops
- Learning stipend
- Company-wide social events
- Team bonding activities
- Happy hours
- Team offsites
