Site Reliability Engineer (Monetization)
Xsolla is a video game commerce company that provides a suite of tools and services—including merchant of record payment processing, tax management, fraud prevention, compliance, refunds, dispute management, and end-user support—to help game developers and publishers launch, grow, and monetize their games globally. It serves video game developers, publishers, and studios of all sizes across global and regional markets.
About Xsolla
Xsolla connects the tools, systems, payments, and web shops used by the video games industry, positioning itself as a global merchant of record supporting over 1,000 payment methods and a cumulative audience of 50 million, with transaction fees around 5%. Its services include tax management, fraud monitoring and prevention, global and regional regulatory compliance, refund and dispute management, and end-user payment support. Xsolla's product lineup includes the Xsolla SDK for native in-app payments on side-loaded apps and alternative app stores, a Buy Button enabling link-out purchases from iOS mobile games in the U.S., and Web Shop for building customized, direct-to-consumer game storefronts. The company works with major gaming industry partners and clients such as Mytona, Ubisoft, MARVEL SNAP, and others, and highlights partner success stories, industry events, and its own culture and hiring initiatives on its site.
Skills
About the Role
You will own application-level infrastructure and reliability for the Monetization domain. You will build and maintain Kubernetes deployments, observability, CI/CD automation, capacity plans, and production-readiness practices. You will investigate incidents, improve runbooks and reliability standards, and bring reliability considerations into product planning and architecture reviews.
Requirements
- 3+ years of SRE, DevOps, or platform engineering experience
- Software development background building and shipping backend services
- Kubernetes experience with Helm, manifests, deployment strategies, and application-level debugging
- Observability experience with monitors, dashboards, SLOs, and SLIs
- Infrastructure as Code experience with Terraform or Terragrunt
- GCP experience, including IAM, networking, and managed services
- CI/CD pipeline experience with GitLab CI or GitHub Actions
- Programming or scripting proficiency with Python, Go, or Bash
- Incident response and post-mortem experience
- Collaboration and communication skills
- Experience with payments, fintech, e-commerce, or gaming transactional systems
Responsibilities
- Own application-level infrastructure, including Helm charts, Terraform configurations, Kubernetes deployments, runtime configuration, networking, and integrations
- Design and implement SLOs, SLIs, monitors, alerts, and dashboards for critical services
- Set up and evolve CI/CD pipelines, including deployment and rollback automation
- Perform capacity planning, load testing, performance tuning, and regression investigations
- Run Production Readiness Reviews and define production-readiness standards
- Support incident response, post-mortems, reliability improvements, and runbook maintenance
- Build automation that reduces operational toil
- Maintain a reliability roadmap with product engineering leads
- Participate in product planning, refinements, and architecture reviews
- Co-author company-wide SLO, SLI, capacity, and operational standards
- Participate in the SRE duty rotation
