Site Reliability Engineer
Xsolla is a global video game commerce company providing tools and services to launch, monetize, and scale games. Its offerings include payments, web shops, publishing, distribution, LiveOps, anti-fraud, subscriptions, SDKs, and creator solutions for developers, publishers, payment providers, creators, and other gaming businesses.
Maintainer signals as of 8/23/2026
Projects
About Xsolla (USA), Inc.
Xsolla operates as a global merchant of record and video game commerce platform serving developers, publishers, resellers, payment providers, creators, and retailers. It provides payment processing across more than 200 countries and regions, 1,000+ payment methods, and 130+ currencies, alongside tax management, compliance, fraud prevention, refunds, dispute management, and end-user support. Its product portfolio includes Web Shop, Publishing Suite, Payments, Xsolla Pay, Mobile Buy Button, SDKs, Subscriptions, game distribution, Partner Network, Offerwall, LiveOps, Anti-Fraud, Login, Site Builder, cloud gaming, and related gaming commerce tools.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will own application infrastructure and service integrations, build observability with SLOs, SLIs, monitors, alerts, and dashboards, and improve CI/CD automation. You will plan capacity, lead production readiness reviews, support incident response, reduce operational toil, maintain reliability roadmaps, and contribute to architecture and operational standards.
Requirements
- 3+ years of SRE DevOps or platform engineering experience
- On-call or incident response experience
- Experience owning SLOs monitoring deployment pipelines and production infrastructure
- Experience building and shipping backend services
- Production-quality automation experience in Go PHP Python or Bash
- Hands-on Kubernetes experience with Helm manifests deployment strategies and debugging
- Observability experience with Datadog Prometheus Grafana or OpenTelemetry
- Infrastructure as Code experience with Terraform or Terragrunt
- GCP experience including IAM networking and managed services
- Experience maintaining GitLab CI or GitHub Actions pipelines
- Practical incident response and post-mortem experience
- Strong collaboration and communication skills
- Experience with payments fintech e-commerce or gaming systems
- Kubernetes certifications are a plus
- Google Cloud Platform certifications are a plus
- HashiCorp certifications are a plus
Responsibilities
- Own Helm charts Terraform configurations Kubernetes deployments and runtime configuration
- Design and implement SLOs SLIs monitors alerts and dashboards
- Evolve CI/CD pipelines with deployment and rollback automation
- Perform capacity planning performance tuning and load testing
- Run Production Readiness Reviews
- Investigate incidents and drive post-mortem improvements
- Maintain runbooks and build operational automation
- Drive the domain reliability roadmap
- Participate in planning refinements and architecture reviews
- Co-author company-wide SLO SLI capacity and operational standards
- Participate in the SRE duty rotation
Benefits
- 100% company-paid medical plans
- 100% company-paid dental plans
- 100% company-paid vision plans
- Unlimited flexible time off
