Operations Engineer
Xsolla is a global video game commerce company providing tools and services to launch, monetize, and scale games. Its offerings include payments, web shops, publishing, distribution, LiveOps, anti-fraud, subscriptions, SDKs, and creator solutions for developers, publishers, payment providers, creators, and other gaming businesses.
Maintainer signals as of 8/23/2026
Projects
About Xsolla (USA), Inc.
Xsolla operates as a global merchant of record and video game commerce platform serving developers, publishers, resellers, payment providers, creators, and retailers. It provides payment processing across more than 200 countries and regions, 1,000+ payment methods, and 130+ currencies, alongside tax management, compliance, fraud prevention, refunds, dispute management, and end-user support. Its product portfolio includes Web Shop, Publishing Suite, Payments, Xsolla Pay, Mobile Buy Button, SDKs, Subscriptions, game distribution, Partner Network, Offerwall, LiveOps, Anti-Fraud, Login, Site Builder, cloud gaming, and related gaming commerce tools.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
Serve as the primary dashboard monitor during shifts, investigate and resolve production incidents, support major incident response, communicate with stakeholders, analyze incident trends, maintain post-incident documentation, build operational automation, develop runbooks, conduct shift handoffs, and publish critical application health reports.
Requirements
- 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment.
- Strong troubleshooting skills across logs, APM traces, infrastructure metrics, databases, and network paths.
- Experience with Datadog or an equivalent observability platform.
- Proficiency in Python, Go, or Bash.
- Clear written and verbal English communication skills.
- Working knowledge of Kubernetes and cloud infrastructure.
- Understanding of SLOs, error budgets, and burn-rate alerting.
- Experience with JIRA Service Management or JIRA, PagerDuty or OpsGenie, Slack, and Confluence.
- Experience with or strong interest in AI/ML-assisted operations.
- Comfort with 24x7 shift-based operations and rotating weekend on-call.
- Gaming, payments, or fintech experience is preferred.
- Familiarity with synthetic monitoring, RUM, databases, CI/CD, and deployment tooling is nice to have.
- JIRA Service Management administration experience or ITIL Foundation certification is nice to have.
Responsibilities
- Monitor the GTO Operational Dashboard in Datadog and detect anomalies.
- Triage, investigate, ticket, route, and resolve production incidents.
- Own lower-severity incidents end-to-end and execute runbook procedures.
- Support the TSO Lead during major incidents and maintain live incident timelines.
- Draft Slack, stakeholder, and customer-facing status page communications.
- Analyze incident trends, recurring issues, and production bugs.
- Compile incident timelines, draft PIR documents, and track action items.
- Build operational automation and contribute to runbook development.
- Conduct structured shift handoffs and knowledge transfer sessions.
- Cover for the TSO Lead during absences, including escalation decisions.
- Publish periodic health reports of critical applications.
