Senior Site Reliability Engineer

Kody is an agentic commerce platform with embedded payments, helping businesses manage payment acceptance, cash flow, customer engagement, and operations.

0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/2/2026

Singapore, SGP
About Kody

Kody provides an all-in-one commerce and payments platform for businesses with physical locations, particularly hospitality and other brick-and-mortar businesses. Its services include payment acceptance, financial and banking services through partners, cash-flow access via KodyCard, management tools, reporting and analytics, customer tracking, loyalty features, and conversational AI capabilities. Kody supports businesses through web- and app-based technology and offers multiple pricing plans for merchants.

View jobs by Kody

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own production observability, incident response, service-level management, and cloud infrastructure reliability for critical payment systems. You will lead SEV1 and SEV2 response, diagnose complex incidents, define SLOs and SLIs, improve monitoring and automation, conduct capacity planning and root-cause analysis, strengthen security and compliance, and mentor engineers.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure
  • Expertise in AWS, Kubernetes, EKS, Terraform, PostgreSQL, Redis, Kafka, Linux, networking, Datadog, Prometheus, and Grafana
  • Deep understanding of distributed systems, high availability, disaster recovery, capacity planning, and microservices orchestration
  • Experience operating payment, banking, fintech, or other highly regulated systems
  • Knowledge of SLO and SLI design, error budget management, alert governance, and toil reduction
  • Based in Hong Kong or Shenzhen
  • Excellent written and spoken English
  • Strong ownership, structured problem-solving, crisis management, and technical mentorship skills

Responsibilities

  • Participate in a follow-the-sun production on-call rotation
  • Lead SEV1 and SEV2 incident management
  • Diagnose, triage, mitigate, and coordinate complex production incidents
  • Define and maintain SLOs, SLIs, error budgets, and alerting standards
  • Drive infrastructure automation, observability improvements, capacity planning, and performance tuning
  • Conduct post-incident root-cause analysis
  • Strengthen architectural resilience, security posture, and operational maturity
  • Mentor junior engineers and reduce operational toil through automation
Senior Site Reliability Engineer at Kody | JobStash