Search...

Site Reliability Engineer

ASX Limited logo
ASX Limited

ASX Limited operates Australian financial markets, providing cash and derivatives trading, listings, clearing and settlement services, market data, and connectivity services for issuers, investors, participants, and data clients.

Distributed
About ASX Limited

ASX Limited operates financial markets in Australia, including cash equity and derivatives markets. Its services include listings, trading, clearing and settlement through platforms and entities such as ASX Clear, ASX Clear (Futures), ASX Settlement, and Austraclear. ASX also provides market data, connectivity and colocation services, issuer services, investor education, and related tools for issuers, investors, market participants, and financial-market businesses.

View jobs by ASX Limited

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will improve the reliability of market-critical services through automation, observability, resilience engineering and operational readiness. You will develop CI/CD and infrastructure-as-code solutions, support capacity and performance planning, manage incidents and changes, and provide rostered 24/7 on-call support.

Requirements

  • 5+ years of experience in a similar SRE role
  • Experience with incident response, post-incident review and problem management
  • Experience with production readiness, release readiness and operational acceptance
  • Experience with capacity, performance and resilience testing
  • Knowledge of observability design
  • Experience supporting high-availability distributed business-critical platforms
  • Scripting and automation skills using Python, PowerShell and shell scripting
  • AWS experience with EC2, S3, Lambda and RDS
  • Understanding of microservices and containerisation with Docker
  • Kubernetes administration experience
  • CI/CD pipeline experience
  • Experience with CloudWatch, Grafana, Prometheus and OpenTelemetry
  • Database operations experience with Oracle and/or Microsoft SQL Server
  • Linux and Unix administration and troubleshooting experience
  • Microsoft Windows Server 2019-2022 skills
  • Troubleshooting, problem-solving and root cause analysis skills
  • AWS certification at Associate level or above
  • Experience with distributed transactions, high availability and performance-critical systems
  • Networking troubleshooting across DNS, TLS, load balancers, firewalls and TCP/IP

Responsibilities

  • Reduce operational toil through automation
  • Improve observability across logs, metrics and traces
  • Support production readiness and non-functional testing
  • Drive resilience, capacity, incident learning and reliability improvement
  • Design and maintain CI/CD pipelines
  • Develop and maintain infrastructure as code
  • Automate operational tasks, deployments and service recovery
  • Conduct production readiness assessments
  • Support capacity planning and performance engineering
  • Validate failover and recovery and participate in resilience exercises
  • Lead or contribute to post-incident reviews
  • Design observability practices and actionable alerts
  • Provide rostered 24/7 on-call support
  • Perform weekend and after-hours installations and upgrades
  • Manage incidents, problems, releases and changes
  • Undertake business-as-usual team work

Benefits

  • Hybrid working
  • Flexible working arrangements