Site Reliability Engineer Low Latency Trading Systems

Ondo helps institutions and individuals access traditional financial assets on blockchain through tokenized US Treasuries and investment products.

Series A0 current maintainers0 active leadsTeam intelligence

Maintainer signals as of 9/2/2026

Cayman Islands
About Ondo

Ondo brings traditional financial assets onchain through institutional-grade platforms and infrastructure. The protocol offers tokenized US Treasuries and investment products with daily yield distributions, while developing Ondo Chain, a Layer 1 blockchain optimized for real-world assets. Products include USDY for non-US investors and OUSG for qualified purchasers, supported by regulated custodians and audited smart contracts.

View jobs by Ondo

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own reliability, observability, and performance for real-time trading services. You will debug production incidents and latency regressions, operate multi-region Kubernetes infrastructure, harden market-data ingestion, build reconciliation tooling, improve deployment safety, and join an on-call rotation covering equity-market and crypto trading operations.

Requirements

  • 5+ years of SRE, production engineering, or infrastructure experience
  • Experience supporting real-time or latency-sensitive systems
  • Go or Rust programming
  • Kubernetes
  • AWS
  • Production experience with stateful latency-sensitive workloads
  • PromQL
  • Structured-log analysis
  • Alert design
  • Linux internals
  • Networking
  • Incident communication

Responsibilities

  • Own production reliability for trading engines, execution gateways, market-data ingestion, and PnL and reconciliation pipelines
  • Operate and evolve multi-region AWS EKS clusters using Flux and SOPS
  • Build observability through Prometheus metrics and alerting, Datadog logs and dashboards, and SLOs
  • Improve deployment safety through progressive rollouts, configuration reload behavior, and deployment guardrails
  • Debug stale feeds, rate limits, WebSocket disconnects, order-lifecycle desynchronization, and latency regressions
  • Harden market-data ingestion with staleness detection, failover, and replay
  • Build reconciliation and data-integrity tooling across gauges, Postgres, and S3 parquet data
  • Participate in on-call coverage for US equity-market hours and 24/7 crypto venues

Benefits

  • Medical, vision, and dental benefits
  • Flexible vacation policy