Search...

Site Reliability Engineer Telemetry

Payward, Inc. logo
Payward, Inc.

Payward, Inc. is a global financial infrastructure company and the parent organization behind Kraken. It provides trading, custody, payments, lending, staking, tokenized assets, derivatives, and market data infrastructure to consumers, professional traders, institutions, enterprises, fintechs, banks, exchanges, asset managers, and onchain platforms.

Recently fundedCompany intelligence
Distributed

Funding history

About Payward, Inc.

Payward, Inc. operates a unified financial infrastructure platform powering a portfolio of products including Kraken, NinjaTrader, Breakout, xStocks, CF Benchmarks, and Payward Services. Its shared architecture provides global liquidity, risk and margin management, collateral and settlement, compliance and licensing, and operational infrastructure across crypto, tokenized assets, and traditional markets. Through Payward Services, the company offers APIs and infrastructure for crypto trading, custody, on/off-ramps, tokenized equities, derivatives, staking and yield, payments, and benchmark data. Payward operates across more than 190 jurisdictions and serves consumers, professional traders, institutional investors, enterprises, fintechs, banks, exchanges, asset managers, and DeFi/onchain protocols.

View jobs by Payward, Inc.

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will operate and improve shared telemetry platforms for metrics, logs, traces, alerting, dashboards, and profiling. You will manage telemetry services and pipelines, troubleshoot production issues, build automation, participate in incident response and on-call, write runbooks, and improve reliability based on incidents.

Requirements

  • 3+ years of experience as a Site Reliability Engineer, Platform Engineer, Infrastructure Engineer, Observability Engineer, or similar
  • Production systems experience at scale
  • Prometheus or a Prometheus-compatible monitoring stack
  • Distributed systems troubleshooting
  • Infrastructure as Code
  • Terraform
  • CI/CD
  • Nomad, Kubernetes, or similar platforms
  • Scripting or programming ability
  • Incident response
  • Documentation and collaboration skills
  • VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry experience is a plus
  • PromQL or LogQL experience is a plus

Responsibilities

  • Operate and improve the shared telemetry platform
  • Maintain metrics collection, storage, querying, dashboards, and alerting
  • Operate log pipelines
  • Operate distributed tracing and profiling capabilities
  • Deploy and manage telemetry services with Terraform and Terragrunt
  • Troubleshoot missing data, slow queries, broken alerts, backpressure, and capacity issues
  • Build reusable configuration and automation
  • Participate in incident response and on-call
  • Write runbooks
  • Improve the platform using incident findings
Site Reliability Engineer Telemetry at Payward, Inc. | JobStash