Site Reliability Engineer Telemetry
Payward, Inc. is a global financial infrastructure company and the parent organization behind Kraken. It provides trading, custody, payments, lending, staking, tokenized assets, derivatives, and market data infrastructure to consumers, professional traders, institutions, enterprises, fintechs, banks, exchanges, asset managers, and onchain platforms.
Funding history
Investors
Projects
About Payward, Inc.
Payward, Inc. operates a unified financial infrastructure platform powering a portfolio of products including Kraken, NinjaTrader, Breakout, xStocks, CF Benchmarks, and Payward Services. Its shared architecture provides global liquidity, risk and margin management, collateral and settlement, compliance and licensing, and operational infrastructure across crypto, tokenized assets, and traditional markets. Through Payward Services, the company offers APIs and infrastructure for crypto trading, custody, on/off-ramps, tokenized equities, derivatives, staking and yield, payments, and benchmark data. Payward operates across more than 190 jurisdictions and serves consumers, professional traders, institutional investors, enterprises, fintechs, banks, exchanges, asset managers, and DeFi/onchain protocols.
Skills
Candidate Availability
Required and preferred rules are kept separate and reflect the wording in the original posting.
About the Role
You will operate and improve shared telemetry platforms for metrics, logs, traces, alerting, dashboards, and profiling. You will manage telemetry services and pipelines, troubleshoot production issues, build automation, participate in incident response and on-call, write runbooks, and improve reliability based on incidents.
Requirements
- 3+ years of experience as a Site Reliability Engineer, Platform Engineer, Infrastructure Engineer, Observability Engineer, or similar
- Production systems experience at scale
- Prometheus or a Prometheus-compatible monitoring stack
- Distributed systems troubleshooting
- Infrastructure as Code
- Terraform
- CI/CD
- Nomad, Kubernetes, or similar platforms
- Scripting or programming ability
- Incident response
- Documentation and collaboration skills
- VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry experience is a plus
- PromQL or LogQL experience is a plus
Responsibilities
- Operate and improve the shared telemetry platform
- Maintain metrics collection, storage, querying, dashboards, and alerting
- Operate log pipelines
- Operate distributed tracing and profiling capabilities
- Deploy and manage telemetry services with Terraform and Terragrunt
- Troubleshoot missing data, slow queries, broken alerts, backpressure, and capacity issues
- Build reusable configuration and automation
- Participate in incident response and on-call
- Write runbooks
- Improve the platform using incident findings
